Memory access method and device, medium and product
By predicting the future access probability of each process in a multi-core RISC-V processor and optimizing the memory replacement strategy, the problems of low memory utilization and low memory access efficiency of the processor are solved, significantly improving the overall performance.
Patent Information
- Application Number
- CN202510600505.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-12
AI Technical Summary
Multi-core RISC-V processors have problems with low memory utilization and low memory memory access efficiency in applications, resulting in low overall performance and inability to achieve the best performance of multi-core processors.
By predicting the future access probability of each replaceable process to the corresponding memory area in the system memory, determining the process to be replaced, and replacing its memory page table from the system memory to the hard disk to free up memory as the target process allocation. If it is predicted that the memory page table of the replaced process needs to be accessed, it will be replaced back to system memory from the hard disk.
The memory utilization and memory memory access efficiency of multi-core RISC-V processors are improved, thereby improving overall performance and avoiding the situation where the memory page table of the process to be replaced is accessed immediately after the hard disk.
Smart Images

Figure CN120104359A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of memory management, and in particular to a memory access method, device, medium and product. Background Art
[0002] Traditional processors mainly include X86 architecture and ARM (Advanced RISC Machine) architecture.
[0003] However, as various industries have increasingly urgent requirements for domestic processors, RISC-V (Reduced Instruction Set Computer V, the fifth-generation reduced instruction set computer) has the advantages of open source, independent controllability and flexibility, as well as the increasingly high single-core performance of RISC-V. The use of multi-core RISC-V processors as server CPUs (Central Processing Units) is becoming more and more widespread.
[0004] However, the single-core performance of RISC-V processors is still far behind that of X86 and ARM architecture processors. Therefore, when RISC-V is used as a server CPU, a multi-core RISC-V interconnection solution is required, such as 64 / 96 / 128 cores. However, multi-core RISC-V processors have the disadvantages of low memory utilization and low memory access efficiency in applications, resulting in low overall performance of multi-core RISC-V processors, which cannot bring out the best performance of multi-core processors.
[0005] It can be seen that how to improve the overall performance of multi-core RISC-V processors is a problem that technical personnel in this field need to solve. Summary of the invention
[0006] The purpose of the embodiments of the present invention is to provide a memory access method, device, medium and product, which improves the overall performance of the multi-core RISC-V processor by improving the memory utilization and memory access efficiency of the multi-core RISC-V processor. The specific scheme is as follows:
[0007] In a first aspect, the present invention provides a memory access method, which is applied to a multi-core RISC-V processor, comprising:
[0008] When allocating a memory area for a target process from the system memory, if the system memory is insufficient, predict the future access probability of each replaceable process to the corresponding memory area in the system memory; the replaceable process is a process that has a corresponding memory area in the system memory;
[0009] Determine the process to be replaced based on the process with the smallest future access probability among all replaceable processes;
[0010] Replacing the memory page table of the process to be replaced from the system memory to the hard disk to release the memory area corresponding to the process to be replaced in the system memory, and completing the memory area allocation for the target process based on the released system memory;
[0011] If it is predicted that the memory page table of the replaced process needs to be accessed, the memory area is reallocated for the replaced process, and the memory page table of the replaced process is replaced from the hard disk to the currently reallocated memory area to complete the access to the memory page table of the replaced process.
[0012] Optionally, the target process is a replaced process or an acquired new process.
[0013] Optionally, the memory access method of the present invention further includes:
[0014] Determine whether the system memory is sufficient based on the size relationship between the current remaining memory in the system memory and the preset memory threshold;
[0015] And / or, determining whether the system memory is sufficient based on a size relationship between the current remaining memory and the memory required by the target process.
[0016] Optionally, predicting the future access probability of each replaceable process to the corresponding memory area in the system memory includes:
[0017] Predicting the target instruction running time of each replaceable process; the target instruction running time is the running time between two adjacent preset instructions; wherein the preset instructions include load instructions and store instructions;
[0018] Based on the execution time of each target instruction, a future access probability of each replaceable process to a corresponding memory area in the system memory is determined.
[0019] Optionally, predict the target instruction execution time of each replaceable process, including:
[0020] Reading unexecuted instructions from the instruction cache of each replaceable process in sequence until a preset number of preset instructions are read;
[0021] Grouping each two adjacent preset instructions and the non-preset instructions between each two adjacent preset instructions into one group to obtain a plurality of groups of instructions; the number of the plurality of groups of instructions is one less than the preset number;
[0022] Count the number of non-preset instructions in each group of instructions;
[0023] Obtaining the time when the non-preset instructions in each group of instructions are executed to determine the time interval between each two adjacent non-preset instructions in each group of instructions;
[0024] The target instruction execution time of each replaceable process is predicted based on the number of non-preset instructions in each group of instructions and the time interval between each two adjacent non-preset instructions in each group of instructions.
[0025] Optionally, based on the number of non-preset instructions in each group of instructions and the time interval between each two adjacent non-preset instructions in each group of instructions, predicting the target instruction running time of each replaceable process includes:
[0026] Calculating an average value based on the number of non-preset instructions in each group of instructions to obtain an average number of instructions;
[0027] Determine a first average time interval between two adjacent non-preset instructions in each group of instructions based on the time interval between each two adjacent non-preset instructions in each group of instructions;
[0028] Calculating an average value of the first average time interval between two adjacent non-preset instructions in each group of instructions to obtain a second average time interval;
[0029] The target instruction execution time of each replaceable process is predicted according to the average number of instructions and the second average time interval.
[0030] Optionally, predicting the target instruction execution time of each replaceable process according to the average number of instructions and the second average time interval includes:
[0031] Obtaining the current target instruction parameter and the current target time parameter of each replaceable process;
[0032] The target instruction running time of each replaceable process is predicted according to the average number of instructions and the current target instruction parameter, the second average time interval and the current target time parameter.
[0033] Optionally, the process of determining the current target instruction parameter and the current target time parameter of each replaceable process includes:
[0034] When the system memory is sufficient, determining the current instruction parameter to be adjusted and the current time parameter to be adjusted of each replaceable process;
[0035] Obtaining the current average number of instructions and the current second average time interval of each replaceable process;
[0036] Determine the predicted instruction execution time of each replaceable process based on the current average number of instructions and the current instruction parameter to be adjusted, the current second average time interval and the current time parameter to be adjusted;
[0037] Determining whether to adjust the current instruction parameter to be adjusted and the current time parameter to be adjusted according to the predicted instruction execution time and the actual instruction execution time of each replaceable process;
[0038] If yes, then based on the current adjusted instruction parameter and the current adjusted time parameter, determine the new current to-be-adjusted instruction parameter and the new current to-be-adjusted time parameter, and jump again to the step of obtaining the current average number of instructions and the current second average time interval of each replaceable process;
[0039] If not, the current target command parameter and the current target time parameter are determined based on the current command parameter to be adjusted and the current time parameter to be adjusted.
[0040] Optionally, determining the current instruction parameter to be adjusted and the current time parameter to be adjusted of each replaceable process includes:
[0041] When the system memory changes from insufficient to sufficient, the current target instruction parameter and the current target time parameter of each replaceable process are determined as the current instruction parameter to be adjusted and the current time parameter to be adjusted of each replaceable process;
[0042] When the system memory is sufficient and the current time is the initial time, the current instruction parameter to be adjusted and the current time parameter to be adjusted of each replaceable process are determined based on the preset instruction initial value and the preset time initial value.
[0043] Optionally, determining whether to adjust the current instruction parameter to be adjusted and the current time parameter to be adjusted according to the predicted instruction execution time and the actual instruction execution time of each replaceable process includes:
[0044] Determine the time difference between the predicted instruction execution time and the actual instruction execution time of each replaceable process;
[0045] By judging whether the time difference satisfies the preset difference condition, it is determined whether to adjust the current instruction parameter to be adjusted and the current time parameter to be adjusted.
[0046] Optionally, based on the execution time of each target instruction, determining the probability of each replaceable process accessing a corresponding memory area in the system memory in the future includes:
[0047] Determining a final instruction execution time of each replaceable process based on a default priority coefficient and a target instruction execution time of each replaceable process;
[0048] Determine, according to the final instruction running time of each replaceable process, the future access probability of each replaceable process to the corresponding memory area in the system memory;
[0049] Among them, the default priority coefficient of the replaceable process is positively correlated with the target instruction execution time; the future access probability of the replaceable process to the corresponding memory area in the system memory is negatively correlated with the final instruction execution time.
[0050] Optionally, determining the final instruction execution time of each replaceable process based on the default priority coefficient and the target instruction execution time of each replaceable process includes:
[0051] When there is a user desired process among the replaceable processes, adjusting the default priority coefficient of the user desired process, and determining the final instruction running time of the user desired process based on the adjusted priority coefficient of the user desired process and the target instruction running time;
[0052] Determining final instruction execution times of the remaining replaceable processes based on default priority coefficients and target instruction execution times of the remaining replaceable processes;
[0053] The remaining replaceable processes are other processes among the replaceable processes except the process desired by the user.
[0054] Optionally, adjust the default priority coefficient of the user's desired process, including:
[0055] If the user's desired process is a process that is not desired to be determined as a process to be replaced, a preset lowering operation is performed on the default priority coefficient of the user's desired process;
[0056] If the user-desired process is a process that is expected to be determined as a process to be replaced, a preset increase operation is performed on the default priority coefficient of the user-desired process.
[0057] Optionally, the memory access method of the present invention further includes:
[0058] By determining whether the target preset instruction exists in the instruction cache of the replaced process, it is predicted whether the memory page table of the replaced process needs to be accessed;
[0059] Among them, the target preset instruction is the preset instruction with the earliest storage time in the instruction cache of the replaced process, the preset instructions include load instructions and store instructions, and the total number of non-preset instructions in the instruction cache of the replaced process whose storage time is earlier than the target preset instruction is the preset total number.
[0060] Optionally, the memory access method of the present invention further includes:
[0061] If it is predicted that the memory page tables of multiple replaced processes all need to be accessed, the memory page tables of the replaced processes are replaced from the hard disk to the system memory in descending order of process priority.
[0062] Optionally, the memory access method of the present invention further includes:
[0063] When any replaceable process executes a preset instruction to be executed, a target virtual address required to be used by any replaceable process is obtained from the preset instruction to be executed; wherein the preset instruction includes a load instruction and a store instruction;
[0064] Obtaining a target cache directory corresponding to a target virtual address from a target mapping table corresponding to any replaceable process; the target mapping table stores a mapping relationship between each virtual address required to be used by any replaceable process and each cache directory of any replaceable process;
[0065] Obtain a target physical address corresponding to the target virtual address from the target cache directory;
[0066] Based on the target physical address, access the target cache corresponding to the target cache directory;
[0067] Among them, the target cache is any cache in the system memory or the multi-level cache of any replaceable process; each cache directory of any replaceable process includes the memory page table of the system memory, the memory page table of each level of cache and the address translation backup buffer of each level of cache.
[0068] Optionally, the target mapping table is updated, including:
[0069] When the first cache updates data to the second cache, a virtual address of the first updated data is determined, and a cache directory corresponding to the virtual address of the first updated data in the target mapping table is updated from a memory page table of the first cache to a memory page table of the second cache;
[0070] When data is updated in the address translation lookaside buffer of any level 1 cache in the multi-level cache, a virtual address of second updated data is obtained, and a cache directory corresponding to the virtual address of the second updated data in the target mapping table is updated to the address translation lookaside buffer of any level 1 cache;
[0071] If the first cache is the system memory, the second cache is the last level cache of any replaceable process; if the first cache is the next level cache of any replaceable process, the second cache is the previous level cache of any replaceable process.
[0072] In a second aspect, the present invention provides an electronic device, comprising:
[0073] Memory for storing computer programs;
[0074] A processor is used to execute a computer program to implement the steps of the aforementioned memory access method.
[0075] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the aforementioned memory access method when the computer program is executed by a processor.
[0076] In a fourth aspect, the present invention provides a computer program product, comprising a computer program / instruction, which implements the steps of the aforementioned memory access method when executed by a processor.
[0077] In the present invention, when a multi-core RISC-V processor allocates a memory area for a target process from a system memory, if the system memory is insufficient, the future access probability of each replaceable process to the corresponding memory area in the system memory is predicted; the replaceable process is a process corresponding to a memory area in the system memory; based on the process with the smallest future access probability among the replaceable processes, a process to be replaced is determined; the memory page table of the process to be replaced is replaced from the system memory to the hard disk to release the memory area corresponding to the process to be replaced in the system memory, and the memory area allocation to the target process is completed based on the released system memory; if it is predicted that the memory page table of the replaced process needs to be accessed, the memory area is reallocated to the replaced process, and the memory page table of the replaced process is replaced from the hard disk to the currently reallocated memory area, so as to complete the access to the memory page table of the replaced process.
[0078] Beneficial effect: When allocating a memory area for a target process from the system memory, if the system memory is insufficient, the present invention predicts the future access probability of each subsequent replaceable process to the corresponding memory area in the system memory, and replaces the memory page table of the process to be replaced with the smallest future access probability from the system memory to the hard disk, thereby releasing a part of the memory to allocate the memory area for the target process. Compared with counting the active behaviors of each replaceable process in the past period of time, the present invention focuses on predicting the future access probability of each subsequent replaceable process, and selects the process that is least likely to be accessed as the process to be replaced, thereby avoiding the situation where the memory page table of the process to be replaced is accessed immediately after being replaced to the hard disk, thereby improving the memory utilization of the multi-core RISC-V processor, and further improving the overall performance of the multi-core RISC-V processor. Furthermore, the present invention determines in advance whether the memory page table of the replaced process needs to be accessed, thereby determining whether to reallocate the memory area for the replaced process from the system memory, and replaces the memory page table of the replaced process from the hard disk back to the system memory. Compared with performing the above operation when the memory page table of the replaced process already needs to be accessed, the efficiency of accessing the memory page table of the process can be significantly improved, that is, the memory access efficiency of the multi-core RISC-V processor is improved, thereby improving the overall performance of the multi-core RISC-V processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] In order to more clearly illustrate the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0080] Figure 1 An architecture diagram of a 64-core RISC-V processor provided in an embodiment of the present invention;
[0081] Figure 2 A schematic diagram of memory release provided by an embodiment of the present invention;
[0082] Figure 3 A schematic diagram of a memory access structure of a multi-core RISC-V processor provided in an embodiment of the present invention;
[0083] Figure 4 A virtual address query flow chart provided by an embodiment of the present invention;
[0084] Figure 5 A flow chart of a memory access method provided by an embodiment of the present invention;
[0085] Figure 6 An architecture diagram of another multi-core RISC-V processor provided in an embodiment of the present invention;
[0086] Figure 7 An internal structure diagram of an instruction information analysis module provided by an embodiment of the present invention;
[0087] Figure 8 An instruction execution flow chart provided for an embodiment of the present invention;
[0088] Fig. 9 A hardware structure diagram involved in updating a mapping table provided in an embodiment of the present invention;
[0089] Fig.10 A structural diagram of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0090] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0091] The terms "including" and "having" in the specification of the present invention and the above-mentioned drawings, as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but may include steps or units that are not listed.
[0092] In order to enable those skilled in the art to better understand the scheme of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0093] The single-core performance of RISC-V processors is relatively different from that of processors of X86 architecture and ARM architecture, so when RISC-V is used as a server CPU, a multi-core RISC-V interconnection solution is needed. However, multi-core RISC-V processors have the disadvantages of low memory utilization and low memory access efficiency in applications, resulting in low overall performance of multi-core RISC-V processors, and the best performance of multi-core processors cannot be brought into play. To this end, the present invention provides a memory access method, which improves the overall performance of multi-core RISC-V processors by improving the memory utilization and memory access efficiency of multi-core RISC-V processors.
[0094] by Figure 1 Taking the 64-core RISC-V processor shown as an example, each RISC-V Core corresponds to L1Cache (first-level cache), L2Cache (second-level cache), and 64 RISC-V Cores share L3Cache (third-level cache). Then the entire 64-core RISC-V CPU is connected to the system memory. Among them, the system memory can use DDR (Double Data Rate Synchronous Dynamic Random Access Memory), and the system memory is connected to the external hard disk. The hard disk contains a SWAP partition (replacement partition). This is also a typical three-layer storage structure, namely Cache—> system memory (DDR)—> hard disk.
[0095] Since multi-core RISC-V processors are used in the server field, they need to run a lot of software processes, and each software process requires a corresponding system memory area to run, so there is often a situation where the system memory is insufficient when a new system memory area needs to be applied for a new software process. The traditional solution is to use virtual memory technology, that is, to divide a part of the SWAP partition of the hard disk as temporary memory, and to judge the software process to replace the data (memory page table) of the least active software process in the past period of time from the system memory to the temporary memory of the SWAP partition, thereby saving some memory areas in the system memory to allocate system memory for the new software process.
[0096] However, the traditional solution has a major drawback. That is, when the system memory is insufficient, the traditional solution is based on statistics on the active behaviors of each software process in the past period of time, rather than predicting the active behaviors of each software process in the future. In other words, the traditional solution cannot solve the problem that the data of the replaced process must be accessed immediately after being replaced to the hard disk (in order to access the memory page table of the replaced process, the memory page table of the replaced process must be replaced from the hard disk back to the system memory before it can be accessed normally, that is, the replaced process can be executed normally), thereby affecting the execution efficiency of the replaced process and the memory access efficiency of the replaced process, and affecting the memory utilization of the multi-core RISC-V processor, thereby causing the overall performance of the multi-core RISC-V processor to decline.
[0097] like Figure 2 As shown in FIG. 1 , the most classic traditional solution is based on the LRU (Least Recently Used) memory elimination algorithm, the principle of which is: when the system memory is insufficient, the software process that has used the least memory recently is eliminated, that is, the memory page table of the software process is replaced from the system memory to the hard disk, which includes writing the memory page table of the software process to the hard disk and releasing the memory area corresponding to the software process in the system memory.
[0098] And, in such Figure 3 In the memory access structure of the multi-core RISC-V processor shown in the figure, the MMU (Memory Management Unit) is responsible for completing the conversion between the virtual address and the physical address of the process. The memory management unit contains the TLB (Translation Lookaside Buffer) of each level of cache. The TLB stores the correspondence between some virtual addresses and physical addresses of each level of cache. When the process needs to obtain memory data from the cache at each level, that is, when the process executes the Load instruction (load instruction) / Store instruction (store instruction), it will first query the TLB of the MMU. If it is not found, it will query the memory page table.
[0099] Specifically, Figure 4 As shown in the figure, after parsing the virtual address from the instruction to be executed by the process, the TLB of the first-level cache is first queried. If the physical address corresponding to the virtual address can be queried, the corresponding memory data is directly obtained from the first-level cache according to the physical address.
[0100] If the virtual address cannot be found in the TLB of the first-level cache, the memory page table of the first-level cache is queried. If the physical address corresponding to the virtual address can be found in the memory page table of the first-level cache, the corresponding memory data is obtained from the first-level cache according to the physical address.
[0101] If the data cannot be found in the memory page table of the first-level cache, it means that the data to be accessed is not in the first-level cache and the TLB of the second-level cache needs to be queried. If the physical address corresponding to the virtual address can be queried from the TLB of the second-level cache, the corresponding memory data in the second-level cache is updated to the first-level cache according to the physical address to obtain the memory data.
[0102] If the virtual address cannot be found in the TLB of the secondary cache, the memory page table of the secondary cache is queried. If the physical address corresponding to the virtual address can be found in the memory page table of the secondary cache, the corresponding memory data in the secondary cache is updated to the primary cache according to the physical address to obtain the memory data.
[0103] If the data cannot be found in the memory page table of the second-level cache, it means that the data to be accessed is not in the second-level cache and the TLB of the third-level cache needs to be queried. If the physical address corresponding to the virtual address can be queried from the TLB of the third-level cache, the corresponding memory data in the third-level cache is first updated to the second-level cache and then updated from the second-level cache to the first-level cache according to the physical address, thereby obtaining the memory data.
[0104] If the query cannot be found from the TLB of the third-level cache, the query is made to the memory page table of the third-level cache. If the physical address corresponding to the virtual address can be queried from the memory page table of the third-level cache, the corresponding memory data in the third-level cache is first updated to the second-level cache and then updated from the second-level cache to the first-level cache according to the physical address, thereby obtaining the memory data.
[0105] If the memory page table of the third-level cache cannot be queried, it means that the data to be accessed is not in the third-level cache, and the physical address corresponding to the virtual address is queried in the memory page table of the system memory (it can definitely be queried from the memory page table of the system memory, because when the relevant process applies for system memory, there must be a corresponding page table record in the system memory), and then the corresponding memory data in the system memory is updated to the third-level cache according to the physical address, then updated from the third-level cache to the second-level cache, and finally from the second-level cache to the first-level cache, so as to obtain the memory data.
[0106] But if Figure 4 The memory access method shown has a major drawback, that is, when the virtual address of the process cannot be hit from the first-level cache, a very complex access process and multiple interactions are required to correctly obtain the memory data, which can easily cause problems such as large memory access latency and low memory access efficiency, thereby causing the execution speed of the process to decrease, greatly reducing the overall performance of the multi-core RISC-V processor. For example,
[0107] If the memory data that the process wants to obtain is in the second-level cache, the multi-core RISC-V processor needs to query the TLB of the first-level cache once, the memory page table of the first-level cache once, the TLB of the second-level cache once, and the memory page table of the second-level cache once.
[0108] If the memory data that the process wants to obtain is in the third-level cache, the multi-core RISC-V processor needs to query the TLB of the first-level cache once, the memory page table of the first-level cache once, the TLB of the second-level cache once, the memory page table of the second-level cache once, the TLB of the third-level cache once, and the memory page table of the third-level cache once.
[0109] If the memory data to be obtained by the process is in the system memory, the multi-core RISC-V processor needs to query the TLB of the first-level cache once, the memory page table of the first-level cache once, the TLB of the second-level cache once, the memory page table of the second-level cache once, the TLB of the third-level cache once, the memory page table of the third-level cache once, and the memory page table of the system memory once.
[0110] Therefore, how to improve the overall performance of the multi-core RISC-V processor by improving the memory utilization and memory access efficiency of the multi-core RISC-V processor is a problem to be achieved and solved by the present invention.
[0111] See also Figure 5 As shown, an embodiment of the present invention provides a memory access method, which is applied to a multi-core RISC-V processor, including:
[0112] Step S11, when allocating a memory area for the target process from the system memory, if the system memory is insufficient, predict the future access probability of each replaceable process to the corresponding memory area in the system memory; the replaceable process is a process that has a corresponding memory area in the system memory.
[0113] The multi-core RISC-V processor proposed in the embodiment of the present invention adopts a single-core single-process, that is, each core in the multi-core RISC-V processor can only run one process at the same time. Accordingly, the multi-core RISC-V processor of the embodiment of the present invention can also be expanded to execute multiple processes on a single core through fast switching, thereby achieving single-core multi-process (at this time, the user feels the illusion that multiple processes are executed in parallel, but in essence, only one process is still running at the same time).
[0114] When the kernel of a multi-core RISC-V processor running a target process needs to allocate a memory area for the target process from the system memory, it first determines whether the system memory is sufficient. If the system memory is sufficient, the memory area can be directly allocated to the target process from the system memory. If the system memory is insufficient, it is necessary to first predict the future access probability of each subsequent replaceable process to the corresponding memory area in the system memory, that is, to predict the access probability of each replaceable process to the corresponding memory area in the system memory at the future moment.
[0115] Among them, a replaceable process refers to a process that has a corresponding memory area in the system memory; the future access probability of the replaceable process to the corresponding memory area in the system memory refers to the probability of accessing the memory area corresponding to the replaceable process in the system memory at a future time, or the possibility of accessing the memory area corresponding to the replaceable process in the system memory at a future time.
[0116] In addition, the system memory is the memory shared by each core in the multi-core RISC-V processor and each process running on the core; and the target process is a replaced process or a new process obtained. It should be noted that when the target process is a new process obtained by the multi-core RISC-V processor, an idle core is first allocated to the new process, and then the idle core allocates a memory area to the new process from the system memory. A replaced process refers to a process that has replaced its own memory page table from the system memory to the hard disk.
[0117] As for judging whether the system memory is sufficient, in a specific implementation, whether the system memory is sufficient can be determined based on the size relationship between the current remaining memory in the system memory and the preset memory threshold; that is, if the current remaining memory in the system memory is greater than the preset memory threshold, it is determined that the system memory is sufficient; if the current remaining memory in the system memory is less than or equal to the preset memory threshold, it is determined that the system memory is insufficient.
[0118] In another specific implementation, whether the system memory is sufficient can be determined based on the size relationship between the current remaining memory in the system memory and the memory required by the target process; that is, if the current remaining memory in the system memory is greater than the memory required by the target process, the system memory is determined to be sufficient, and if the current remaining memory in the system memory is less than or equal to the memory required by the target process, the system memory is determined to be insufficient. Of course, there are other ways to determine whether the system memory is sufficient, which will not be repeated here.
[0119] The prediction of the future access probability of each replaceable process to the corresponding memory area in the system memory can specifically include: predicting the target instruction running time of each replaceable process, wherein the target instruction running time is the running time between two adjacent preset instructions; based on the target instruction running time of each replaceable process, determining the future access probability of each replaceable process to the corresponding memory area in the system memory.
[0120] It should be noted that the preset instructions include load instructions and store instructions; both load instructions and store instructions are memory access instructions, which are used to complete corresponding operations in the memory access phase. Load instructions are used to read data from system memory / cache at all levels, and store instructions are used to write data to system memory / cache at all levels.
[0121] It should also be noted that the probability of a replaceable process accessing the corresponding memory area in the system memory in the future is negatively correlated with the target instruction execution time of the replaceable process. That is, the shorter the target instruction execution time of the replaceable process, the more active the replaceable process is, and accordingly, the greater the probability of the replaceable process accessing the corresponding memory area in the system memory in the future; on the contrary, the longer the target instruction execution time of the replaceable process, the less active the replaceable process is, and accordingly, the smaller the probability of the replaceable process accessing the corresponding memory area in the system memory in the future.
[0122] In the process of predicting the target instruction running time of each replaceable process, for each replaceable process in each replaceable process, first read the unexecuted instructions from the instruction cache of each replaceable process in sequence until a preset number of preset instructions are read; group each two adjacent preset instructions and the non-preset instructions between each two adjacent preset instructions into a group to obtain a plurality of groups of instructions; wherein the number of the plurality of groups of instructions is one less than the preset number; then count the number of non-preset instructions in each group of instructions, and obtain the time when the non-preset instructions in each group of instructions are executed to determine the time interval between each two adjacent non-preset instructions in each group of instructions; based on the number of non-preset instructions in each group of instructions and the time interval between each two adjacent non-preset instructions in each group of instructions, predict the target instruction running time of each replaceable process.
[0123] It should be noted that the instruction cache of each replaceable process is located in the first-level cache of each replaceable process, and the first-level cache of each replaceable process is divided into two parts, one part is the instruction cache (I-Cache, Instruction Cache), which is used to store unexecuted instructions, and the other part is the data cache (D-Cache, Data Cache), which is used to store recently used data. The second-level cache and third-level cache of each replaceable process generally do not need to distinguish between instruction cache and data cache.
[0124] Among them, the unexecuted instructions include preset instructions and non-preset instructions. The preset instructions include storage instructions and load instructions. The non-preset instructions include but are not limited to arithmetic operation instructions, such as addition instructions, subtraction instructions, multiplication instructions, etc., logical operation instructions, such as logical AND instructions, logical OR instructions, etc., and control flow instructions, such as branch instructions, etc.
[0125] Taking the preset number Q+1 (Q is a positive integer) as an example, unexecuted instructions are read from the instruction cache of each replaceable process in sequence until Q+1 preset instructions are read; the first preset instruction, the second preset instruction, and the non-preset instructions between the first preset instruction and the second preset instruction are divided into the first group; the second preset instruction, the third preset instruction, and the non-preset instructions between the second preset instruction and the third preset instruction are divided into the second group; and so on until the Qth preset instruction, the Q+1th preset instruction, and the non-preset instructions between the Qth preset instruction and the Q+1th preset instruction are divided into the Qth group, thereby obtaining Q groups of instructions.
[0126] Assuming that the number of non-preset instructions in a certain group of instructions is N+1 (N is a positive integer), the time when each non-preset instruction in the group of instructions is executed is obtained to determine the time interval T1 between the first non-preset instruction and the second non-preset instruction in the group of instructions, the time interval T2 between the second non-preset instruction and the third non-preset instruction, and so on to the time interval TN between the Nth non-preset instruction and the N+1th non-preset instruction, thereby obtaining N time intervals.
[0127] After obtaining the number of non-preset instructions in each group of instructions and the time interval between each adjacent two non-preset instructions in each group of instructions, an average value is calculated based on the number of non-preset instructions in each group of instructions to obtain an average number of instructions; based on the time interval between each adjacent two non-preset instructions in each group of instructions, a first average time interval between two adjacent non-preset instructions in each group of instructions is determined; an average value is calculated for the first average time interval between two adjacent non-preset instructions in each group of instructions to obtain a second average time interval; and based on the average number of instructions and the second average time interval, a target instruction running time of each replaceable process is predicted.
[0128] Assume that there are Q groups of instructions, and the number of non-preset instructions in each group of instructions is recorded as NUM1, NUM2, ..., NUMQ, respectively. In this case, the average number of instructions = (NUM1 + NUM2 + ... + NUMQ) / Q. Assume that the number of non-preset instructions in a certain group of instructions is N + 1 (N is a positive integer), then the first average time interval between two adjacent non-preset instructions in the group of instructions = (T1 + T2 + ... + TN) / N, and the second average time interval = (the first average time interval of the first group of instructions + the first average time interval of the second group of instructions + ... + the first average time interval of the Qth group of instructions) / Q.
[0129] After obtaining the average number of instructions and the second average time interval of each replaceable process, the current target instruction parameter and the current target time parameter of each replaceable process are obtained, and the target instruction running time of each replaceable process is predicted based on the average number of instructions and the current target instruction parameter of each replaceable process, the second average time interval and the current target time parameter.
[0130] In a specific implementation, a first product result between the average number of instructions of each replaceable process and the current target instruction parameter is determined, and a second product result between the second average time interval of each replaceable process and the current target time parameter is determined, and then the product result between the first product result and the second product result is determined as the target instruction execution time of each replaceable process. That is, the target instruction execution time of each replaceable process = average number of instructions × current target instruction parameter × second average time interval × current target time parameter.
[0131] It should be noted that the current target instruction parameter and the current target time parameter of each replaceable process are adaptive adjustment parameters, that is, the current target instruction parameter and the current target time parameter of each replaceable process support adaptive feedback adjustment. Specifically, the multi-core RISC-V processor adaptively adjusts the current target instruction parameter and the current target time parameter of each replaceable process when the system memory is sufficient, so as to predict the target instruction running time of each replaceable process when the system memory is insufficient.
[0132] Specifically, the process of determining the current target instruction parameters and the current target time parameters of each replaceable process includes: when the system memory is sufficient, determining the current instruction parameters to be adjusted and the current time parameters to be adjusted of each replaceable process, and obtaining the current average number of instructions and the current second average time interval of each replaceable process; based on the current average number of instructions and the current instruction parameters to be adjusted, the current second average time interval and the current time parameters to be adjusted of each replaceable process, determining the predicted instruction running time of each replaceable process; according to the predicted instruction running time and the actual instruction running time of each replaceable process, determining whether to adjust the current instruction parameters to be adjusted and the current time parameters to be adjusted. If adjustment is required, the current instruction parameters to be adjusted and the current time parameters to be adjusted are adjusted, and based on the current adjusted instruction parameters and the current adjusted time parameters, new current instruction parameters to be adjusted and new current time parameters to be adjusted are determined, and then the process jumps again to the step of obtaining the current average number of instructions and the current second average time interval of each replaceable process; if adjustment is not required, it means that the current instruction parameters to be adjusted and the current time parameters to be adjusted are already relatively accurate, and the current target instruction parameters and current target time parameters of each replaceable process can be determined based on the current instruction parameters to be adjusted and the current time parameters to be adjusted.
[0133] When the system memory is sufficient, the current instruction parameters to be adjusted and the current time parameters to be adjusted of each replaceable process are determined, including the following two situations: the first situation is that when the system memory changes from insufficient to sufficient, the current target instruction parameters and the current target time parameters of each replaceable process are determined as the current instruction parameters to be adjusted and the current time parameters to be adjusted of each replaceable process; the second situation is that when the system memory is sufficient and the current moment is the initial moment, based on the preset instruction initial value and the preset time initial value, the current instruction parameters to be adjusted and the current time parameters to be adjusted of each replaceable process are determined; wherein, the preset instruction initial value and the preset time initial value are generally set to one.
[0134] And according to the predicted instruction running time and the actual instruction running time of each replaceable process, determining whether to adjust the current instruction parameter to be adjusted and the current time parameter to be adjusted can specifically include: first determining the time difference between the predicted instruction running time and the actual instruction running time of each replaceable process; and then determining whether to adjust the current instruction parameter to be adjusted and the current time parameter to be adjusted by judging whether the time difference meets the preset difference condition.
[0135] Specifically, if the time difference satisfies the preset difference condition, the current instruction parameter to be adjusted and the current time parameter to be adjusted are not adjusted; if the time difference does not satisfy the preset difference condition, the current instruction parameter to be adjusted and the current time parameter to be adjusted are adjusted. The preset difference condition includes but is not limited to the time difference being less than the preset difference.
[0136] According to one specific example, if it is determined not to adjust the current instruction parameters to be adjusted and the current time parameters to be adjusted based on the predicted instruction running time and the actual instruction running time of each replaceable process, then the current instruction parameters to be adjusted and the current time parameters to be adjusted of each replaceable process are directly determined as the current target instruction parameters and the current target time parameters of each replaceable process.
[0137] Furthermore, after predicting the target instruction execution time of each replaceable process, the final instruction execution time of each replaceable process is determined based on the default priority coefficient and the target instruction execution time of each replaceable process; then, based on the final instruction execution time of each replaceable process, the future access probability of each replaceable process to the corresponding memory area in the system memory is determined.
[0138] Among them, the default priority coefficient of the replaceable process is positively correlated with the target instruction execution time; the future access probability of the replaceable process to the corresponding memory area in the system memory is negatively correlated with the final instruction execution time.
[0139] That is, the shorter the target instruction running time of the replaceable process, the more active the replaceable process is, the higher the priority of the replaceable process is, and the smaller the default priority coefficient of the replaceable process is. In this case, the final instruction running time determined based on the default priority coefficient of the replaceable process and the target instruction running time is also correspondingly smaller, and the probability of the replaceable process accessing the corresponding memory area in the system memory in the future will be greater. On the contrary, the longer the target instruction running time of the replaceable process is, the less active the replaceable process is, the lower the priority of the replaceable process is, and the larger the default priority coefficient of the replaceable process is. In this case, the final instruction running time determined based on the default priority coefficient of the replaceable process and the target instruction running time is also correspondingly larger, and the probability of the replaceable process accessing the corresponding memory area in the system memory in the future will be smaller.
[0140] According to one specific example, the final instruction execution time of the replaceable process = the target instruction execution time of the replaceable process × the default priority coefficient of the replaceable process; wherein the default priority coefficient may adopt a default setting or may be configured by the user.
[0141] However, the embodiments of the present invention take into account that the user can also set an expected process, referred to as the user expected process. The user expected process can be a process that is not expected to be determined as a process to be replaced, or a process that is expected to be determined as a process to be replaced. When there is a user expected process, the default priority coefficient of the user expected process can be adjusted so that the user expected process finally meets the user's expectations.
[0142] Specifically, when there is a user-expected process among each replaceable process, the default priority coefficient of the user-expected process is adjusted, and the final instruction running time of the user-expected process is determined based on the adjusted priority coefficient of the user-expected process and the target instruction running time; and for other processes except the user-expected process among each replaceable process, referred to as the remaining replaceable processes, the final instruction running time of the remaining replaceable processes is determined based on the default priority coefficients and the target instruction running time of the remaining replaceable processes.
[0143] In the process of adjusting the default priority coefficient of the user's desired process, it may specifically include: if the user's desired process is a process that is not expected to be determined as a process to be replaced, the default priority coefficient of the user's desired process is preset to be lowered, and accordingly, the final instruction running time of the user's desired process will be reduced, thereby increasing the probability that the user's desired process will access the corresponding memory area in the system memory in the future. If the user's desired process is a process that is expected to be determined as a process to be replaced, the default priority coefficient of the user's desired process is preset to be raised, and accordingly, the final instruction running time of the user's desired process will be increased, thereby decreasing the probability that the user's desired process will access the corresponding memory area in the system memory in the future.
[0144] Among them, the preset lowering operation includes but is not limited to adjusting the default priority coefficient of the user's desired process to a preset minimum coefficient that tends to be infinitely small; the preset raising operation includes but is not limited to adjusting the default priority coefficient of the user's desired process to a preset maximum coefficient that tends to be infinitely large.
[0145] According to the final instruction running time of each replaceable process, the future access probability of each replaceable process to the corresponding memory area in the system memory is determined, including the following two determination methods: the first determination method is to determine the time sum of the final instruction running time of each replaceable process, and determine the ratio of the final instruction running time of the replaceable process to the time sum, and then determine the difference between 1 and the ratio as the future access probability of the replaceable process to the corresponding memory area in the system memory; that is, the future access probability of the replaceable process to the corresponding memory area in the system memory = 1-final instruction running time of the replaceable process / time sum. The second determination method is to construct a monotonically decreasing function based on the negative correlation between the future access probability of the replaceable process to the corresponding memory area in the system memory and the final instruction running time, and substitute the final instruction running time of the replaceable process into the monotonically decreasing function to determine the future access probability of the replaceable process to the corresponding memory area in the system memory.
[0146] Step S12: Determine a process to be replaced based on the process with the smallest future access probability among all replaceable processes.
[0147] In the embodiment of the present invention, after determining the future access probability of each replaceable process to the corresponding memory area in the system memory, the process with the smallest future access probability is queried from the replaceable processes, and the process with the smallest future access probability is used as the process to be replaced.
[0148] It should be noted that when there are multiple processes with the lowest probability of future access among the replaceable processes, any one of the processes with the lowest probability of future access can be used as the process to be replaced, or candidate processes can be first determined from the processes with the lowest probability of future access, and the size of the memory area corresponding to the candidate process in the system memory plus the current remaining memory of the system memory must be greater than the memory required by the target process, and then any one of the candidate processes can be used as the process to be replaced.
[0149] Step S13: Replace the memory page table of the process to be replaced from the system memory to the hard disk to release the memory area corresponding to the process to be replaced in the system memory, and complete the memory area allocation to the target process based on the released system memory.
[0150] In an embodiment of the present invention, after determining the process to be replaced from various replaceable processes, the memory page table of the process to be replaced is replaced from the system memory to a specific area of the hard disk to release the memory area corresponding to the process to be replaced in the system memory, and then the memory area allocation for the target process can be completed based on the released system memory, that is, the memory area is allocated to the target process from the released system memory.
[0151] Step S14: If it is predicted that the memory page table of the replaced process needs to be accessed, the memory area is reallocated for the replaced process, and the memory page table of the replaced process is replaced from the hard disk to the currently reallocated memory area to complete the access to the memory page table of the replaced process.
[0152] The replaced process in the embodiment of the present invention is a process that has replaced its own memory page table from the system memory to the hard disk. If it is predicted that the memory page table of the replaced process is about to be accessed, a memory area is reallocated from the system memory for the replaced process, and the memory page table of the replaced process is replaced from the hard disk to the currently reallocated memory area, thereby completing the access to the memory page table of the replaced process and achieving normal execution of the replaced process.
[0153] Specifically, by determining whether there is a target preset instruction in the instruction cache of the replaced process, it is predicted whether the memory page table of the replaced process needs to be accessed; wherein, the target preset instruction is the preset instruction with the earliest storage time in the instruction cache of the replaced process, the preset instructions include load instructions and store instructions, and the total number of non-preset instructions in the instruction cache of the replaced process whose storage time is earlier than the target preset instruction is the preset total number.
[0154] That is, if the target preset instruction exists in the instruction cache of the replaced process, it is predicted that the memory page table of the replaced process needs to be accessed; if the target preset instruction does not exist in the instruction cache of the replaced process, it is predicted that the memory page table of the replaced process does not need to be accessed.
[0155] In other words, assuming the preset total number is P, by analyzing the unexecuted instructions in the instruction cache of the replaced process, and when it is found that there are P non-preset instructions that are about to execute the preset instructions, it is predicted that the memory page table of the replaced process needs to be accessed.
[0156] It should be noted that if it is predicted that the memory page tables of multiple replaced processes all need to be accessed, the memory page tables of the replaced processes are replaced from the hard disk to the system memory in order of process priority from high to low. The higher the process priority, the smaller the default priority coefficient of the process.
[0157] Beneficial effect: When allocating a memory area for a target process from the system memory, if the system memory is insufficient, the present invention predicts the future access probability of each subsequent replaceable process to the corresponding memory area in the system memory, and replaces the memory page table of the process to be replaced with the smallest future access probability from the system memory to the hard disk, thereby releasing a part of the memory to allocate the memory area for the target process. Compared with counting the active behaviors of each replaceable process in the past period of time, the present invention focuses on predicting the future access probability of each subsequent replaceable process, and selects the process that is least likely to be accessed as the process to be replaced, thereby avoiding the situation where the memory page table of the process to be replaced is accessed immediately after being replaced to the hard disk, thereby improving the memory utilization of the multi-core RISC-V processor, and further improving the overall performance of the multi-core RISC-V processor. Furthermore, the present invention determines in advance whether the memory page table of the replaced process needs to be accessed, thereby determining whether to reallocate the memory area for the replaced process from the system memory, and replaces the memory page table of the replaced process from the hard disk back to the system memory. Compared with performing the above operation when the memory page table of the replaced process already needs to be accessed, the efficiency of accessing the memory page table of the process can be significantly improved, that is, the memory access efficiency of the multi-core RISC-V processor is improved, thereby improving the overall performance of the multi-core RISC-V processor.
[0158] See also Figure 6 As shown, taking a multi-core RISC-V processor as a 64-core RISC-V processor and a three-layer storage structure of a three-level cache, system memory and hard disk as an example, a memory access method proposed in an embodiment of the present invention is described in detail.
[0159] First of all, in the hardware structure of the multi-core RISC-V processor, it is necessary to add an instruction information analysis module for each core, and add a replacement scheduling analysis module as a whole, and connect the instruction information analysis module of each core to the replacement scheduling analysis module, and the instruction information analysis module of each core is connected to each core and the first-level cache of each core.
[0160] Secondly, in the software logic of the multi-core RISC-V processor, when the kernel running the target process of the multi-core RISC-V processor needs to allocate a memory area for the target process from the system memory, if the system memory is insufficient, it is necessary to use the instruction information analysis module of each target kernel to predict the future access probability of the replaceable process to the corresponding memory area in the system memory; among which, the target kernel is the kernel running the replaceable process, and the multi-core RISC-V processor adopts a single-core single-process.
[0161] Take the instruction information analysis module of any target kernel as an example. Figure 7As shown, the interface module in the instruction information analysis module sequentially reads unexecuted instructions from the instruction cache of the first-level cache, and analyzes the instruction types of the read unexecuted instructions through the instruction parsing module in the instruction information analysis module, and the instruction types include preset instructions and non-preset instructions, until a preset number of preset instructions are read from the instruction cache of the first-level cache. The analysis and statistics module in the instruction information analysis module groups each adjacent two preset instructions and the non-preset instructions between each adjacent two preset instructions into a group to obtain a number of groups of instructions, and then counts the number of non-preset instructions in each group of instructions, and obtains the time when the non-preset instructions in each group of instructions are executed (the interface module determines the time when the non-preset instructions are executed by monitoring the value operation of the instruction fetch module in the kernel on the instruction cache), and determines the time interval between each adjacent two non-preset instructions in each group of instructions, and finally predicts the target instruction running time of the replaceable process based on the number of non-preset instructions in each group of instructions and the time interval between each adjacent two non-preset instructions in each group of instructions, and determines the future access probability of the replaceable process to the corresponding memory area in the system memory based on the target instruction running time of the replaceable process.
[0162] Furthermore, the instruction information analysis module of each target core transmits the predicted future access probability to the replacement scheduling analysis module. Accordingly, the replacement scheduling analysis module determines the minimum future access probability from each future access probability, and determines the process corresponding to the minimum future access probability as the process to be replaced, and then replaces the memory page table of the process to be replaced from the system memory to the replacement partition in the hard disk, thereby releasing the memory area corresponding to the process to be replaced in the system memory, and completes the memory area allocation for the target process based on the released system memory.
[0163] It should be noted that for the kernels running the replaced processes, the instruction information analysis modules of these kernels will still monitor the instruction cache in the first-level cache to analyze the unexecuted instructions in the instruction cache, and when it is found that there are P non-preset instructions that are about to execute the preset instructions, it is predicted that the memory page table of the replaced process will soon need to be accessed, and an indication signal is generated at this time, and the indication signal is transmitted to the replacement scheduling analysis module.
[0164] After receiving the upcoming use indication signal sent by a kernel, the replacement scheduling analysis module reallocates the memory area from the system memory for the replaced process running on the kernel, and replaces the memory page table of the replaced process from the replacement partition of the hard disk back to the system memory, thereby completing the access to the memory page table of the replaced process and realizing the normal execution of the replaced process.
[0165] Beneficial effect: When allocating a memory area for a target process from the system memory, if the system memory is insufficient, the present invention predicts the future access probability of each subsequent replaceable process to the corresponding memory area in the system memory, and replaces the memory page table of the process to be replaced with the smallest future access probability from the system memory to the hard disk, thereby releasing a part of the memory to allocate the memory area for the target process. Compared with counting the active behaviors of each replaceable process in the past period of time, the present invention focuses on predicting the future access probability of each subsequent replaceable process, and selects the process that is least likely to be accessed as the process to be replaced, thereby avoiding the situation where the memory page table of the process to be replaced is accessed immediately after being replaced to the hard disk, thereby improving the memory utilization of the multi-core RISC-V processor, and further improving the overall performance of the multi-core RISC-V processor. Furthermore, the present invention determines in advance whether the memory page table of the replaced process needs to be accessed, thereby determining whether to reallocate the memory area for the replaced process from the system memory, and replaces the memory page table of the replaced process from the hard disk back to the system memory. Compared with performing the above operation when the memory page table of the replaced process already needs to be accessed, the efficiency of accessing the memory page table of the process can be significantly improved, that is, the memory access efficiency of the multi-core RISC-V processor is improved, thereby improving the overall performance of the multi-core RISC-V processor.
[0166] See also Figure 8 As shown, an embodiment of the present invention provides a process for executing a preset instruction by a process, which is applied to a multi-core RISC-V processor, including:
[0167] Step S21, when any replaceable process executes a preset instruction to be executed, a target virtual address required to be used by any replaceable process is obtained from the preset instruction to be executed; wherein the preset instruction includes a load instruction and a store instruction.
[0168] In an embodiment of the present invention, when any replaceable process executes a preset instruction to be executed, the multi-core RISC-V processor parses the preset instruction to be executed to obtain a target virtual address required to be used by any replaceable process.
[0169] Step S22: Obtain the target cache directory corresponding to the target virtual address from the target mapping table corresponding to any replaceable process; the target mapping table stores the mapping relationship between each virtual address required to be used by any replaceable process and each cache directory of any replaceable process.
[0170] In the embodiment of the present invention, the target mapping table corresponding to any replaceable process is queried through the memory management unit to obtain the target cache directory corresponding to the target virtual address. Different replaceable processes correspond to different mapping tables, and the mapping table stores the mapping relationship between each virtual address required to be used by the replaceable process and each cache directory of the replaceable process.
[0171] Specifically, the mapping table may also store the mapping relationship between each virtual address required to be used by the replaceable process and the directory identifier of each cache directory of the replaceable process, and different cache directories correspond to different directory identifiers.
[0172] Step S23: Obtain a target physical address corresponding to the target virtual address from the target cache directory.
[0173] In the embodiment of the present invention, the corresponding relationship between each virtual address and each physical address is stored in the cache directory of the replaceable process. Therefore, by querying the target cache directory, the target physical address corresponding to the target virtual address can be obtained.
[0174] Step S24, based on the target physical address, access the target cache corresponding to the target cache directory; wherein the target cache is any cache in the system memory or the multi-level cache of any replaceable process; each cache directory of any replaceable process includes the memory page table of the system memory, the memory page table of each level of cache and the address translation backup buffer of each level of cache.
[0175] In an embodiment of the present invention, after obtaining the target physical address corresponding to the target virtual address, the target cache corresponding to the target cache directory is accessed based on the target physical address, that is, the memory space in the target cache corresponding to the target physical address is accessed, such as data writing or reading.
[0176] It should be noted that the target cache can be either the system memory or any cache in the multi-level cache of any replaceable process, and the multi-level cache can use a three-level cache. Accordingly, each cache directory of any replaceable process includes the memory page table of the system memory, the memory page table of each level of cache, and the address translation backup buffer of each level of cache.
[0177] Furthermore, the update process of the target mapping table corresponding to any replaceable process may specifically include: when the first cache updates data to the second cache, the virtual address of the first updated data is determined, and the cache directory corresponding to the virtual address of the first updated data in the target mapping table is updated from the memory page table of the first cache to the memory page table of the second cache; wherein, if the first cache is the system memory, the second cache is the last level cache of any replaceable process; if the first cache is the next level cache of any replaceable process, the second cache is the previous level cache of any replaceable process. When data is updated in the address translation backup buffer of any level cache in the multi-level cache, the virtual address of the second updated data is obtained, and the cache directory corresponding to the virtual address of the second updated data in the target mapping table is updated to the address translation backup buffer of any level cache.
[0178] It should be noted that when the first cache updates data to the second cache, it not only includes updating the data corresponding to a physical address in the first cache to the second cache, but also includes updating the memory page table of the first cache and the memory page table of the second cache, that is, updating the correspondence between the virtual address and the physical address in the memory page table.
[0179] Take the multi-level cache of any replaceable process as an example, using a three-level cache. Fig. 9 As shown, by adding a first-level update module, a second-level update module, a third-level update module and a mapping table update module in a multi-core RISC-V processor, the creation and update of the target mapping table corresponding to any replaceable process can be realized. It should be noted that, at the beginning, each virtual address required to be used by any replaceable process in the target mapping table corresponds to the memory page table of the system memory.
[0180] Among them, the third-level update module is connected to the system memory to determine the virtual address of the updated data when the system memory updates the data to the third-level cache, and updates the cache directory corresponding to the virtual address of the updated data in the target mapping table from the memory page table of the system memory to the memory page table of the third-level cache.
[0181] The second-level update module is connected to the third-level cache to determine the virtual address of the updated data when the third-level cache updates the data to the second-level cache, and updates the cache directory corresponding to the virtual address of the updated data in the target mapping table from the memory page table of the third-level cache to the memory page table of the second-level cache.
[0182] The first-level update module is connected to the second-level cache to determine the virtual address of the updated data when the second-level cache updates the data to the first-level cache, and updates the cache directory corresponding to the virtual address of the updated data in the target mapping table from the memory page table of the second-level cache to the memory page table of the first-level cache.
[0183] The mapping table update module is connected to the first-level update module, the second-level update module, the third-level update module and the address translation backup buffer of the first-level cache, the second-level cache and the third-level cache in the memory management unit, and is used to receive request information sent by the first-level update module, the second-level update module and the third-level update module to update the target mapping table, and monitor the address translation backup buffer of the first-level cache, the second-level cache and the third-level cache.
[0184] When the mapping table update module detects that data is updated in the address translation backup buffer of any first-level cache, it obtains the virtual address of the updated data and updates the cache directory corresponding to the virtual address of the updated data in the target mapping table to the address translation backup buffer of any first-level cache. In this way, the mapping table update module can have a real-time updated target mapping table.
[0185] Beneficial effect: The embodiment of the present invention queries the target mapping table corresponding to the replaceable process to obtain the target cache directory corresponding to the virtual address required to be used by the replaceable process, and then directly queries the physical address corresponding to the virtual address required to be used by the replaceable process in the target cache directory, and finally accesses the target cache corresponding to the target cache directory based on the queried physical address. Compared with the traditional solution that requires a more complicated access process and multiple data interactions, the embodiment of the present invention can greatly reduce the query steps for the physical address corresponding to the virtual address, thereby reducing the memory access steps, speeding up the memory access efficiency, and thus improving the overall performance of the multi-core RISC-V processor.
[0186] Furthermore, the present application also discloses an electronic device. Fig.10 This is a structural diagram of an electronic device according to an exemplary embodiment, and the content in the diagram cannot be regarded as any limitation on the scope of use of this application. The electronic device may specifically include: at least one processor 11, at least one memory 12, a power supply 13, a communication interface 14, an input / output interface 15, and a communication bus 16. Among them, the memory 12 is used to store a computer program, and the computer program is loaded and executed by the processor 11 to implement the relevant steps in the memory access method disclosed in any of the aforementioned embodiments. In addition, the electronic device in this embodiment may specifically be an electronic computer.
[0187] In this embodiment, the power supply 13 is used to provide working voltage for each hardware device on the electronic device; the communication interface 14 can create a data transmission channel between the electronic device and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 15 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0188] In addition, the memory 12, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon can include an operating system 121, a computer program 122, etc., and the storage method can be temporary storage or permanent storage.
[0189] The operating system 121 is used to manage and control various hardware devices on the electronic device and the computer program 122, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the memory access method performed by the electronic device disclosed in any of the aforementioned embodiments, the computer program 122 can further include a computer program that can be used to complete other specific tasks.
[0190] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein the computer program, when executed by a processor, implements the aforementioned disclosed memory access method. The specific steps of the method can refer to the corresponding contents disclosed in the aforementioned embodiments, and will not be repeated here.
[0191] Furthermore, the present application also discloses a computer program product, including a computer program / instruction; wherein the computer program / instruction, when executed by a processor, implements the aforementioned disclosed memory access method. For the specific steps of the method, reference may be made to the corresponding contents disclosed in the aforementioned embodiments, and no further description will be given here.
[0192] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0193] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0194] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0195] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0196] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A memory access method, characterized in that: Applicable to multi-core RISC-V processors, including: When allocating a memory area for a target process from a system memory, if the system memory is insufficient, predicting the future access probability of each replaceable process to a corresponding memory area in the system memory; the replaceable process is a process having a corresponding memory area in the system memory; Determining a process to be replaced based on the process with the smallest future access probability among the replaceable processes; Replacing the memory page table of the process to be replaced from the system memory to the hard disk to release the memory area corresponding to the process to be replaced in the system memory, and completing the memory area allocation to the target process based on the released system memory; If it is predicted that the memory page table of the replaced process needs to be accessed, the memory area is reallocated for the replaced process, and the memory page table of the replaced process is replaced from the hard disk to the currently reallocated memory area to complete the access to the memory page table of the replaced process.
2. The memory access method according to claim 1, characterized in that: The target process is the replaced process or the acquired new process.
3. The memory access method according to claim 1, characterized in that: Also includes: Determining whether the system memory is sufficient based on a size relationship between the current remaining memory in the system memory and a preset memory threshold; And / or, determining whether the system memory is sufficient based on a size relationship between the current remaining memory and the memory required by the target process.
4. The memory access method according to claim 1, characterized in that: The predicting of the future access probability of each replaceable process to the corresponding memory area in the system memory includes: Predicting the target instruction running time of each replaceable process; the target instruction running time is the running time between two adjacent preset instructions; wherein the preset instructions include load instructions and store instructions; Based on the execution time of each target instruction, a future access probability of each replaceable process to a corresponding memory area in the system memory is determined.
5. The memory access method according to claim 4, characterized in that: The predicting of the target instruction execution time of each replaceable process includes: Reading unexecuted instructions from the instruction cache of each replaceable process in sequence until a preset number of preset instructions are read; Grouping each two adjacent preset instructions and the non-preset instructions between each two adjacent preset instructions into one group to obtain a plurality of groups of instructions; the number of the plurality of groups of instructions is one less than the preset number; Count the number of non-preset instructions in each group of instructions; Acquire the time when the non-preset instructions in each group of instructions are executed to determine the time interval between each two adjacent non-preset instructions in each group of instructions; The target instruction execution time of each replaceable process is predicted based on the number of non-preset instructions in each group of instructions and the time interval between each two adjacent non-preset instructions in each group of instructions.
6. The memory access method according to claim 5, characterized in that: The predicting the target instruction running time of each replaceable process based on the number of non-preset instructions in each group of instructions and the time interval between each two adjacent non-preset instructions in each group of instructions includes: Calculating an average value based on the number of non-preset instructions in each group of instructions to obtain an average number of instructions; Determine a first average time interval between two adjacent non-preset instructions in each group of instructions based on the time interval between each two adjacent non-preset instructions in each group of instructions; Calculating an average value of the first average time interval between two adjacent non-preset instructions in each group of instructions to obtain a second average time interval; The target instruction execution time of each replaceable process is predicted according to the average number of instructions and the second average time interval.
7. The memory access method according to claim 6, characterized in that: The step of predicting the target instruction execution time of each replaceable process according to the average number of instructions and the second average time interval includes: Obtaining current target instruction parameters and current target time parameters of each replaceable process; The target instruction execution time of each replaceable process is predicted according to the average number of instructions and the current target instruction parameter, the second average time interval and the current target time parameter.
8. The memory access method according to claim 7, characterized in that: The process of determining the current target instruction parameters and the current target time parameters of each replaceable process includes: When the system memory is sufficient, determining the current instruction parameter to be adjusted and the current time parameter to be adjusted of each replaceable process; Obtaining the current average number of instructions and the current second average time interval of each replaceable process; Determining a predicted instruction execution time of each replaceable process based on the current average number of instructions and the current instruction parameter to be adjusted, the current second average time interval and the current time parameter to be adjusted; Determining whether to adjust the current instruction parameter to be adjusted and the current time parameter to be adjusted according to the predicted instruction execution time and the actual instruction execution time of each replaceable process; If yes, then based on the current adjusted instruction parameter and the current adjusted time parameter, determine the new current instruction parameter to be adjusted and the new current time parameter to be adjusted, and jump again to the step of obtaining the current average number of instructions and the current second average time interval of each replaceable process; If not, the current target command parameter and the current target time parameter are determined based on the current command parameter to be adjusted and the current time parameter to be adjusted.
9. The memory access method according to claim 8, characterized in that: The determining of the current instruction parameter to be adjusted and the current time parameter to be adjusted of each replaceable process includes: When the system memory changes from insufficient to sufficient, determining the current target instruction parameter and the current target time parameter of each replaceable process as the current instruction parameter to be adjusted and the current time parameter to be adjusted of each replaceable process; When the system memory is sufficient and the current moment is the initial moment, the current instruction parameter to be adjusted and the current time parameter to be adjusted of each replaceable process are determined based on the preset instruction initial value and the preset time initial value.
10. The memory access method according to claim 8, characterized in that: The determining whether to adjust the current instruction parameter to be adjusted and the current time parameter to be adjusted according to the predicted instruction execution time and the actual instruction execution time of each replaceable process includes: Determine the time difference between the predicted instruction execution time and the actual instruction execution time of each replaceable process; Whether to adjust the current instruction parameter to be adjusted and the current time parameter to be adjusted is determined by judging whether the time difference satisfies a preset difference condition.
11. The memory access method according to claim 4, characterized in that: The determining, based on the execution time of each target instruction, the probability of each replaceable process accessing the corresponding memory area in the system memory in the future includes: Determining a final instruction execution time of each of the replaceable processes based on a default priority coefficient of each of the replaceable processes and the target instruction execution time; Determining, according to the final instruction execution time of each replaceable process, a future access probability of each replaceable process to a corresponding memory area in the system memory; Among them, the default priority coefficient of the replaceable process is positively correlated with the target instruction execution time; the future access probability of the replaceable process to the corresponding memory area in the system memory is negatively correlated with the final instruction execution time.
12. The memory access method according to claim 11, characterized in that: The determining the final instruction execution time of each replaceable process based on the default priority coefficient of each replaceable process and the target instruction execution time includes: When there is a user desired process in each of the replaceable processes, adjusting the default priority coefficient of the user desired process, and determining the final instruction execution time of the user desired process based on the adjusted priority coefficient of the user desired process and the target instruction execution time; Determining a final instruction execution time of the remaining replaceable processes based on the default priority coefficients of the remaining replaceable processes and the target instruction execution time; The remaining replaceable processes are other processes among the replaceable processes except the user desired process.
13. The memory access method according to claim 12, characterized in that: The adjusting the default priority coefficient of the process desired by the user includes: If the user desired process is a process that is not desired to be determined as the process to be replaced, a preset lowering operation is performed on the default priority coefficient of the user desired process; If the user-desired process is a process expected to be determined as the process to be replaced, a preset increase operation is performed on the default priority coefficient of the user-desired process.
14. The memory access method according to claim 1, characterized in that: Also includes: By determining whether there is a target preset instruction in the instruction cache of the replaced process, it is pre-determined whether the memory page table of the replaced process needs to be accessed; Among them, the target preset instruction is the preset instruction with the earliest storage time in the instruction cache of the replaced process, the preset instructions include load instructions and storage instructions, and the total number of non-preset instructions in the instruction cache of the replaced process whose storage time is earlier than the target preset instruction is the preset total number.
15. The memory access method according to claim 14, characterized in that: Also includes: If it is predicted that the memory page tables of the multiple replaced processes all need to be accessed, the memory page tables of the replaced processes are replaced from the hard disk to the system memory in order of process priority from high to low.
16. The memory access method according to any one of claims 1 to 15, characterized in that: Also includes: When any replaceable process executes a preset instruction to be executed, a target virtual address required to be used by any replaceable process is obtained from the preset instruction to be executed; wherein the preset instruction includes a load instruction and a store instruction; Obtaining a target cache directory corresponding to the target virtual address from a target mapping table corresponding to any replaceable process; the target mapping table stores a mapping relationship between each virtual address required to be used by any replaceable process and each cache directory of any replaceable process; Acquire a target physical address corresponding to the target virtual address from the target cache directory; Based on the target physical address, access the target cache corresponding to the target cache directory; Among them, the target cache is any cache in the multi-level cache of the system memory or any replaceable process; each cache directory of any replaceable process includes the memory page table of the system memory, the memory page tables of each level of cache and the address translation backup buffer of each level of cache.
17. The memory access method according to claim 16, characterized in that: The updating process of the target mapping table includes: When the first cache updates data to the second cache, determining a virtual address of the first updated data, and updating a cache directory corresponding to the virtual address of the first updated data in the target mapping table from a memory page table of the first cache to a memory page table of the second cache; When data is updated in the address translation lookaside buffer of any one level cache in the multi-level cache, a virtual address of second updated data is obtained, and a cache directory corresponding to the virtual address of the second updated data in the target mapping table is updated to the address translation lookaside buffer of any one level cache; Among them, if the first cache is the system memory, the second cache is the last level cache of any replaceable process; if the first cache is the next level cache of any replaceable process, the second cache is the previous level cache of any replaceable process.
18. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to execute the computer program to implement the steps of the memory access method as claimed in any one of claims 1 to 17.
19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the memory access method according to any one of claims 1 to 17 are implemented.
20. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the memory access method described in any one of claims 1 to 17 are implemented.
Citation Information
Patent Citations
Memory optimization processing method and device
CN111177024A
Memory page replacement method, system and equipment and computer readable storage medium
CN112764929A
Memory management method and electronic equipment
CN116860429A
Memory access optimization method and device, equipment, medium and program product
CN118051189A
Memory management method and device, storage medium and terminal
CN118245398A
Cited By
Translation lookaside buffer table item processing method and system, medium and product
CN121833555A
A method, system, medium and product for processing a translation lookaside buffer entry
CN121833555B