A memory access method, device, medium and product

By predicting the future access probability of the permutable processes in a multi-core RISC-V processor and optimizing memory allocation and access, the problems of low memory utilization and memory access efficiency are solved, and the overall performance of the processor is improved.

CN120104359BActive Publication Date: 2025-07-11SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510600505.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-11
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

RISC-V processors have problems with low memory utilization and low memory memory access efficiency in server CPU applications, resulting in low overall performance.

Method used

By predicting the future access probability of the permutable process in a multi-core RISC-V processor, select the process with the smallest access in the future to replace its memory page table to the hard disk, free up system memory, and replace the memory page table of the permuted process from the hard disk back to system memory when needed, optimizing memory allocation and access.

Benefits of technology

Improves the memory utilization and memory memory access efficiency of multi-core RISC-V processors, thereby improving overall performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104359B_ABST
    Figure CN120104359B_ABST
Patent Text Reader

Abstract

The present invention discloses a memory access method, device, medium and product, relating to the field of memory management, including: when a multi-core RISC-V processor allocates a memory area for a target process from the system memory, if the system memory is insufficient, it predicts the future access probabilities of each replaceable process to the corresponding memory area in the system memory, and takes the process with the smallest future access probability as the process to be replaced; replaces the memory page table of the process to be replaced from the system memory to the hard disk to release the corresponding memory area in the system memory, and completes the allocation of the memory area for the target process based on the released system memory; if it is predicted that the memory page table of the replaced process needs to be accessed, re-allocates a memory area for the replaced process, and replaces the memory page table of the replaced process from the hard disk to the currently re-allocated memory area to complete the access to the memory page table of the replaced process. This application can improve the overall performance of the multi-core RISC-V processor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of memory management, and in particular, to a memory access method, device, medium, and product. Background Art

[0002] Traditional processors mainly include the X86 architecture and the ARM (Advanced RISC Machine) architecture.

[0003] However, as the requirements for domestic processors in various industries become more urgent, RISC-V (Reduced Instruction Set Computer V) has the advantages of being open-source, autonomously controllable, and flexible, and the single-core performance of RISC-V is getting higher and higher. The application of using multi-core RISC-V processors as server CPUs (Central Processing Units) is becoming more and more widespread.

[0004] However, compared with the processors of the X86 architecture and the ARM architecture, the single-core performance of the RISC-V processor still has a large gap. Therefore, when using RISC-V as a server CPU, a multi-core RISC-V interconnection scheme is required, such as 64 / 96 / 128 cores, etc. However, multi-core RISC-V processors have the disadvantages of low memory utilization and low memory access efficiency in applications, resulting in low overall performance of multi-core RISC-V processors and unable to exert the best performance of multi-core processors.

[0005] It can be seen that how to improve the overall performance of multi-core RISC-V processors is a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0006] The purpose of the embodiments of the present invention is to provide a memory access method, device, medium, and product, which improve the memory utilization and memory access efficiency of multi-core RISC-V processors, thereby improving the overall performance of multi-core RISC-V processors. The specific solutions are as follows:

[0007] In a first aspect, the present invention provides a memory access method, which is applied to a multi-core RISC-V processor and includes:

[0008] When allocating a memory area for a target process from the system memory, if the system memory is insufficient, predict the future access probabilities of each replaceable process for the corresponding memory area in the system memory; a replaceable process is a process that has a corresponding memory area in the system memory;

[0009] Based on the process with the smallest future access probability among the replaceable processes, determine the process to be replaced;

[0010] Replace the memory page table of the process to be replaced from the system memory to the hard disk to release the corresponding memory area of the process to be replaced in the system memory, and complete the memory area allocation for the target process based on the released system memory;

[0011] If it is predicted that the memory page table of the replaced process needs to be accessed, re-allocate a memory area for the replaced process, and replace the memory page table of the replaced process from the hard disk to the currently re-allocated memory area to complete the access to the memory page table of the replaced process.

[0012] Optionally, the target process is a replaced process or a newly obtained process.

[0013] Optionally, the memory access method of the present invention further includes:

[0014] Determine whether the system memory is sufficient based on the size relationship between the current remaining memory in the system memory and a preset memory threshold;

[0015] And / or, determine whether the system memory is sufficient based on the size relationship between the current remaining memory and the memory required by the target process.

[0016] Optionally, predicting the future access probability of each replaceable process to the corresponding memory area in the system memory includes:

[0017] Predict the target instruction running time of each replaceable process; the target instruction running time is the running time between two adjacent preset instructions; wherein, the preset instructions include load instructions and store instructions;

[0018] Based on the target instruction running time of each, determine the future access probability of each replaceable process to the corresponding memory area in the system memory.

[0019] Optionally, predicting the target instruction running time of each replaceable process includes:

[0020] Sequentially read the unexecuted instructions from the instruction cache of each replaceable process until a preset number of preset instructions are read;

[0021] Group every two adjacent preset instructions and the non-preset instructions between every two adjacent preset instructions to obtain several groups of instructions; the number of groups of several groups of instructions is one less than the preset number;

[0022] Count the number of non-preset instructions in each group of instructions;

[0023] Obtain the time when the non-preset instructions in each group of instructions are executed to determine the time interval between every two adjacent non-preset instructions in each group of instructions;

[0024] Predict the target instruction running time of each replaceable process based on the number of non - preset instructions in each group of instructions and the time interval between every two adjacent non - preset instructions in each group of instructions.

[0025] Optionally, predicting the target instruction running time of each replaceable process based on the number of non - preset instructions in each group of instructions and the time interval between every two adjacent non - preset instructions in each group of instructions includes:

[0026] Calculate the average value based on the number of non - preset instructions in each group of instructions to obtain the average number of instructions;

[0027] Based on the time interval between every two adjacent non - preset instructions in each group of instructions, determine the first average time interval between two adjacent non - preset instructions in each group of instructions;

[0028] Calculate the average value of the first average time interval between two adjacent non - preset instructions in each group of instructions to obtain the second average time interval;

[0029] Predict the target instruction running time of each replaceable process according to the average number of instructions and the second average time interval.

[0030] Optionally, predicting the target instruction running time of each replaceable process according to the average number of instructions and the second average time interval includes:

[0031] Obtain the current target instruction parameters and current target time parameters of each replaceable process;

[0032] Predict the target instruction running time of each replaceable process according to the average number of instructions and the current target instruction parameters, the second average time interval and the current target time parameters.

[0033] Optionally, the determination process of the current target instruction parameters and current target time parameters of each replaceable process includes:

[0034] When the system memory is sufficient, determine the current instruction parameters to be adjusted and the current time parameters to be adjusted of each replaceable process;

[0035] Obtain the current average number of instructions and the current second average time interval of each replaceable process;

[0036] Based on the current average number of instructions and the current instruction parameters to be adjusted, the current second average time interval and the current time parameters to be adjusted, determine the predicted instruction running time of each replaceable process;

[0037] According to the predicted instruction running time and the actual instruction running time of each replaceable process, determine whether to adjust the current instruction parameters to be adjusted and the current time parameters to be adjusted;

[0038] If so, based on the current adjusted instruction parameter and the current adjusted time parameter, determine a new current instruction parameter to be adjusted and a new current time parameter to be adjusted, and then jump back to the step of obtaining the current average instruction count and the current second average time interval of each replaceable process;

[0039] If not, based on the current instruction parameter to be adjusted and the current time parameter to be adjusted, determine the current target instruction parameter and the current target time parameter.

[0040] Optionally, determining the current instruction parameter to be adjusted and the current time parameter to be adjusted for each replaceable process includes:

[0041] When the system memory changes from insufficient to sufficient, determine the current target instruction parameter and the current target time parameter of each replaceable process as the current instruction parameter to be adjusted and the current time parameter to be adjusted for each replaceable process;

[0042] When the system memory is sufficient and the current moment is the initial moment, based on the preset initial instruction value and the preset initial time value, determine the current instruction parameter to be adjusted and the current time parameter to be adjusted for each replaceable process.

[0043] Optionally, determining whether to adjust the current instruction parameter to be adjusted and the current time parameter to be adjusted according to the predicted instruction running time and the actual instruction running time of each replaceable process includes:

[0044] Determine the time difference between the predicted instruction running time and the actual instruction running time of each replaceable process;

[0045] By determining whether the time difference meets the preset difference condition, determine whether to adjust the current instruction parameter to be adjusted and the current time parameter to be adjusted.

[0046] Optionally, based on the target instruction running times, determine the future access probabilities of each replaceable process for the corresponding memory areas in the system memory, including:

[0047] Based on the default priority coefficients and target instruction running times of each replaceable process, determine the final instruction running times of each replaceable process;

[0048] According to the final instruction running times of each replaceable process, determine the future access probabilities of each replaceable process for the corresponding memory areas in the system memory;

[0049] Among them, the default priority coefficient of the replaceable process is positively correlated with the target instruction running time; the future access probability of the replaceable process for the corresponding memory area in the system memory is negatively correlated with the final instruction running time.

[0050] Optionally, based on the default priority coefficients and target instruction running times of the replaceable processes, determine the final instruction running times of the replaceable processes, including:

[0051] When there is a user-expected process among the replaceable processes, adjust the default priority coefficient of the user-expected process, and determine the final instruction running time of the user-expected process based on the adjusted priority coefficient and the target instruction running time of the user-expected process;

[0052] Based on the default priority coefficients and target instruction running times of the remaining replaceable processes, determine the final instruction running times of the remaining replaceable processes;

[0053] Wherein, the remaining replaceable processes are other processes among the replaceable processes except the user-expected process.

[0054] Optionally, adjusting the default priority coefficient of the user-expected process includes:

[0055] If the user-expected process is a process that does not expect to be determined as a replaceable process, perform a preset reduction operation on the default priority coefficient of the user-expected process;

[0056] If the user-expected process is a process that expects to be determined as a replaceable process, perform a preset increase operation on the default priority coefficient of the user-expected process.

[0057] Optionally, the memory access method of the present invention further includes:

[0058] By determining whether there is a target preset instruction in the instruction cache of the replaced process, to predict whether the memory page table of the replaced process needs to be accessed;

[0059] Wherein, the target preset instruction is the preset instruction with the earliest storage time in the instruction cache of the replaced process, the preset instructions include load instructions and store instructions, and the total number of non-preset instructions in the instruction cache of the replaced process whose storage time is earlier than the target preset instruction is the preset total number.

[0060] Optionally, the memory access method of the present invention further includes:

[0061] If it is predicted that the memory page tables of multiple replaced processes all need to be accessed, then in the order from high to low of the process priorities, successively replace the memory page tables of the replaced processes from the hard disk to the system memory.

[0062] Optionally, the memory access method of the present invention further includes:

[0063] When any replaceable process executes a to-be-executed preset instruction, obtain the target virtual address required by any replaceable process from the to-be-executed preset instruction; wherein, the preset instructions include load instructions and store instructions;

[0064] Obtain a target cache directory corresponding to a target virtual address from a target mapping table corresponding to any replaceable process; the target mapping table stores the mapping relationship between each virtual address required by any replaceable process and each cache directory of any replaceable process;

[0065] Obtain a target physical address corresponding to the target virtual address from the target cache directory;

[0066] Access the target cache corresponding to the target cache directory based on the target physical address;

[0067] Wherein, the target cache is any one of the system memory or the multi-level caches of any replaceable process; each cache directory of any replaceable process includes a memory page table of the system memory, memory page tables of each level of cache, and an address translation buffer of each level of cache.

[0068] Optionally, the update process of the target mapping table includes:

[0069] When updating data from the first cache to the second cache, determine the virtual address of the first updated data, and update the cache directory corresponding to the virtual address of the first updated data in the target mapping table from the memory page table of the first cache to the memory page table of the second cache;

[0070] When data is updated in the address translation buffer of any level of cache in the multi-level cache, obtain the virtual address of the second updated data, and update the cache directory corresponding to the virtual address of the second updated data in the target mapping table to the address translation buffer of any level of cache;

[0071] Wherein, if the first cache is the system memory, the second cache is the last-level cache of any replaceable process; if the first cache is the next-level cache of any replaceable process, the second cache is the upper-level cache of any replaceable process.

[0072] In a second aspect, the present invention provides an electronic device, including:

[0073] A memory for storing a computer program;

[0074] A processor for executing the computer program to implement the steps of the foregoing memory access method.

[0075] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the foregoing memory access method are implemented.

[0076] Fourthly, the present invention provides a computer program product, including computer programs / instructions, which, when executed by a processor, implement the steps of the aforementioned memory access method.

[0077] In the present invention, when a multi-core RISC-V processor allocates a memory area for a target process from the system memory, if the system memory is insufficient, it predicts the future access probabilities of each replaceable process to the corresponding memory areas in the system memory; a replaceable process is a process that has a corresponding memory area in the system memory; based on the process with the smallest future access probability among the replaceable processes, it determines the process to be replaced; it replaces the memory page table of the process to be replaced from the system memory to the hard disk to release the corresponding memory area of the process to be replaced in the system memory, and based on the released system memory, it completes the allocation of the memory area for the target process; if it is predicted that the memory page table of the replaced process needs to be accessed, it re-allocates a memory area for the replaced process and replaces the memory page table of the replaced process from the hard disk to the currently re-allocated memory area to complete the access to the memory page table of the replaced process.

[0078] Beneficial effects: When the present invention allocates a memory area for a target process from the system memory, if the system memory is insufficient, it predicts the future access probabilities of each subsequent replaceable process to the corresponding memory areas in the system memory, and replaces the memory page table of the process to be replaced with the smallest future access probability from the system memory to the hard disk, thereby releasing a part of the memory to allocate a memory area for the target process. Compared with statistically analyzing the active behaviors of each replaceable process in the past period of time, the present invention focuses on predicting the future access probabilities of each subsequent replaceable process and selects the process least likely to be accessed as the process to be replaced, thereby avoiding the situation that the memory page table of the process to be replaced needs to be accessed immediately after being replaced to the hard disk, improving the memory utilization rate of the multi-core RISC-V processor, and further enhancing the overall performance of the multi-core RISC-V processor. Further, the present invention determines whether to re-allocate a memory area for the replaced process from the system memory in advance by predicting whether the memory page table of the replaced process needs to be accessed, and replaces the memory page table of the replaced process from the hard disk back to the system memory. Compared with performing the above operations when the memory page table of the replaced process already needs to be accessed, it can significantly improve the access efficiency of the process's memory page table, that is, improve the memory access efficiency of the multi-core RISC-V processor, and further enhance the overall performance of the multi-core RISC-V processor. Description of the Drawings

[0079] In order to more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0080] Figure 1 It is an architecture diagram of a 64-core RISC-V processor provided by an embodiment of the present invention;

[0081] Figure 2 It is a schematic diagram of memory release provided by an embodiment of the present invention;

[0082] Figure 3 It is a schematic diagram of the memory access structure of a multi-core RISC-V processor provided by an embodiment of the present invention;

[0083] Figure 4 It is a flow chart of virtual address query provided by an embodiment of the present invention;

[0084] Figure 5 It is a flow chart of a memory access method provided by an embodiment of the present invention;

[0085] Figure 6 It is another architecture diagram of a multi-core RISC-V processor provided by an embodiment of the present invention;

[0086] Figure 7 It is an internal structure diagram of an instruction information analysis module provided by an embodiment of the present invention;

[0087] Figure 8 It is a flow chart of instruction execution provided by an embodiment of the present invention;

[0088] Figure 9 It is a hardware structure diagram involved in mapping table update provided by an embodiment of the present invention;

[0089] Figure 10 It is a structure diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0090] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0091] The terms "including" and "having" in the specification of the present invention and any deformations related to "including" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may include steps or units that are not listed.

[0092] To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0093] Compared with the processors of the X86 architecture and the ARM architecture, the single-core performance of the RISC-V processor has a large gap. Therefore, when the RISC-V is used as a server CPU, a multi-core RISC-V interconnection solution is required. However, the multi-core RISC-V processor has the disadvantages of low memory utilization and low memory access efficiency in applications, resulting in the overall low performance of the multi-core RISC-V processor and being unable to exert the best performance of the multi-core processor. For this reason, the present invention provides a memory access method to improve the memory utilization and memory access efficiency of the multi-core RISC-V processor, thereby enhancing the overall performance of the multi-core RISC-V processor.

[0094] Take Figure 1 the 64-core RISC-V processor shown as an example. Each RISC-V Core (kernel) corresponds to an L1Cache (Level 1 cache) and an L2 Cache (Level 2 cache). The 64 RISC-V Cores share an L3 Cache (Level 3 cache). Then, the entire 64-core RISC-V CPU is externally connected to the system memory. Among them, the system memory can be selected as DDR (Double Data Rate Synchronous Dynamic Random Access Memory), and the system memory is externally connected to the hard disk, and the hard disk contains a SWAP partition (swap partition), which is also a typical three-layer storage structure, that is, Cache -> system memory (DDR) -> hard disk.

[0095] Since the multi-core RISC-V processor is to be used in the server field and there are a very large number of software processes to be run, and each software process requires a corresponding system memory area to run, there is often a situation where the system memory is insufficient when a new system memory area needs to be applied for a new software process. The traditional solution is to adopt virtual memory technology, that is, to divide a part of the area in the SWAP partition of the hard disk as temporary memory, and judge the software process to replace the data (memory page table) of the least active software process in the past period of time from the system memory to the temporary memory in the SWAP partition, so as to save a part of the memory area in the system memory to allocate system memory for the new software process.

[0096] However, traditional solutions have significant drawbacks. When the system memory is insufficient, traditional solutions are based on statistical analysis of the active behavior of each software process over a period of time in the past, rather than predicting the active behavior of each software process at future moments. That is to say, traditional solutions cannot solve the problem that the data of the swapped-out process needs to be accessed immediately after being swapped to the hard disk (to access the memory page table of the swapped-out process, the memory page table of the swapped-out process needs to be swapped back from the hard disk to the system memory before it can be accessed normally, that is, the swapped-out process can be executed normally), thus affecting the execution efficiency of the swapped-out process and the memory access efficiency of the swapped-out process, and affecting the memory utilization rate of the multi-core RISC-V processor, resulting in a decline in the overall performance of the multi-core RISC-V processor.

[0097] As Figure 2 shown, one of the most classic ways in traditional solutions is based on the LRU (Least Recently Used) memory eviction algorithm. Its principle is: when the system memory is insufficient, the software process that has used the memory least recently is evicted, that is, the memory page table of the software process is swapped from the system memory to the hard disk, which includes writing the memory page table of the software process to the hard disk and releasing the corresponding memory area of the software process in the system memory.

[0098] Moreover, in the memory access structure of the multi-core RISC-V processor as Figure 3 shown, the MMU (Memory Management Unit) is responsible for completing the conversion between the virtual address and the physical address of the process. The memory management unit contains TLBs (Translation Lookaside Buffers) of each level of cache, and the TLBs store the corresponding relationships between partial virtual addresses and physical addresses of each level of cache. When a process needs to obtain memory data from each level of cache, that is, when the process executes a Load instruction (load instruction) / Store instruction (store instruction), it will first query in the TLB of the MMU. If not found, it will then query in the memory page table.

[0099] Specifically, as Figure 4 shown, after parsing the virtual address from the instruction to be executed by the process, it first queries in the TLB of the first-level cache. If the physical address corresponding to the virtual address can be found, the corresponding memory data is directly obtained from the first-level cache according to the physical address.

[0100] If not found in the TLB of the first-level cache, it will query in the memory page table of the first-level cache. If the physical address corresponding to the virtual address can be found from the memory page table of the first-level cache, the corresponding memory data is obtained from the first-level cache according to the physical address.

[0101] If the data to be accessed cannot be found in the memory page table of the first-level cache, it means that the data is not in the first-level cache. It is necessary to query the TLB of the second-level cache. If the physical address corresponding to the virtual address can be found in the TLB of the second-level cache, the corresponding memory data in the second-level cache is updated to the first-level cache according to the physical address, so as to obtain the memory data.

[0102] If the physical address cannot be found in the TLB of the second-level cache, query the memory page table of the second-level cache. If the physical address corresponding to the virtual address can be found in the memory page table of the second-level cache, the corresponding memory data in the second-level cache is updated to the first-level cache according to the physical address, so as to obtain the memory data.

[0103] If the physical address cannot be found in the memory page table of the second-level cache, it means that the data to be accessed is not in the second-level cache. It is necessary to query the TLB of the third-level cache. If the physical address corresponding to the virtual address can be found in the TLB of the third-level cache, the corresponding memory data in the third-level cache is first updated to the second-level cache and then from the second-level cache to the first-level cache according to the physical address, so as to obtain the memory data.

[0104] If the physical address cannot be found in the TLB of the third-level cache, query the memory page table of the third-level cache. If the physical address corresponding to the virtual address can be found in the memory page table of the third-level cache, the corresponding memory data in the third-level cache is first updated to the second-level cache and then from the second-level cache to the first-level cache according to the physical address, so as to obtain the memory data.

[0105] If the memory page table of the third-level cache cannot find it, it means that the data to be accessed is not in the third-level cache, and query the physical address corresponding to the virtual address in the memory page table of the system memory (it must be possible to find it in the memory page table of the system memory, because when the relevant process applies for system memory, there must be corresponding page table records in the system memory), and then update the corresponding memory data in the system memory to the third-level cache first, then from the third-level cache to the second-level cache, and finally from the second-level cache to the first-level cache according to the physical address, so as to obtain the memory data.

[0106] However Figure 4 This memory access method as shown has a major drawback, that is, when the virtual address of the process cannot be hit in the first-level cache, a very complex access process and multiple interactions are required to correctly obtain the memory data, which is likely to cause problems such as large memory access latency and low access efficiency, and then cause the execution speed of the process to decrease, greatly reducing the overall performance of the multi-core RISC-V processor. For example,

[0107] If the memory data to be retrieved by a process is in the secondary cache, the multi-core RISC-V processor needs to query the TLB of the first-level cache once, the memory page table of the first-level cache once, the TLB of the secondary cache once, and the memory page table of the secondary cache once.

[0108] If the memory data to be retrieved by a process is in the tertiary cache, the multi-core RISC-V processor needs to query the TLB of the first-level cache once, the memory page table of the first-level cache once, the TLB of the secondary cache once, the memory page table of the secondary cache once, the TLB of the tertiary cache once, and the memory page table of the tertiary cache once.

[0109] If the memory data to be retrieved by a process is in the system memory, the multi-core RISC-V processor needs to query the TLB of the first-level cache once, the memory page table of the first-level cache once, the TLB of the secondary cache once, the memory page table of the secondary cache once, the TLB of the tertiary cache once, the memory page table of the tertiary cache once, and the memory page table of the system memory once.

[0110] Therefore, how to improve the memory utilization rate and memory access efficiency of the multi-core RISC-V processor, so as to enhance the overall performance of the multi-core RISC-V processor, is the problem to be achieved and solved by the present invention.

[0111] See Figure 5 As shown, an embodiment of the present invention provides a memory access method, which is applied to a multi-core RISC-V processor and includes:

[0112] Step S11, when allocating a memory area for a target process from the system memory, if the system memory is insufficient, predict the future access probability of each replaceable process to the corresponding memory area in the system memory; a replaceable process is a process that has a corresponding memory area in the system memory.

[0113] The multi-core RISC-V processor proposed in the embodiment of the present invention adopts a single-core single-process, that is, each core in the multi-core RISC-V processor can only run one process at the same time. Correspondingly, the multi-core RISC-V processor of the embodiment of the present invention can also be extended to execute multiple processes on a single core by rapid switching, so as to achieve single-core multi-process (at this time, it makes the user feel the illusion that multiple processes are executing in parallel, but in essence, only one process is running at the same time).

[0114] When the kernel of a multi-core RISC-V processor needs to allocate a memory area for a target process from the system memory during the operation of the target process, it first determines whether the system memory is sufficient. If the system memory is sufficient, it can directly allocate a memory area for the target process from the system memory. If the system memory is insufficient, it is necessary to first predict the future access probabilities of each replaceable process to the corresponding memory areas in the system memory, that is, to predict the access probabilities of each replaceable process to the corresponding memory areas in the system memory at future times.

[0115] Among them, a replaceable process refers to a process that has a corresponding memory area in the system memory; the future access probability of a replaceable process to the corresponding memory area in the system memory refers to the probability of accessing the corresponding memory area of the replaceable process in the system memory at a future time, or the possibility of accessing the corresponding memory area of the replaceable process in the system memory at a future time.

[0116] Moreover, the system memory is the memory shared by each core in the multi-core RISC-V processor and each process running on the core; and the target process is a replaced process or a newly obtained process. It should be noted that when the target process is a newly obtained process by the multi-core RISC-V processor, a free core is first allocated to the new process, and then the free core allocates a memory area for the new process from the system memory. A replaced process refers to a process that has replaced its own memory page table from the system memory to the hard disk.

[0117] Regarding the determination of whether the system memory is sufficient, in a specific implementation, it can be determined whether the system memory is sufficient based on the size relationship between the current remaining memory in the system memory and a preset memory threshold; that is, if the current remaining memory in the system memory is greater than the preset memory threshold, it is determined that the system memory is sufficient, and if the current remaining memory in the system memory is less than or equal to the preset memory threshold, it is determined that the system memory is insufficient.

[0118] In another specific implementation, it can be determined whether the system memory is sufficient based on the size relationship between the current remaining memory in the system memory and the memory required by the target process; that is, if the current remaining memory in the system memory is greater than the memory required by the target process, it is determined that the system memory is sufficient, and if the current remaining memory in the system memory is less than or equal to the memory required by the target process, it is determined that the system memory is insufficient. Of course, there can be other ways to determine whether the system memory is sufficient, which will not be elaborated here.

[0119] For the prediction of the future access probability of each replaceable process to the corresponding memory area in the system memory, it may specifically include: predicting the running time of the target instruction of each replaceable process, where the running time of the target instruction is the running time between two adjacent preset instructions; and determining the future access probability of each replaceable process to the corresponding memory area in the system memory based on the running time of the target instruction of each replaceable process.

[0120] It should be noted that the preset instructions include load instructions and store instructions; both load instructions and store instructions are memory access instructions used to complete corresponding operations in the memory access stage. A load instruction is used to read data from the system memory / each level of cache, and a store instruction is used to write data to the system memory / each level of cache.

[0121] It should also be noted that the future access probability of a replaceable process to the corresponding memory area in the system memory is negatively correlated with the running time of the target instruction of the replaceable process. That is, the shorter the running time of the target instruction of a replaceable process, the more active the replaceable process is. Correspondingly, the future access probability of the replaceable process to the corresponding memory area in the system memory is greater; on the contrary, the longer the running time of the target instruction of the replaceable process, the less active the replaceable process is. Correspondingly, the future access probability of the replaceable process to the corresponding memory area in the system memory is smaller.

[0122] In the process of predicting the running time of the target instruction of each replaceable process, for each replaceable process among the replaceable processes, first, unexecuted instructions are sequentially read from the instruction cache of each replaceable process until a preset number of preset instructions are read; each pair of adjacent preset instructions and the non-preset instructions between each pair of adjacent preset instructions are grouped to obtain several groups of instructions; where the number of groups of the several groups of instructions is one less than the preset number; then, the number of non-preset instructions in each group of instructions is counted, and the time when the non-preset instructions in each group of instructions are executed is obtained to determine the time interval between each pair of adjacent non-preset instructions in each group of instructions; and the running time of the target instruction of each replaceable process is predicted based on the number of non-preset instructions in each group of instructions and the time interval between each pair of adjacent non-preset instructions in each group of instructions.

[0123] It should be noted that the instruction cache of each replaceable process is located in the first-level cache of each replaceable process, and the first-level cache of each replaceable process is divided into two parts. One part is the instruction cache (I-Cache, Instruction Cache) for storing unexecuted instructions, and the other part is the data cache (D-Cache, Data Cache) for storing recently used data. Generally, it is not necessary to distinguish between the instruction cache and the data cache for the second-level cache and the third-level cache of each replaceable process.

[0124] Among them, the unexecuted instructions include preset instructions and non-preset instructions. The preset instructions include storage instructions and load instructions. The non-preset instructions include, but are not limited to, arithmetic operation instructions such as addition instructions, subtraction instructions, multiplication instructions, etc., logical operation instructions such as logical AND instructions, logical OR instructions, etc., and control flow instructions such as branch instructions, etc.

[0125] Taking the preset quantity as Q + 1 (Q is a positive integer) as an example, unexecuted instructions are sequentially read from the instruction cache of each replaceable process until Q + 1 preset instructions are read; the first preset instruction, the second preset instruction, and the non-preset instructions between the first preset instruction and the second preset instruction are divided into a first group; the second preset instruction, the third preset instruction, and the non-preset instructions between the second preset instruction and the third preset instruction are divided into a second group; and so on until the Qth preset instruction, the (Q + 1)th preset instruction, and the non-preset instructions between the Qth preset instruction and the (Q + 1)th preset instruction are divided into the Qth group, thus obtaining Q groups of instructions.

[0126] Assuming that the number of non-preset instructions in a certain group of instructions is N + 1 (N is a positive integer), then the time when each non-preset instruction in this group of instructions is executed is obtained to determine the time interval T1 between the first non-preset instruction and the second non-preset instruction in this group of instructions, the time interval T2 between the second non-preset instruction and the third non-preset instruction, and so on until the time interval TN between the Nth non-preset instruction and the (N + 1)th non-preset instruction, thus obtaining N time intervals.

[0127] After obtaining the number of non-preset instructions in each group of instructions and the time interval between every two adjacent non-preset instructions in each group of instructions, an average value is calculated based on the number of non-preset instructions in each group of instructions to obtain the average number of instructions; based on the time interval between every two adjacent non-preset instructions in each group of instructions, the first average time interval between every two adjacent non-preset instructions in each group of instructions is determined; the average value of the first average time interval between every two adjacent non-preset instructions in each group of instructions is calculated to obtain the second average time interval; according to the average number of instructions and the second average time interval, the target instruction running time of each replaceable process is predicted.

[0128] Suppose there are Q groups of instructions, and the number of non-preset instructions in each group of instructions is denoted as NUM1, NUM2, …, NUMQ respectively. At this time, the average number of instructions = (NUM1 + NUM2 + … + NUMQ) / Q. Suppose the number of non-preset instructions in a certain group of instructions is N + 1 (N is a positive integer), then the first average time interval between two adjacent non-preset instructions in this group of instructions = (T1 + T2 + … + TN) / N, and the second average time interval = (the first average time interval of the first group of instructions + the first average time interval of the second group of instructions + … + the first average time interval of the Qth group of instructions) / Q.

[0129] After obtaining the average number of instructions and the second average time interval of each replaceable process, obtain the current target instruction parameter and the current target time parameter of each replaceable process, and predict the target instruction running time of each replaceable process according to the average number of instructions and the current target instruction parameter of each replaceable process, the second average time interval and the current target time parameter.

[0130] In a specific embodiment, determine the first product result between the average number of instructions and the current target instruction parameter of each replaceable process, and determine the second product result between the second average time interval and the current target time parameter of each replaceable process, and then determine the product result between the first product result and the second product result as the target instruction running time of each replaceable process. That is, the target instruction running time of each replaceable process = average number of instructions × current target instruction parameter × second average time interval × current target time parameter.

[0131] It should be noted that the current target instruction parameter and the current target time parameter of each replaceable process are adaptive adjustment parameters, that is, the current target instruction parameter and the current target time parameter of each replaceable process support adaptive feedback adjustment. Specifically, the multi-core RISC-V processor adaptively adjusts the current target instruction parameter and the current target time parameter of each replaceable process when the system memory is sufficient, so as to predict the target instruction running time of each replaceable process when the system memory is insufficient.

[0132] Specifically, the determination process of the current target instruction parameter and the current target time parameter for each replaceable process includes: when the system memory is sufficient, determining the current instruction parameter to be adjusted and the current time parameter to be adjusted for each replaceable process, and obtaining the current average number of instructions and the current second average time interval for each replaceable process; based on the current average number of instructions, the current instruction parameter to be adjusted, the current second average time interval, and the current time parameter to be adjusted for each replaceable process, determining the predicted instruction running time for each replaceable process; according to the predicted instruction running time and the actual instruction running time of each replaceable process, determining whether to adjust the current instruction parameter to be adjusted and the current time parameter to be adjusted. If an adjustment is needed, adjusting the current instruction parameter to be adjusted and the current time parameter to be adjusted, and based on the current adjusted instruction parameter and the current adjusted time parameter, determining the new current instruction parameter to be adjusted and the new current time parameter to be adjusted, and then re-jumping to the step of obtaining the current average number of instructions and the current second average time interval for each replaceable process; if no adjustment is needed, it means that the current instruction parameter to be adjusted and the current time parameter to be adjusted are already relatively accurate, and the current target instruction parameter and the current target time parameter for each replaceable process can be determined based on the current instruction parameter to be adjusted and the current time parameter to be adjusted.

[0133] When the system memory is sufficient, determining the current instruction parameter to be adjusted and the current time parameter to be adjusted for each replaceable process includes the following two cases: the first case is when the system memory changes from insufficient to sufficient, determining the current target instruction parameter and the current target time parameter of each replaceable process as the current instruction parameter to be adjusted and the current time parameter to be adjusted for each replaceable process; the second case is when the system memory is sufficient and the current moment is the initial moment, determining the current instruction parameter to be adjusted and the current time parameter to be adjusted for each replaceable process based on the preset initial instruction value and the preset initial time value; among them, the preset initial instruction value and the preset initial time value are generally set to one.

[0134] And according to the predicted instruction running time and the actual instruction running time of each replaceable process, determining whether to adjust the current instruction parameter to be adjusted and the current time parameter to be adjusted can specifically include: first determining the time difference between the predicted instruction running time and the actual instruction running time of each replaceable process; then by judging whether the time difference meets the preset difference condition to determine whether to adjust the current instruction parameter to be adjusted and the current time parameter to be adjusted.

[0135] Specifically, if the time difference meets the preset difference condition, the current instruction parameter to be adjusted and the current time parameter to be adjusted are not adjusted; if the time difference does not meet the preset difference condition, the current instruction parameter to be adjusted and the current time parameter to be adjusted are adjusted. Among them, the preset difference condition includes but is not limited to that the time difference is less than the preset difference.

[0136] According to one specific example, if it is determined, based on the predicted instruction running time and the actual instruction running time of each replaceable process, not to adjust the current instruction parameter to be adjusted and the current time parameter to be adjusted for each replaceable process, then directly determine the current instruction parameter to be adjusted and the current time parameter to be adjusted for each replaceable process as the current target instruction parameter and the current target time parameter for each replaceable process.

[0137] Further, after predicting the target instruction running time of each replaceable process, based on the default priority coefficient and the target instruction running time of each replaceable process, determine the final instruction running time of each replaceable process; then according to the final instruction running time of each replaceable process, determine the future access probability of each replaceable process to the corresponding memory area in the system memory.

[0138] Among them, the default priority coefficient of the replaceable process is positively correlated with the target instruction running time; the future access probability of the replaceable process to the corresponding memory area in the system memory is negatively correlated with the final instruction running time.

[0139] That is to say, the shorter the target instruction running time of the replaceable process, the more active the replaceable process is, the higher the priority of the replaceable process, and at this time, the smaller the default priority coefficient of the replaceable process. In this case, the final instruction running time determined based on the default priority coefficient and the target instruction running time of the replaceable process is correspondingly smaller, and the future access probability of the replaceable process to the corresponding memory area in the system memory will be greater. On the contrary, the longer the target instruction running time of the replaceable process, the less active the replaceable process is, the lower the priority of the replaceable process, and at this time, the larger the default priority coefficient of the replaceable process. In this case, the final instruction running time determined based on the default priority coefficient and the target instruction running time of the replaceable process is correspondingly larger, and the future access probability of the replaceable process to the corresponding memory area in the system memory will be smaller.

[0140] According to one specific example, the final instruction running time of the replaceable process = the target instruction running time of the replaceable process × the default priority coefficient of the replaceable process; among them, the default priority coefficient can be set by default or configured by the user independently.

[0141] However, in the embodiments of the present invention, it is considered that the user can also set an expected process, abbreviated as the user-expected process. The user-expected process can be a process that is not expected to be determined as a process to be replaced, or a process that is expected to be determined as a process to be replaced. When there is a user-expected process, the default priority coefficient of the user-expected process can be adjusted so that the user-expected process finally meets the user's expectations.

[0142] Specifically, when there is a user-expected process among the replaceable processes, the default priority coefficient of the user-expected process is adjusted, and based on the adjusted priority coefficient of the user-expected process and the target instruction running time, the final instruction running time of the user-expected process is determined; for the other processes among the replaceable processes except the user-expected process, abbreviated as the remaining replaceable processes, the final instruction running time of the remaining replaceable processes is determined based on the default priority coefficient and the target instruction running time of the remaining replaceable processes.

[0143] In the process of adjusting the default priority coefficient of the user-expected process, it can specifically include: if the user-expected process is a process that is not expected to be determined as a process to be replaced, then a preset reduction operation is performed on the default priority coefficient of the user-expected process. Correspondingly, the final instruction running time of the user-expected process will become smaller, so that the future access probability of the user-expected process to the corresponding memory area in the system memory becomes larger. If the user-expected process is a process that is expected to be determined as a process to be replaced, then a preset increase operation is performed on the default priority coefficient of the user-expected process. Correspondingly, the final instruction running time of the user-expected process will become larger, so that the future access probability of the user-expected process to the corresponding memory area in the system memory becomes smaller.

[0144] Among them, the preset reduction operation includes, but is not limited to, adjusting the default priority coefficient of the user-expected process to a preset minimum coefficient that tends to be infinitely small; the preset increase operation includes, but is not limited to, adjusting the default priority coefficient of the user-expected process to a preset maximum coefficient that tends to be infinitely large.

[0145] For determining the future access probability of each replaceable process to the corresponding memory area in the system memory according to the final instruction running time of each replaceable process, there are the following two determination methods: The first determination method is to determine the total sum of the final instruction running times of each replaceable process, and determine the ratio of the final instruction running time of the replaceable process to the total sum of time. Then, the difference between 1 and the ratio is determined as the future access probability of the replaceable process to the corresponding memory area in the system memory; that is, the future access probability of the replaceable process to the corresponding memory area in the system memory = 1 - the final instruction running time of the replaceable process / the total sum of time. The second determination method is to construct a monotonically decreasing function based on the negative correlation between the future access probability of the replaceable process to the corresponding memory area in the system memory and the final instruction running time, and substitute the final instruction running time of the replaceable process into the monotonically decreasing function to determine the future access probability of the replaceable process to the corresponding memory area in the system memory.

[0146] Step S12: Determine the process to be replaced based on the process with the minimum future access probability among each replaceable process.

[0147] In the embodiment of the present invention, after determining the future access probability of each replaceable process to the corresponding memory area in the system memory, query the process with the minimum future access probability from each replaceable process, and use the process with the minimum future access probability as the process to be replaced.

[0148] It should be noted that when there are multiple processes with the minimum future access probability among each replaceable process, any one of the processes with the minimum future access probability can be used as the process to be replaced, or candidate processes can be first determined from the processes with the minimum future access probability. The size of the memory area corresponding to the candidate process in the system memory plus the current remaining memory of the system memory should be greater than the memory required by the target process, and then any one of the candidate processes can be used as the process to be replaced.

[0149] Step S13: Replace the memory page table of the process to be replaced from the system memory to the hard disk to release the memory area corresponding to the process to be replaced in the system memory, and complete the memory area allocation for the target process based on the released system memory.

[0150] In the embodiment of the present invention, after determining the process to be replaced from each replaceable process, replace the memory page table of the process to be replaced from the system memory to a specific area of the hard disk to release the memory area corresponding to the process to be replaced in the system memory, and then the memory area allocation for the target process can be completed based on the released system memory, that is, allocate a memory area for the target process from the released system memory.

[0151] Step S14: If it is predicted that the memory page table of the swapped-out process needs to be accessed, re-allocate a memory area for the swapped-out process, and swap the memory page table of the swapped-out process from the hard disk to the currently re-allocated memory area, so as to complete the access to the memory page table of the swapped-out process.

[0152] In the embodiment of the present invention, the swapped-out process is a process that has swapped its own memory page table from the system memory to the hard disk. If it is predicted that the memory page table of the swapped-out process is about to be accessed, re-allocate a memory area for the swapped-out process from the system memory, and swap the memory page table of the swapped-out process from the hard disk to the currently re-allocated memory area, so as to complete the access to the memory page table of the swapped-out process and realize the normal execution of the swapped-out process.

[0153] Specifically, determine whether there is a target preset instruction in the instruction cache of the swapped-out process to predict whether the memory page table of the swapped-out process needs to be accessed; wherein, the target preset instruction is the preset instruction with the earliest storage time in the instruction cache of the swapped-out process, the preset instructions include load instructions and store instructions, and the total number of non-preset instructions in the instruction cache of the swapped-out process with a storage time earlier than the target preset instruction is the preset total number.

[0154] That is, if there is a target preset instruction in the instruction cache of the swapped-out process, it is predicted that the memory page table of the swapped-out process needs to be accessed; if there is no target preset instruction in the instruction cache of the swapped-out process, it is predicted that the memory page table of the swapped-out process does not need to be accessed.

[0155] In other words, assuming that the preset total number is P, analyze the unexecuted instructions in the instruction cache of the swapped-out process, and when it is found that there are still P non-preset instructions to execute the preset instruction, predict that the memory page table of the swapped-out process needs to be accessed.

[0156] It should be noted that if it is predicted that the memory page tables of multiple swapped-out processes all need to be accessed, swap the memory page tables of the swapped-out processes from the hard disk to the system memory in order from the highest to the lowest process priority. Among them, the higher the process priority, the smaller the default priority coefficient of the process.

[0157] Beneficial effects: When the present invention allocates a memory area for a target process from the system memory, if the system memory is insufficient, it predicts the future access probability of each replaceable process to the corresponding memory area in the system memory, and replaces the memory page table of the replaceable process with the minimum future access probability from the system memory to the hard disk, thereby releasing a part of the memory to allocate a memory area for the target process. Compared with counting the active behaviors of each replaceable process in the past period of time, the present invention focuses on predicting the future access probability of each replaceable process, and selects the process least likely to be accessed as the replaceable process, thereby avoiding the situation that the memory page table of the replaceable process needs to be accessed immediately after being replaced to the hard disk, improving the memory utilization rate of the multi-core RISC-V processor, and further improving the overall performance of the multi-core RISC-V processor. Further, the present invention determines whether to re-allocate a memory area for the replaced process from the system memory by predicting in advance whether the memory page table of the replaced process needs to be accessed, and replaces the memory page table of the replaced process from the hard disk back to the system memory. Compared with performing the above operations when the memory page table of the replaced process already needs to be accessed, it can significantly improve the access efficiency of the memory page table of the process, that is, improve the memory access efficiency of the multi-core RISC-V processor, and further improve the overall performance of the multi-core RISC-V processor.

[0158] See Figure 6 As shown in the figure, taking a multi-core RISC-V processor as a 64-core RISC-V processor and adopting a three-layer storage structure of three-level cache, system memory and hard disk as an example, a memory access method proposed in an embodiment of the present invention will be described in detail.

[0159] First, in the hardware structure of the multi-core RISC-V processor, it is necessary to add a new instruction information analysis module for each core, add a replacement scheduling analysis module as a whole, and connect the instruction information analysis module of each core to the replacement scheduling analysis module, and the instruction information analysis module of each core is connected to each core and the first-level cache of each core.

[0160] Secondly, in the software logic of the multi-core RISC-V processor, when the core running the target process in the multi-core RISC-V processor needs to allocate a memory area for the target process from the system memory, if the system memory is insufficient, it is necessary to predict the future access probability of the replaceable process to the corresponding memory area in the system memory through the instruction information analysis module of each target core; where the target core is the core running the replaceable process, and the multi-core RISC-V processor adopts single-core single-process.

[0161] Taking the instruction information analysis module of any target core as an example for description, as Figure 7As shown, the interface module in the instruction information analysis module sequentially reads unexecuted instructions from the instruction cache of the first-level cache, and analyzes the instruction types of the read unexecuted instructions through the instruction parsing module in the instruction information analysis module, and the instruction types include preset instructions and non-preset instructions, until a preset number of preset instructions are read from the instruction cache of the first-level cache. The analysis and statistics module in the instruction information analysis module groups each adjacent two preset instructions and the non-preset instructions between each adjacent two preset instructions into a group to obtain a number of groups of instructions, and then counts the number of non-preset instructions in each group of instructions, and obtains the time when the non-preset instructions in each group of instructions are executed (the interface module determines the time when the non-preset instructions are executed by monitoring the value operation of the instruction fetch module in the kernel on the instruction cache), and determines the time interval between each adjacent two non-preset instructions in each group of instructions, and finally predicts the target instruction running time of the replaceable process based on the number of non-preset instructions in each group of instructions and the time interval between each adjacent two non-preset instructions in each group of instructions, and determines the future access probability of the replaceable process to the corresponding memory area in the system memory based on the target instruction running time of the replaceable process.

[0162] Furthermore, the instruction information analysis module of each target core transmits the predicted future access probability to the replacement scheduling analysis module. Accordingly, the replacement scheduling analysis module determines the minimum future access probability from each future access probability, and determines the process corresponding to the minimum future access probability as the process to be replaced, and then replaces the memory page table of the process to be replaced from the system memory to the replacement partition in the hard disk, thereby releasing the memory area corresponding to the process to be replaced in the system memory, and completes the memory area allocation for the target process based on the released system memory.

[0163] It should be noted that for the kernels running the replaced processes, the instruction information analysis modules of these kernels will still monitor the instruction cache in the first-level cache to analyze the unexecuted instructions in the instruction cache, and when it is found that there are P non-preset instructions that are about to execute the preset instructions, it is predicted that the memory page table of the replaced process will soon need to be accessed, and an indication signal is generated at this time, and the indication signal is transmitted to the replacement scheduling analysis module.

[0164] After receiving the upcoming use indication signal sent by a kernel, the replacement scheduling analysis module reallocates the memory area from the system memory for the replaced process running on the kernel, and replaces the memory page table of the replaced process from the replacement partition of the hard disk back to the system memory, thereby completing the access to the memory page table of the replaced process and realizing the normal execution of the replaced process.

[0165] Beneficial effects: When the present invention allocates a memory area for a target process from the system memory, if the system memory is insufficient, it predicts the future access probability of each replaceable process to the corresponding memory area in the system memory, and replaces the memory page table of the replaceable process with the minimum future access probability from the system memory to the hard disk, thereby releasing a part of the memory to allocate a memory area for the target process. Compared with counting the active behaviors of each replaceable process in the past period of time, the present invention focuses on predicting the future access probability of each replaceable process, and selects the process least likely to be accessed as the replaceable process, thus avoiding the situation that the memory page table of the replaceable process needs to be accessed immediately after being replaced to the hard disk, improving the memory utilization rate of the multi-core RISC-V processor, and further improving the overall performance of the multi-core RISC-V processor. Further, the present invention determines whether to re-allocate a memory area for the replaced process from the system memory in advance by predicting whether the memory page table of the replaced process needs to be accessed, and replaces the memory page table of the replaced process from the hard disk back to the system memory. Compared with performing the above operations when the memory page table of the replaced process already needs to be accessed, it can significantly improve the access efficiency of the memory page table of the process, that is, improve the memory access efficiency of the multi-core RISC-V processor, and further improve the overall performance of the multi-core RISC-V processor.

[0166] See Figure 8 As shown, an embodiment of the present invention provides an execution process of a process for a preset instruction, which is applied to a multi-core RISC-V processor, including:

[0167] Step S21, when any replaceable process executes a preset instruction to be executed, obtain a target virtual address required by any replaceable process from the preset instruction to be executed; wherein, the preset instruction includes a load instruction and a store instruction.

[0168] In an embodiment of the present invention, when a multi-core RISC-V processor executes a preset instruction to be executed by any replaceable process, it parses the preset instruction to be executed to obtain a target virtual address required by any replaceable process.

[0169] Step S22, obtain a target cache directory corresponding to the target virtual address from the target mapping table corresponding to any replaceable process; the target mapping table stores the mapping relationship between each virtual address required by any replaceable process and each cache directory of any replaceable process.

[0170] In an embodiment of the present invention, the memory management unit queries the target mapping table corresponding to any replaceable process to obtain the target cache directory corresponding to the target virtual address. Among them, different replaceable processes correspond to different mapping tables, and the mapping table stores the mapping relationship between each virtual address required by the replaceable process and each cache directory of the replaceable process.

[0171] Specifically, the mapping table may also store the mapping relationship between each virtual address required by the swappable process and the directory identifier of each cache directory of the swappable process, and different cache directories correspond to different directory identifiers.

[0172] Step S23: Obtain the target physical address corresponding to the target virtual address from the target cache directory.

[0173] In the embodiment of the present invention, the corresponding relationship between each virtual address and each physical address is stored in the cache directory of the swappable process. Therefore, by querying the target cache directory, the target physical address corresponding to the target virtual address can be obtained.

[0174] Step S24: Access the target cache corresponding to the target cache directory based on the target physical address; wherein, the target cache is any one of the system memory or the multi-level caches of any swappable process; each cache directory of any swappable process includes the memory page table of the system memory, the memory page tables of each level of cache, and the address translation lookaside buffer of each level of cache.

[0175] In the embodiment of the present invention, after obtaining the target physical address corresponding to the target virtual address, access the target cache corresponding to the target cache directory based on the target physical address, that is, access the memory space corresponding to the target physical address in the target cache, such as data writing or reading.

[0176] It should be noted that the target cache can be either the system memory or any one of the multi-level caches of any swappable process, and the multi-level cache can adopt a three-level cache. Correspondingly, each cache directory of any swappable process includes the memory page table of the system memory, the memory page tables of each level of cache, and the address translation lookaside buffer of each level of cache.

[0177] Further, for the update process of the target mapping table corresponding to any swappable process, it may specifically include: when the first cache updates data to the second cache, determine the virtual address of the first updated data, and update the cache directory corresponding to the virtual address of the first updated data in the target mapping table from the memory page table of the first cache to the memory page table of the second cache; wherein, if the first cache is the system memory, the second cache is the last-level cache of any swappable process; if the first cache is the next-level cache of any swappable process, the second cache is the upper-level cache of any swappable process. When the data in the address translation lookaside buffer of any level of cache in the multi-level cache is updated, obtain the virtual address of the second updated data, and update the cache directory corresponding to the virtual address of the second updated data in the target mapping table to the address translation lookaside buffer of any level of cache.

[0178] It should be noted that when updating data from the first cache to the second cache, it not only includes updating the data corresponding to a certain physical address in the first cache to the second cache, but also includes updating the memory page table of the first cache and the memory page table of the second cache, that is, updating the correspondence between the virtual address and the physical address in the memory page table.

[0179] Taking a three-level cache of any replaceable process as an example, as Figure 9 shown, by adding a first-level update module, a second-level update module, a third-level update module, and a mapping table update module in a multi-core RISC-V processor, the creation and update of the target mapping table corresponding to any replaceable process can be achieved. It should be noted that in the initial stage of the target mapping table, each virtual address required by any replaceable process corresponds to the memory page table of the system memory.

[0180] Among them, the third-level update module is connected to the system memory to determine the virtual address of the updated data when the system memory updates the data to the third-level cache, and update the cache directory corresponding to the virtual address of the updated data in the target mapping table from the memory page table of the system memory to the memory page table of the third-level cache.

[0181] The second-level update module is connected to the third-level cache to determine the virtual address of the updated data when the third-level cache updates the data to the second-level cache, and update the cache directory corresponding to the virtual address of the updated data in the target mapping table from the memory page table of the third-level cache to the memory page table of the second-level cache.

[0182] The first-level update module is connected to the second-level cache to determine the virtual address of the updated data when the second-level cache updates the data to the first-level cache, and update the cache directory corresponding to the virtual address of the updated data in the target mapping table from the memory page table of the second-level cache to the memory page table of the first-level cache.

[0183] The mapping table update module is connected to the address translation lookaside buffers of the first-level cache, the second-level cache, and the third-level cache in the first-level update module, the second-level update module, the third-level update module, and the memory management unit, and is used to receive the request information for updating the target mapping table sent by the first-level update module, the second-level update module, and the third-level update module, and monitor the address translation lookaside buffers of the first-level cache, the second-level cache, and the third-level cache.

[0184] When the mapping table update module detects that the address translation lookaside buffer of any first-level cache has data updated, it obtains the virtual address of the updated data, and updates the cache directory corresponding to the virtual address of the updated data in the target mapping table to the address translation lookaside buffer of any first-level cache. In this way, the mapping table update module can have a target mapping table that is updated in real time.

[0185] Beneficial effects: In the embodiments of the present invention, by querying the target mapping table corresponding to the replaceable process, the target cache directory corresponding to the virtual address required by the replaceable process is obtained, and then the physical address corresponding to the virtual address required by the replaceable process is directly queried in the target cache directory, and finally, based on the queried physical address, the access to the target cache corresponding to the target cache directory is realized. Compared with the traditional scheme that requires a more complex access process and multiple data interactions, the embodiments of the present invention can greatly reduce the query steps for the physical address corresponding to the virtual address, thereby reducing the steps of memory access and accelerating the efficiency of memory access, thus improving the overall performance of the multi-core RISC-V processor.

[0186] Further, the embodiments of the present application also disclose an electronic device. Figure 10 It is a structural diagram of an electronic device shown according to an exemplary embodiment, and the content in the figure cannot be considered as any limitation to the scope of use of the present application. The electronic device may specifically include: at least one processor 11, at least one memory 12, a power supply 13, a communication interface 14, an input / output interface 15, and a communication bus 16. Among them, the memory 12 is used to store computer programs, and the computer programs are loaded and executed by the processor 11 to implement the relevant steps in the memory access method disclosed in any of the foregoing embodiments. In addition, the electronic device in this embodiment may specifically be an electronic computer.

[0187] In this embodiment, the power supply 13 is used to provide operating voltage for each hardware device on the electronic device; the communication interface 14 can create a data transmission channel between the electronic device and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and specific limitations are not imposed here; the input / output interface 15 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and no specific limitations are imposed here.

[0188] In addition, as a carrier for resource storage, the memory 12 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc., and the resources stored thereon may include an operating system 121, a computer program 122, etc., and the storage method may be short-term storage or permanent storage.

[0189] Among them, the operating system 121 is used to manage and control each hardware device and the computer program 122 on the electronic device, and it may be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the memory access method executed by the electronic device disclosed in any of the foregoing embodiments, the computer program 122 may further include computer programs that can be used to complete other specific tasks.

[0190] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the foregoing disclosed memory access method is implemented. For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0191] Furthermore, the present application also discloses a computer program product, including a computer program / instructions; wherein, when the computer program / instructions are executed by a processor, the foregoing disclosed memory access method is implemented. For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0192] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference may be made to each other. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference may be made to the description in the method part for the relevant parts.

[0193] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0194] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0195] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0196] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this text to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A memory access method, characterized in that, Applied to a multi-core RISC-V processor, including: When allocating a memory area for a target process from the system memory, if the system memory is insufficient, predict the future access probability of each replaceable process to the corresponding memory area in the system memory; the replaceable process is a process corresponding to a memory area in the system memory; Based on the process with the minimum future access probability among each of the replaceable processes, determine the process to be replaced; Replace the memory page table of the process to be replaced from the system memory to the hard disk to release the corresponding memory area of the process to be replaced in the system memory, and complete the allocation of the memory area for the target process based on the released system memory; If it is predicted that the memory page table of the replaced process needs to be accessed, re-allocate a memory area for the replaced process, and replace the memory page table of the replaced process from the hard disk to the currently re-allocated memory area to complete the access to the memory page table of the replaced process; Among them, the predicting the future access probability of each replaceable process to the corresponding memory area in the system memory includes: Predict the target instruction running time of each replaceable process; the target instruction running time is the running time between two adjacent preset instructions; among them, the preset instructions include load instructions and store instructions; Based on each of the target instruction running times, determine the future access probability of each replaceable process to the corresponding memory area in the system memory.

2. The memory access method according to claim 1, wherein The target process is the replaced process or a newly obtained process.

3. The memory access method according to claim 1, wherein Also includes: Based on the size relationship between the current remaining memory in the system memory and a preset memory threshold, determine whether the system memory is sufficient; And / or, based on the size relationship between the current remaining memory and the memory required by the target process, determine whether the system memory is sufficient.

4. The memory access method according to claim 1, wherein The predicting the target instruction running time of each replaceable process includes: Sequentially read unexecuted instructions from the instruction cache of each replaceable process until a preset number of preset instructions are read; Divide each adjacent two preset instructions and the non-preset instructions between each adjacent two preset instructions into a group to obtain several groups of instructions; the number of groups of the several groups of instructions is one less than the preset number; Count the number of non-preset instructions in each group of instructions; Obtain the time when the non-preset instructions in each group of instructions are executed to determine the time interval between each adjacent two non-preset instructions in each group of instructions; Based on the number of non-preset instructions in each group of instructions and the time interval between each adjacent two non-preset instructions in each group of instructions, predict the target instruction running time of each replaceable process.

5. The memory access method according to claim 4, wherein The predicting the target instruction running time of each replaceable process based on the number of non-preset instructions in each group of instructions and the time interval between each adjacent two non-preset instructions in each group of instructions includes: Calculate the average value based on the number of non-preset instructions in each group of instructions to obtain the average number of instructions; Determine a first average time interval between two adjacent non-preset instructions in each set of instructions based on the time interval between every two adjacent non-preset instructions in each set of instructions; Calculate the average value of the first average time intervals between two adjacent non-preset instructions in each set of instructions to obtain a second average time interval; Predict the target instruction running time of each replaceable process according to the average number of instructions and the second average time interval.

6. The memory access method according to claim 5, wherein The predicting the target instruction running time of each replaceable process according to the average number of instructions and the second average time interval includes: Obtain the current target instruction parameters and current target time parameters of each replaceable process; Predict the target instruction running time of each replaceable process according to the average number of instructions, the current target instruction parameters, the second average time interval, and the current target time parameters.

7. The memory access method according to claim 6, characterized in that The determining process of the current target instruction parameters and current target time parameters of each replaceable process includes: When the system memory is sufficient, determine the current instruction parameters to be adjusted and current time parameters to be adjusted of each replaceable process; Obtain the current average number of instructions and current second average time interval of each replaceable process; Determine the predicted instruction running time of each replaceable process based on the current average number of instructions, the current instruction parameters to be adjusted, the current second average time interval, and the current time parameters to be adjusted; Determine whether to adjust the current instruction parameters to be adjusted and the current time parameters to be adjusted according to the predicted instruction running time and actual instruction running time of each replaceable process; If so, determine new current instruction parameters to be adjusted and new current time parameters to be adjusted based on the current adjusted instruction parameters and current adjusted time parameters, and re-jump to the step of obtaining the current average number of instructions and current second average time interval of each replaceable process; If not, determine the current target instruction parameters and current target time parameters based on the current instruction parameters to be adjusted and the current time parameters to be adjusted.

8. The memory access method according to claim 7, wherein The determining the current instruction parameters to be adjusted and current time parameters to be adjusted of each replaceable process includes: When the system memory changes from insufficient to sufficient, determine the current target instruction parameters and current target time parameters of each replaceable process as the current instruction parameters to be adjusted and current time parameters to be adjusted of each replaceable process; When the system memory is sufficient and the current time is the initial time, determine the current instruction parameters to be adjusted and current time parameters to be adjusted of each replaceable process based on the preset instruction initial value and preset time initial value.

9. The memory access method according to claim 7, wherein The determining whether to adjust the current instruction parameters to be adjusted and the current time parameters to be adjusted according to the predicted instruction running time and actual instruction running time of each replaceable process includes: Determine the time difference between the predicted instruction running time and actual instruction running time of each replaceable process; Whether to adjust the current instruction parameter to be adjusted and the current time parameter to be adjusted is determined by judging whether the time difference satisfies a preset difference condition.

10. The memory access method according to claim 1, wherein The determining, based on the execution time of each target instruction, the probability of each replaceable process accessing the corresponding memory area in the system memory in the future includes: Determining a final instruction execution time of each of the replaceable processes based on a default priority coefficient of each of the replaceable processes and the target instruction execution time; Determining, according to the final instruction execution time of each replaceable process, a future access probability of each replaceable process to a corresponding memory area in the system memory; Among them, the default priority coefficient of the replaceable process is positively correlated with the target instruction execution time; the future access probability of the replaceable process to the corresponding memory area in the system memory is negatively correlated with the final instruction execution time.

11. The memory access method according to claim 10, wherein The determining the final instruction execution time of each replaceable process based on the default priority coefficient of each replaceable process and the target instruction execution time includes: When there is a user desired process in each of the replaceable processes, adjusting the default priority coefficient of the user desired process, and determining the final instruction execution time of the user desired process based on the adjusted priority coefficient of the user desired process and the target instruction execution time; Determining a final instruction execution time of the remaining replaceable processes based on the default priority coefficients of the remaining replaceable processes and the target instruction execution time; The remaining replaceable processes are other processes among the replaceable processes except the user desired process.

12. The memory access method according to claim 11, wherein The adjusting the default priority coefficient of the process desired by the user includes: If the user desired process is a process that is not desired to be determined as the process to be replaced, a preset lowering operation is performed on the default priority coefficient of the user desired process; If the user-desired process is a process expected to be determined as the process to be replaced, a preset increase operation is performed on the default priority coefficient of the user-desired process.

13. The memory access method according to claim 1, characterized in that, Also includes: By determining whether there is a target preset instruction in the instruction cache of the replaced process, it is pre-determined whether the memory page table of the replaced process needs to be accessed; Among them, the target preset instruction is the preset instruction with the earliest storage time in the instruction cache of the replaced process, the preset instructions include load instructions and storage instructions, and the total number of non-preset instructions in the instruction cache of the replaced process whose storage time is earlier than the target preset instruction is the preset total number.

14. The memory access method according to claim 13, wherein Also includes: If it is predicted that the memory page tables of the multiple replaced processes all need to be accessed, the memory page tables of the replaced processes are replaced from the hard disk to the system memory in order of process priority from high to low.

15. The memory access method according to any one of claims 1 to 14, characterized in that, Also includes: When any replaceable process executes a preset instruction to be executed, a target virtual address required to be used by any replaceable process is obtained from the preset instruction to be executed; wherein the preset instruction includes a load instruction and a store instruction; Obtain the target cache directory corresponding to the target virtual address from the target mapping table corresponding to any of the replaceable processes; the target mapping table stores the mapping relationship between each virtual address required by any of the replaceable processes and each cache directory of any of the replaceable processes; Obtain the target physical address corresponding to the target virtual address from the target cache directory; Access the target cache corresponding to the target cache directory based on the target physical address; Wherein, the target cache is any one of the system memory or the multi-level caches of any of the replaceable processes; each cache directory of any of the replaceable processes includes the memory page table of the system memory, the memory page tables of each level of cache, and the address translation lookaside buffer of each level of cache.

16. The memory access method according to claim 15, wherein The update process of the target mapping table includes: When the first cache updates data to the second cache, determine the virtual address of the first updated data, and update the cache directory corresponding to the virtual address of the first updated data in the target mapping table from the memory page table of the first cache to the memory page table of the second cache; When the address translation lookaside buffer of any level of cache in the multi-level cache has data updated, obtain the virtual address of the second updated data, and update the cache directory corresponding to the virtual address of the second updated data in the target mapping table to the address translation lookaside buffer of any level of cache; Wherein, if the first cache is the system memory, the second cache is the last-level cache of any of the replaceable processes; if the first cache is the next-level cache of any of the replaceable processes, the second cache is the upper-level cache of any of the replaceable processes.

17. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for executing the computer program to implement the steps of the memory access method according to any one of claims 1 to 16.

18. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the steps of the memory access method according to any one of claims 1 to 16 are implemented.

19. A computer program product comprising computer programs / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the memory access method according to any one of claims 1 to 16 are implemented.

Citation Information

Patent Citations

  • Memory management method and electronic equipment

    CN116860429A

  • Dynamic instruction conversion memory access conflict optimization method based on memory partition

    CN119847609A