Method and device for increasing storage information density of translation bypass cache, electronic equipment and computer program product

By merging multiple consecutive address mapping data into a merged entry and introducing address region buffer identifiers, the problem of wasted TLB storage space is solved, and the storage information density of TLB and system processing performance are improved.

CN121029643AActive Publication Date: 2025-11-28GUANGDONG LEAPFIVE TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511575343.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2025-11-28
Estimated Expiration
2045-10-31

AI Technical Summary

Technical Problem

In existing technologies, the TLB uses a single entry to store a single address mapping, which limits the effective capacity improvement under limited chip area, resulting in wasted storage space and affecting system processing performance.

Method used

By merging multiple consecutive address mapping data into a single merged entry, sharing the same high-order portion of storage, and introducing an address region buffer identifier to replace the actual high-order address storage method, storage space is compressed.

Benefits of technology

While maintaining the same physical TLB capacity, the storage density of effective address mapping was significantly increased, subsequent address translation latency was reduced, and system processing performance was improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029643A_ABST
    Figure CN121029643A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for increasing the storage information density of a translation bypass cache, electronic equipment and a computer program product.The method comprises the steps that under the condition that an address translation request of a first virtual address is received, multiple pieces of continuous address mapping data are obtained according to the first virtual address, each piece of address mapping data comprises a mapping relationship between a virtual address and a physical address; under the condition that the multiple pieces of continuous address mapping data meet a preset merging condition, merging the multiple pieces of continuous address mapping data into a merged entry; and storing the merged entry into a translation bypass cache, and returning a target physical address corresponding to the first virtual address to the address request unit according to the merged entry. Therefore, multiple pieces of continuous address mapping data are merged into a single entry and share and store the same high-order part, so that the storage density of effective address mapping is remarkably improved under the condition that the physical capacity of a translation bypass cache is kept unchanged.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of chips, and particularly relates to a method and device for increasing the storage information density of a translation lookaside buffer, an electronic device, and a computer program product. BACKGROUND

[0002] With the rapid development of high-performance computing, big data, and artificial intelligence applications, higher requirements are put forward for the memory access efficiency of a processor. A translation lookaside buffer (TLB) in a memory management unit (MMU) is a key component that determines the address translation efficiency, and the capacity and efficiency of the TLB directly affect the overall performance of the system.

[0003] In related technologies, the TLB adopts a structure of storing a single address mapping in a single entry, and each entry independently stores complete high bits of a virtual address and high bits of a physical address. When there are continuous virtual address mappings in an address space, this structure causes a large amount of same high bit address information to be repeatedly stored in different entries, resulting in significant waste of storage space, limiting the improvement of the effective capacity of the TLB under limited chip area, and further affecting the processing performance of the system. SUMMARY

[0004] The embodiments of the application provide a method and device for increasing the storage information density of a translation lookaside buffer, an electronic device, and a computer program product, which can significantly improve the storage information density of the TLB.

[0005] A first aspect of the embodiments of the application provides a method for increasing the storage information density of a translation lookaside buffer, comprising: in a case where an address translation request of a first virtual address is received, obtaining a plurality of continuous address mapping data according to the first virtual address, wherein the request source corresponding to the first virtual address is an address request unit, and each address mapping data contains a mapping relationship between a virtual address and a physical address; in a case where the plurality of continuous address mapping data satisfy a preset merging condition, merging the plurality of continuous address mapping data into one merged entry, wherein the merged entry stores the same high bit part of the virtual address, the same high bit part of the physical address, and the address low bit part and the valid bit corresponding to each address mapping data in the plurality of continuous address mapping data, respectively; storing the merged entry into the translation lookaside buffer, and returning a target physical address corresponding to the first virtual address to the address request unit according to the merged entry.

[0006] In the technical solution of the application, the plurality of continuous address mapping data are merged into a single entry and the same high bit part is shared, so that the storage density of the effective address mapping is significantly improved while the physical capacity of the TLB remains unchanged.

[0007] Optionally, in a possible implementation manner of the first aspect, before the plurality of continuous address mapping data are merged into one merged entry under the condition that the plurality of continuous address mapping data satisfy the preset merging condition, the method further includes: determining a to-be-checked bit range of a virtual address (VA) and a physical address (PA) in the plurality of continuous address mapping data according to a page size corresponding to the address translation request; determining whether the high-order parts of all virtual addresses in the to-be-checked bit range are same and whether the high-order parts of all physical addresses in the to-be-checked bit range are same based on the plurality of continuous address mapping data; and determining that the plurality of continuous address mapping data satisfy the merging condition in a case that the high-order parts of all virtual addresses in the to-be-checked bit range are same and the high-order parts of all physical addresses in the to-be-checked bit range are same. In this way, by dynamically determining the to-be-checked bit range according to the page size and performing corresponding high-order part consistency verification, it is ensured that the merging operation can adapt to address mapping scenarios of different page specifications, and the generality and reliability of the scheme are improved.

[0008] Optionally, in another possible implementation manner of the first aspect, before the merged entry is stored into the translation bypass cache, the method further includes: sending the same virtual address high-order part and the same physical address high-order part in the merged entry to an address region buffer in the translation bypass cache, where the same virtual address high-order part and the same physical address high-order part in the address region buffer are associated with a set address region buffer identifier (ARB); and replacing the same virtual address high-order part and the same physical address high-order part in the merged entry with the address region buffer identifier. In this way, by introducing the address region buffer and replacing the storage of the actual high-order address value with the identifier, the storage space occupation of the merged entry is further compressed, and a two-level compression effect is achieved.

[0009] Optionally, in still another possible implementation manner of the first aspect, the method further includes: in a case that an address translation request of a second virtual address is received, the translation bypass cache is queried, and the merged entry is hit, reading the address region buffer identifier and the address high-order part corresponding to the second virtual address from the merged entry; acquiring the physical address high-order part corresponding to the address region buffer identifier from the address region buffer according to the address region buffer identifier; and obtaining the physical address corresponding to the second virtual address by combining the physical address high-order part and the address high-order part corresponding to the second virtual address. In this way, by using the ARB identifier to restore the complete physical address when the address translation is queried, the correct decoding and use of the compressed storage data are achieved while the query efficiency is maintained.

[0010] Optionally, in a further possible implementation of the first aspect, the storing the merged entry into the translation lookaside buffer comprises: extracting a bit sequence of a preset width from the first virtual address as an index; using the index to address a target set (SET) in the translation lookaside buffer; and storing the merged entry into the target set. In this way, by extracting a configurable bit-width sequence from the virtual address as an index, different TLB organization structures can be flexibly adapted, and efficient addressing and storage distribution of the merged entry in the TLB are ensured.

[0011] Optionally, in a further possible implementation of the first aspect, after the storing the merged entry into the translation lookaside buffer, the method further comprises: in a case where an address translation request of a third virtual address is received, reading a physical address corresponding to the third virtual address from the merged entry in the translation lookaside buffer, wherein the third virtual address is one of the remaining virtual addresses in the merged entry other than the first virtual address. In this way, by directly responding to a translation request of another virtual address in the entry after storing the merged entry, a mapping prefetch effect is achieved by using address space locality, and the delay of subsequent address translation is greatly reduced.

[0012] Optionally, in a further possible implementation of the first aspect, before the acquiring, according to the first virtual address, the plurality of continuous address mapping data, in a case where an address translation request of the first virtual address is received, the method further comprises: configuring a capacity proportion of the translation lookaside buffer available for storing the merged entry; and dividing the translation lookaside buffer into a merged entry storage area according to the capacity proportion of the merged entry, wherein the merged entry storage area is used to perform merging and storing of the merged entry. In this way, by partitioning the TLB storage area and limiting the storage range of the merged entry, a flexible area and performance trade-off mechanism is provided in a hardware resource limited environment.

[0013] The second aspect of the embodiment of the present application provides a device for increasing storage information density of a translation lookaside buffer, comprising: The acquiring module is configured to, in a case where an address translation request of a first virtual address is received, acquire, according to the first virtual address, a plurality of continuous address mapping data, wherein a request source corresponding to the first virtual address is an address request unit, and each address mapping data contains a mapping relationship between a virtual address and a physical address.

[0014] The merging module is configured to, in a case where the plurality of continuous address mapping data meet a preset merging condition, merge the plurality of continuous address mapping data into one merged entry, wherein the merged entry stores a same high-bit part of a virtual address, a same high-bit part of a physical address, and address low-bit parts and valid bits of each address mapping data respectively.

[0015] The processing module is configured to store the merged entry into the translation bypass cache, and return the target physical address corresponding to the first virtual address to the address request unit according to the merged entry.

[0016] A third aspect of the embodiments of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the processor implements the method for increasing the information density of the translation bypass cache storage of the first aspect when executing the computer program.

[0017] A fourth aspect of the embodiments of the present application provides a computer program product, which, when running on an electronic device, causes the electronic device to execute the method for increasing the information density of the translation bypass cache storage of the first aspect.

[0018] It can be understood that the beneficial effects of the second aspect to the fourth aspect can be referred to the related description in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0020] Figure 1 is a system architecture schematic diagram; Figure 2 is a TLB basic structure schematic diagram; Figure 3 is a virtual address and physical address mapping schematic diagram; Figure 4 is a system architecture schematic diagram provided by the embodiments of the present application; Figure 5 is a flowchart of the method for increasing the information density of the translation bypass cache storage provided by the embodiments of the present application; Figure 6 is a virtual address and physical address mapping schematic diagram provided by the embodiments of the present application; Figure 7 is an address region cache schematic diagram provided by the embodiments of the present application; Figure 8 is a TLB structure schematic diagram provided by the embodiments of the present application; Figure 9 is an entry comparison example diagram provided by the embodiments of the present application; Figure 10 is an entry example diagram provided by the embodiments of the present application; Figure 11 is a structural schematic diagram of a device for increasing the information density of a translation-lookaside buffer provided by an embodiment of the present application; Figure 12 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0021] In the following description, for purposes of explanation and not limitation, specific details are set forth, such as particular sequences of steps, techniques, etc., in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known methods, devices, circuits, and

[0022] It is to be understood that the terminology "includes", "has", "holds", "contains" or variants thereof, when used in this specification and in the following claims, indicates the presence of the described features, integers, steps, operations, elements, and / or components but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0023] It is also to be understood that the terminology "and / or" when used in this specification and in the following claims, refers to at least one of the items, or any combination of the items, listed after the term in the various aspects.

[0024] As used in this specification and in the claims, the term "if" can be interpreted as meaning "when", or "once", or "in response to determining", or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [a described condition or event] is detected" can be interpreted to mean "once it is determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]", depending on the context.

[0025] In addition, in the description of the specification and the appended claims, the terms "first", "second", "third", etc. are used only to distinguish descriptions, and cannot be understood as indicating or implying relative importance.

[0026] Reference in the specification to "one embodiment" or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearances of the phrase "in one embodiment" or "in some embodiments" in various places in the specification are not necessarily all referring to the same embodiment, although it can. The terms "including," "comprising," "having" and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items.

[0027] It should be understood that the size of the serial number of each step in the embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0028] Figure 1 A system architecture diagram is shown. As shown in the figure, the system includes: Figure 1 Memory management unit (MMU), a module inside the processor for managing virtual address translation physical address, used for processing access requests for virtual address translation physical address; Page table walk (PTW), a module for performing virtual address translation to physical address; Translation lookaside buffer (TLB), a storage cache for storing virtual address and physical address mapping relationship, which can speed up the processing of address translation requests compared to the page table walk unit; Instruction fetch unit (IFU), a module inside the processor for fetching instructions, specifically for fetching instructions from memory, in a typical system, accessing the instruction cache (ICache) in this unit requires obtaining a physical address, which will initiate a request for virtual address translation to physical address; Instruction cache (ICache), a storage unit for storing instructions, a smaller storage unit inside the processor for storing instructions, in this typical system, accessing this component requires obtaining the mapping of virtual address and physical address; Instruction translation lookaside buffer (ITLB), a storage unit for storing instruction virtual address and physical address mapping relationship, a smaller TLB than the translation lookaside buffer (TLB), used only for storing the mapping of instruction virtual address and instruction physical address; ​Load Store Unit (LSU) for executing load and store instructions, in a typical system, accessing the Data Translation Lookaside Buffer (DTLB) requires obtaining a physical address, which initiates a request for virtual to physical address translation; Data Cache (DCache) is a smaller storage unit inside the processor for storing data, in the typical system, accessing this component requires obtaining the mapping between virtual and physical addresses; Data Translation Lookaside Buffer (DTLB) is a storage unit for storing the mapping between data virtual and physical addresses, a smaller TLTB inside the LSU than the Translation Lookaside Buffer (TLB), only used for storing the mapping between data virtual and physical addresses.

[0029] Level 2 Cache (L2 Cache) is a larger cache, the MMU requests data for address translation from it.

[0030] Taking the instruction side as an example, the address generation unit (VA generator) generates a virtual address for accessing the ITLB and ICache. The ITLB is usually integrated inside the instruction fetch unit and is relatively small, usually containing 32 to 128 entries, and is commonly referred to as the L1 ITLB in processor microarchitecture. In a single clock cycle, the system performs two operations in parallel: on the one hand, it queries the ITLB, and if the address matches successfully, it outputs the corresponding physical address; on the other hand, it accesses the ICache to obtain the physical address and corresponding instruction content stored therein. When the physical addresses obtained by the two operations are consistent, the instruction output by the ICache is determined to be valid and is transmitted to the subsequent processing module.

[0031] When the ITLB query misses, the system sends an address translation request to the MMU. The MMU first queries the TLTB for instructions and data sharing stored inside it, which is a larger cache, usually containing more than 2000 entries. If a physical address is successfully matched in this level of cache, the result is immediately returned; if it still misses, the request is forwarded to the PTW to perform a complete address translation process. After the PTW completes the address translation, it stores the newly generated address mapping in the TLTB and returns the translated physical address to the original request initiator.

[0032] In the address translation process, the translation process of virtual address to physical address is based on page. Taking a typical 4KB page as an example, the virtual address VA[47:12] is translated into the physical address PA[47:12], and the address low bit VA[11:0] is not involved in the translation process as a page offset. The TLB adopts a set-associative structure organization, and its organization structure is determined by the index bit width parameter n. When the address translation request arrives, the system uses the VA[n-1:12] bit segment of the virtual address as the Index to select the target SET, and each SET contains y ways (WAY). After determining the SET, the system reads all the entries (the minimum storage unit for storing virtual address and physical address mapping) in the SET in parallel, and determines whether there is a correct address mapping by comparing the VA[47:n] bit segment with the address tag stored in each Entry. If there is a match, the corresponding physical address is returned, otherwise the request is sent to the PTW to perform complete address translation. The total capacity calculation formula of the TLB is Entries, wherein represents the number of SETs, and y represents the number of WAYs contained in each SET.

[0033] Figure 2 A basic structure diagram of a TLB is shown. As Figure 2 shown, a specific example with n=3, y=4 is illustrated. The TLB contains 8 SETs (SET0-SET7) in total, each equipped with 4 WAYs (WAY0-WAY3), forming a storage structure with a total capacity of 32 address mapping Entries. When the IFU or LSU initiates an address translation request, the virtual address carried is divided into three parts: VA[14:12] is used as the Index to select the SET, VA[47:15] is used as the tag (TAG) to compare with the tag stored in all the Entries in the selected SET, and VA[11:0] is used as the page offset. For example, when VA[14:12] value is 000, SET0 is accessed, value is 001, SET1 is accessed, and the rest of the SETs are accessed accordingly. Each Entry stores a complete virtual address to physical address mapping relationship, and the specific storage information includes: the virtual address tag of the VA[47:15] bit segment, the physical address page number of the PA[47:12] bit segment, and the valid (Valid) state bit indicating whether the Entry is valid. Figure 3 A virtual address and physical address mapping diagram is shown. As Figure 3 shown, the virtual address and physical address mapping is stored in the position of the TLB as Figure 2As shown, Entry0 stores the simplified information: VA0[47:15], PA0[47:12], Valid; Entry1 stores the simplified information: VA1[47:15], PA1[47:12], Valid; Entry2 stores the simplified information: VA2[47:15], PA2[47:12], Valid; and Entry3 stores the simplified information: VA3[47:15], PA3[47:12], Valid.

[0034] In the related art, the TLB adopts a structure of storing a single address mapping in a single entry, and each entry independently stores a complete virtual address high bit and a physical address high bit. When there are continuous virtual address mappings in the address space, this structure causes a large amount of same high bit address information to be repeatedly stored in different entries, resulting in significant waste of storage space, limiting the improvement of the effective capacity of the TLB under limited chip area, and further affecting the system processing performance.

[0035] Therefore, embodiments of the present application provide a method, device, electronic equipment and computer program product for increasing the storage information density of a translation lookaside buffer. First, in a case where a first virtual address is received, a plurality of continuous address mapping data are obtained according to the first virtual address, wherein the request source corresponding to the first virtual address is an address request unit, and each address mapping data contains a mapping relationship between a virtual address and a physical address. Then, in a case where the plurality of continuous address mapping data satisfy a preset merging condition, the plurality of continuous address mapping data are merged into a merged entry, wherein the merged entry stores a same virtual address high bit part, a same physical address high bit part in the plurality of continuous address mapping data, and an address low bit part and a valid bit corresponding to each address mapping data, respectively. Finally, the merged entry is stored in the translation lookaside buffer, and a target physical address corresponding to the first virtual address is returned to the address request unit according to the merged entry. In this way, by merging the plurality of continuous address mapping data into a single entry and sharing the storage of the same high bit part, the storage density of the effective address mapping is significantly improved while keeping the physical capacity of the TLB unchanged.

[0036] In order to illustrate the technical solutions of the present application, specific embodiments are described below.

[0037] Figure 4 is a system architecture schematic diagram provided by an embodiment of the present application. As shown in Figure 4 , compared with Figure 1The system shown in the embodiment of the present application adds a merging detection module in the second level cache, and the merging detection module is a control module for detecting the merging of the virtual address and the physical address. In addition, the system also adds an instruction translation bypass cache area buffer ITLB ARB, a data translation bypass cache area buffer DTLB ARB, and a second level translation bypass cache area buffer L2TLB ARB, which are respectively used for storing the high bits of the virtual address and the physical address of the ITLB, the DTLB, and the L2 TLB.

[0038] On the basis of the system architecture shown in the above Figure 4 Figure 5 A flowchart of a method for increasing the information density of the translation bypass cache provided by the embodiment of the present application is shown. The steps shown in the following are introduced. Figure 5

[0039] Step 501, in the case of receiving an address translation request of a first virtual address, a plurality of continuous address mapping data are acquired according to the first virtual address.

[0040] The request source corresponding to the first virtual address is an address request unit, and each address mapping data contains a mapping relationship between a virtual address and a physical address. It should be understood that the number of address mapping data needs to conform to the power of 2, such as 2, 4, 8, etc.

[0041] In the embodiment of the present application, the address request unit refers to a functional module that initiates an address translation request, which can be an instruction fetch unit IFU or a load store unit LSU. When the address request unit sends a translation request for the first virtual address, the system will acquire a plurality of continuous address mapping data based on the virtual address. These address mapping data are derived from the page table traversal process, and each address mapping data contains a complete corresponding relationship between a virtual address and a physical address. These virtual addresses are arranged continuously in the address space.

[0042] In one embodiment, when the IFU or the LSU needs to convert a virtual address into a physical address (PA), the respective L1 TLB (ITLB or DTLB) is first accessed. If the L1 TLB does not hit (i.e., no corresponding address mapping is found), the request is sent to the MMU. After receiving the request, the MMU checks the shared L2 TLB. If the L2 TLB hits, the physical address is directly returned. If it does not hit, the PTW is started to obtain address mapping data from the L2 Cache. Then the merging detection module intervenes to perform step 501.

[0043] ​​It should be noted that when the merging detection module receives a 64-bit data request from the MMU, the module will continuously obtain N 64-bit data units from the second level cache for the merging detection operation. The 64-bit data length is an industry standard, and its structural features are that a specific format of data storage area is accessed through a virtual address, and the corresponding physical address information is encoded in the 64-bit data unit. The final required physical address will be parsed and extracted from the 64-bit data unit.

[0044] In one embodiment, before step 501, the method further comprises, before obtaining the plurality of continuous address mapping data according to the first virtual address, configuring a capacity proportion of the translation bypass cache available for storing the merging entries; and dividing the translation bypass cache into a merging entry storage area according to the capacity proportion of the merging entries, wherein the merging entry storage area is used for subsequent execution of merging and storage of the merging entries. Thus, by partitioning the TLB storage area and limiting the storage range of the merging entries, a flexible area and performance trade-off mechanism is provided in a hardware resource limited environment.

[0045] It should be noted that in a conventional implementation, the entire TLB can be designed to support the structure of merging entries. It should be understood that the embodiments of the present application do not limit the proportion structure of the merging entries in the overall TLB capacity, and the TLB can contain a mixed architecture of merging entries and non-merging entries. In actual implementation, the number of merging entries can be configured to occupy a certain proportion of the total number of TLB entries, for example, a capacity allocation scheme of 50%. This design reflects the technical trade-off between storage density optimization and hardware overhead: although merging four address mappings into one merging entry can improve storage efficiency, merging operations will introduce additional circuit area overhead, and different application scenarios have different access patterns of address mappings. It should also be understood that if all entries in the TLB are merged, but there is a certain proportion of entries that cannot be merged, these entries that cannot be merged can only be placed in the merged entries, which will waste some resources, so there is also a design that fixes the TLB into two formats to balance this situation. The TLB can flexibly configure a mixed proportion storage scheme according to actual needs.

[0046] Step 502, in the case that the plurality of continuous address mapping data satisfies a preset merging condition, merging the plurality of continuous address mapping data into one merging entry.

[0047] The merging entry stores the same high bit part of the virtual address, the same high bit part of the physical address, and the low bit part and the valid bit corresponding to each address mapping data in the plurality of continuous address mapping data.

[0048] In the embodiments of the present application, after obtaining the plurality of continuous address mapping data, the system detects whether the data meets the preset merging condition. The merging condition mainly includes that the specific high bit part of all virtual addresses is completely same, and the specific high bit part of all physical addresses is also completely same. When the conditions are met, the system merges the address mapping data into one merged entry. In the merged entry, the high bit part of the virtual address and the high bit part of the physical address shared by all the mapping data are stored only once, and the low bit part and the valid bit of each mapping data are stored respectively, thereby forming a compressed storage structure.

[0049] In one embodiment, when the L2 Cache receives the address translation read data request, the merging detection module will perform the merging detection and merging operation of the page mapping. The processing flow is performed synchronously in the address translation process. In the traditional scheme, the MMU only obtains 64-bit data for single address translation each time; and when the method for increasing the information density of the translation bypass cache storage provided in the embodiments of the present application is used, 256-bit data (taking four merging as an example) is obtained at one time in the translation process, and then the merging condition detection is performed. If the detection meets the merging condition, the data is organized according to the specific merging format and the merged page mapping result is returned, and finally the MMU completes the actual page translation work and fills the merged page content back to the TLB / ITLB / DTLB. Through the merging mechanism which is completed synchronously in the translation process, the technical goal of storing the mapping relationship of multiple virtual addresses and physical addresses in a single Entry is achieved. Compared with the processing mode of four independent translation operations in the traditional technology, the embodiments of the present application can obtain four address mapping relationships in one translation process.

[0050] In one embodiment, it is necessary to determine whether the merging condition is met before merging. Specifically, according to the page size corresponding to the address translation request, the to-be-checked bit range of the virtual address and the physical address in the plurality of continuous address mapping data is determined; based on the plurality of continuous address mapping data, it is determined whether the high bit part of all the virtual addresses in the to-be-checked bit range is same, and whether the high bit part of all the physical addresses in the to-be-checked bit range is same; in the case that the high bit part of all the virtual addresses in the to-be-checked bit range is same, and the high bit part of all the physical addresses in the to-be-checked bit range is same, it is determined that the plurality of continuous address mapping data meets the merging condition.

[0051] Among them, please refer to Figure 6The virtual address and physical address mapping diagram is shown. Before implementing the address mapping merging operation, the system first needs to determine the specific merging detection rule according to the page size parameter provided by the MMU. When the MMU issues a 4KB page request, the 4KB page merging detection process is started. Taking the merging configuration of N=4 as an example, the system needs to verify whether the four consecutive address mapping data meet the specific structural characteristics: the lowest two bits VA[13:12] of the virtual addresses of these mappings must present consecutive numerical arrangement (00, 01, 10, 11 in turn), and the high bit part VA[47:14] of all virtual addresses remains completely the same - this feature is guaranteed by the consecutive address acquisition mechanism. In terms of physical address verification, the high bit part PA[47:14] of all corresponding physical addresses must be detected as consistent, while the low bit part PA[13:12] of the physical addresses is allowed to appear any value and does not require to maintain a specific order. Therefore, by dynamically determining the bit range to be checked according to the page size and performing the corresponding high bit consistency verification, it is ensured that the merging operation can adapt to different page size address mapping scenarios, improving the generality and reliability of the scheme.

[0052] It should be understood that if the MMU sends a 2MB page request, 2MB merging check is performed, and the difference between 2MB merging check and 4KB merging check is that the check bits are different, and the 2MB check VA and PA addresses are [47:21].

[0053] It should also be understood that the address size of the merged mapping can be 4KB / 16KB / 2MB, etc. The address size supported by the standard.

[0054] In step 503, the merging entry is stored in the translation bypass cache, and the target physical address corresponding to the first virtual address is returned to the address request unit according to the merging entry.

[0055] In the embodiments of the present application, the system stores the generated merging entry in the translation bypass cache, and completes the update of the address mapping information. At the same time, according to the merging entry stored in the foregoing embodiments, the system queries the target physical address corresponding to the first virtual address, and returns the physical address to the address request unit which originally initiated the request, completing the processing flow of the current address translation request.

[0056] It should be noted that, for example, Figure 7The address region cache diagram shown stores the high bits of the virtual address and the high bits of the physical address in a buffer, which is small, for example, 16 entries, each Entry stores an address high bit, here, 47:30 is taken as an address region, the stored Entry corresponds to a fixed ID, and finally the ID is stored in the TLB to reduce the number of stored VA and PA bits. The addresses stored here do not distinguish between VA and PA, different values exist in different Entries, and the IDs obtained for the same address are consistent.

[0057] In one embodiment, before storing the merged entry into the translation bypass cache, the method further includes: sending the same virtual address high bit part and the same physical address high bit part in the merged entry to an address region buffer in the translation bypass cache, wherein the same virtual address high bit part and the same physical address high bit part in the address region buffer are associated with a set address region buffer identifier ARB; and replacing the same virtual address high bit part and the same physical address high bit part in the merged entry with the address region buffer identifier. Thus, by introducing the address region buffer and replacing the storage of the actual high bit address value with the identifier, the storage space occupation of the merged entry is further compressed, achieving a two-level compression effect.

[0058] In one embodiment, after introducing the address region buffer and replacing the actual high bit address value with the identifier, if the address translation request of the second virtual address is received and the translation bypass cache is queried, and the merged entry is hit, the address region buffer identifier and the address bit part corresponding to the second virtual address are read from the merged entry; the physical address high bit part corresponding to the address region buffer identifier is obtained from the address region buffer according to the address region buffer identifier; and the physical address corresponding to the second virtual address is obtained by combining the physical address high bit part and the address bit part corresponding to the second virtual address. Thus, by using the ARB identifier to restore the complete physical address when querying the address translation, the correct decoding and use of the compressed storage data are realized while maintaining the query efficiency.

[0059] It should be noted that, for example, Figure 8The TLB structure diagram is shown. When the merging detection module returns the merging data to the MMU, the MMU stores the high bits of the virtual address and the high bits of the physical address into the L2 ARB, and respectively obtains the corresponding virtual address ARB ID and the physical address ARB ID. Taking a 4KB page as an example, the merging entry stored in the TLB specifically includes the following contents: an ARB identifier, four physical address low bits PA0[13:12] to PA3[13:12], and four valid bits Valid0 to Valid3. Through this storage structure, the mapping relationship of four virtual addresses to physical addresses is stored in a single entry. The same mechanism is adopted for the processing of a 2MB page, and after the system obtains the corresponding ARB identifier, the merging result is stored in the corresponding entry of the 2MB TLB. In addition, the 4KB page and the 2MB page are stored in separate structures, which can reduce the redundant information that needs to be stored for the 2MB page. For specific entry comparison, see, for example, Figure 9 The entry comparison example diagram is shown.

[0060] In the embodiments of the present application, under the same area budget, taking the merging of four address mapping data as an example, one entry can store the address mapping data that originally requires four entries to store. With the ARB technology, the area required for each page address mapping will be lower on average.

[0061] For example, the traditional Entry bit width is 100, the 4KB TLB with four merging and carrying the ARB technology has a bit width of 77, and the Entry bit width of the 2MB TLB with four merging and carrying the ARB technology is 58. If four merging is completely implemented, the average bit width of each mapping for a 4KB page is 19.25, and the average bit width of each page mapping for a 2MB page is 14.5. Compared with the traditional 100-bit width, the area is greatly reduced (the information density of storage is increased).

[0062] In an embodiment, a bit sequence of a preset width is extracted from a first virtual address as an index; the target group (SET) in the translation lookaside buffer is addressed using the index; and the merging entry is stored in the target group. In this way, by extracting a configurable bit width sequence from the virtual address as an index, different TLB organization structures can be flexibly adapted, and efficient addressing and storage distribution of the merging entry in the TLB are ensured. It should be understood that the specific structure of the TLB is not limited in the embodiments of the present application, for example, the number of WAYs, the number of bits of the index, or whether the L2 TLB itself distinguishes instruction L2 TLB and data L2 TLB, which can be determined in combination with actual application scenarios and requirements. In addition, the embodiments of the present application do not limit whether the L2 TLB is a paging structure, i.e., different pages are stored in separate structures.

[0063] In the embodiments of the present application, in the data merging process, since there are some completely same address information in the four address mappings, the system stores these same address parts to realize information compression. The specific stored information after merging includes: shared virtual address high bits VA[47:17], shared physical address high bits PA[47:14], four independent physical address low bits PA0[13:12] to PA3[13:12], and corresponding four valid bits Valid0 to Valid3. Among them, the shared virtual address bit segment VA[16:14] is used as the index of the TLB. The adjustment of the index bit width is derived from the characteristic requirements of the merging technology. In the traditional TLB structure, the virtual address bit segment VA[13:12] is used as the index, and the merging technology stores the four continuous mappings in a single Entry, so that VA[13:12] can reflect the difference within the Entry. In order to keep the TLB organization structure unchanged, it is necessary to select the higher bit VA[16:14] as the new index. The selection of this index bit width has configurability, and in actual implementation, other bit segments such as VA[18:14] can also be used. In the traditional scheme without merging, the system uses VA[14:12] as the index, and stores VA[47:15] in the Entry for matching comparison. In the merging scheme, the system uses VA[16:14] as the index, and completes the address matching by comparing the VA[47:17] and VA[13:12] of the access address with the corresponding bit segments stored in the Entry. Both schemes finally ensure that the complete virtual address VA[47:12] of the access is consistent with the stored address information, and the different parameter selection only affects the specific storage position of the address mapping in the TLB, without changing the correctness basis of the address translation.

[0064] It should be noted that part of the same virtual address bit segment needs to be used as an index. Since multiple entries are merged into a single storage entry, the indexes corresponding to these entries must be consistent. In fact, the merged address structure contains three distinct parts: the first part is the bit segment that allows differences, which in the case of four-merging is embodied in the fact that the two low bits of the virtual address can be different; the second part is the same address bit segment allocated as an index; and the third part is the same address bit segment stored inside the TLB entry. The bit width of the index can be flexibly selected according to design requirements, and different bit width values will directly affect the overall capacity of the TLB. Regarding the basic principle of address matching, the TLB access mechanism requires that the virtual address used when querying must be completely consistent with the stored virtual address. For example, the TLB stores two sets of mapping relationships of VA0-PA0 and VA1-PA1. When the system initiates an address query again, only when the input address is equal to VA0 or VA1 can the corresponding physical address PA0 or PA1 be successfully obtained. This mechanism guarantees the accuracy and reliability of the address translation result and is the basic condition for the normal operation of the TLB.

[0065] In an embodiment, in addition to being arranged in the second-level cache, the merging detection module can also be arranged in the MMU or in the DCache, and can be arranged in accordance with actual requirements. For example, when the module is placed inside the MMU, its processing flow is basically the same as the scheme of being arranged in the L2 Cache, the main difference being the transmission path from the unmerged data source to the processing module: at this time, the unmerged data obtained from the L2 Cache is directly transmitted to the merging detection module inside the MMU for processing, and the subsequent merging operation and data backfilling process remain unchanged. If the merging detection module is integrated in the DCache, the overall working process is exactly the same as the scheme of being arranged in the L2 Cache, at this time the main difference between the DCache and the L2 Cache lies in the storage capacity and access speed: the DCache has a smaller storage capacity but faster access speed, and the L2 Cache provides a larger storage capacity but relatively slower access speed. This setting scheme in different positions reflects the comprehensive trade-off of performance indicators and timing requirements in the chip design process.

[0066] In the embodiments of the present application, the system returns the merged data to the ITLB and the DTLB respectively after processing the address translation request. In the data returning process, the L2 ARB needs to be accessed to obtain the corresponding virtual address and physical address high bit information. When the data is returned to the ITLB or the DTLB, the address high bit information will be stored to the respective independent ITLB ARB or DTLB ARB. The three ARBs (L2 ARB, ITLB ARB and DTLB ARB) are completely independent in the hardware structure. In the specific execution process, the system performs data routing according to the request source identification: if the request comes from the ITLB, the merged data is returned to the ITLB; if the request comes from the DTLB, the data is returned to the DTLB. Each request contains a source ID to identify the initiator, and the MMU accurately judges the target object to which the data should be sent by identifying the source ID. The entry examples of the L1 ITLB and the L1 DTLB are shown in the following figure. Figure 10

[0067] In one embodiment, after storing the merged entry into the translation bypass cache, the physical address corresponding to the third virtual address can be read from the merged entry in the translation bypass cache in the case that an address translation request of the third virtual address is received, wherein the third virtual address is one of the remaining virtual addresses in the merged entry except the first virtual address. Thus, by directly responding to the translation request of the other virtual addresses in the entry after storing the merged entry, the mapping prefetch effect is achieved by using the address space locality, which greatly reduces the delay of subsequent address translation.

[0068] It should be noted that in actual operation, only the first virtual address in the merged address mapping data is immediately needed by the current request, and the remaining three address mappings are not the target of this access, but due to the spatial locality feature of the computer system access mode, that is, the processor tends to access adjacent memory space in a short period of time, these pre-fetched address mappings have a high probability of being used in subsequent operations. By merging four address mappings into the TLB entry in each address translation process, the system realizes the pre-storage of address mapping, and this mechanism uses the spatial locality principle to produce the prefetch effect, thereby significantly reducing the access delay of the subsequent address translation process.

[0069] ​The method for increasing the storage information density of the translation-lookaside buffer disclosed in the above embodiments of the present application first acquires a plurality of continuous address mapping data according to a first virtual address in the case of receiving an address translation request of the first virtual address, wherein the request source corresponding to the first virtual address is an address request unit, and each address mapping data contains a mapping relationship between a virtual address and a physical address; then merges the plurality of continuous address mapping data into one merged entry in the case that the plurality of continuous address mapping data satisfies a preset merging condition, wherein the merged entry stores the same high-bit part of the virtual address, the same high-bit part of the physical address, and the address low-bit part and the valid bit corresponding to each address mapping data respectively in the plurality of continuous address mapping data; and finally stores the merged entry into the translation-lookaside buffer and returns the target physical address corresponding to the first virtual address to the address request unit according to the merged entry. In this way, the plurality of continuous address mapping data is merged into a single entry and the same high-bit part is shared, thereby significantly improving the storage density of the effective address mapping while keeping the physical capacity of the TLB unchanged.

[0070] Referring to Figure 11 , a structure schematic diagram of a device for increasing the storage information density of the translation-lookaside buffer provided by the embodiments of the present application is shown, and only the parts related to the embodiments of the present application are shown for the convenience of description.

[0071] The device for increasing the storage information density of the translation-lookaside buffer can specifically include the following modules: The obtaining module 1101 is configured to acquire a plurality of continuous address mapping data according to a first virtual address in the case of receiving an address translation request of the first virtual address, wherein the request source corresponding to the first virtual address is an address request unit, and each address mapping data contains a mapping relationship between a virtual address and a physical address.

[0072] The merging module 1102 is configured to merge the plurality of continuous address mapping data into one merged entry in the case that the plurality of continuous address mapping data satisfies a preset merging condition, wherein the merged entry stores the same high-bit part of the virtual address, the same high-bit part of the physical address, and the address low-bit part and the valid bit corresponding to each address mapping data respectively in the plurality of continuous address mapping data.

[0073] The processing module 1103 is configured to store the merged entry into the translation-lookaside buffer and return the target physical address corresponding to the first virtual address to the address request unit according to the merged entry.

[0074] The device for increasing storage information density of translation bypass cache disclosed in the above embodiments of the present application first acquires a plurality of continuous address mapping data according to the first virtual address in the case of receiving an address translation request of the first virtual address, wherein the request source corresponding to the first virtual address is an address request unit, and each address mapping data contains a mapping relationship between a virtual address and a physical address; then merges the plurality of continuous address mapping data into one merged entry in the case that the plurality of continuous address mapping data meet a preset merging condition, wherein the merged entry stores the same high-bit part of the virtual address, the same high-bit part of the physical address, and the address low-bit part and the valid bit corresponding to each address mapping data respectively in the plurality of continuous address mapping data; and finally stores the merged entry into the translation bypass cache, and returns the target physical address corresponding to the first virtual address to the address request unit according to the merged entry. Thus, the plurality of continuous address mapping data are merged into a single entry and the same high-bit part is shared, thereby significantly improving the storage density of effective address mapping while keeping the physical capacity of the TLB unchanged.

[0075] Further, in a possible implementation manner of the embodiment of the present application, the device for increasing storage information density of translation bypass cache can further include the following modules: The first determining module is configured to determine the to-be-checked bit range of the virtual address and the physical address in the plurality of continuous address mapping data according to the page size corresponding to the address translation request.

[0076] The second determining module is configured to determine whether the address high-bit part of all the virtual addresses in the to-be-checked bit range is the same and whether the address high-bit part of all the physical addresses in the to-be-checked bit range is the same based on the plurality of continuous address mapping data.

[0077] The third determining module is configured to determine that the plurality of continuous address mapping data meet the merging condition in the case that the address high-bit part of all the virtual addresses in the to-be-checked bit range is the same and the address high-bit part of all the physical addresses in the to-be-checked bit range is the same.

[0078] Thus, the to-be-checked bit range is dynamically determined according to the page size and the corresponding high-bit part consistency verification is performed, thereby ensuring that the merging operation can adapt to different address mapping scenarios of different page specifications and improving the generality and reliability of the scheme.

[0079] Further, in another possible implementation manner of the embodiment of the present application, the device for increasing storage information density of translation bypass cache can further include the following modules: The first sending module is configured to send the same virtual address high part and the same physical address high part in the merged entry to an address region buffer in the translation bypass cache, wherein the same virtual address high part and the same physical address high part in the address region buffer are associated with a set address region buffer identifier.

[0080] The first processing module is configured to replace the same virtual address high part and the same physical address high part in the merged entry with the address region buffer identifier.

[0081] Therefore, by introducing the address region buffer and replacing the storage of the actual high address value with the identifier, the storage space occupation of the merged entry is further compressed, and a two-level compression effect is achieved.

[0082] Further, in another possible implementation of the embodiment of the application, the device for increasing the storage information density of the translation bypass cache can further include the following modules: The first obtaining module is configured to, in the case that the address translation request of the second virtual address is received and the translation bypass cache is queried and the merged entry is hit, read the address region buffer identifier and the address location part corresponding to the second virtual address from the merged entry.

[0083] The second obtaining module is configured to obtain the physical address high part corresponding to the address region buffer identifier from the address region buffer according to the address region buffer identifier.

[0084] The second processing module is configured to obtain the physical address corresponding to the second virtual address by combining the physical address high part and the address location part corresponding to the second virtual address.

[0085] Therefore, by using the ARB identifier to restore the complete physical address when the address translation is queried, the correct decoding and use of the compressed storage data are achieved while the query efficiency is maintained.

[0086] Further, in another possible implementation of the embodiment of the application, the processing module 1103 can include the following units: The third processing module is configured to extract a bit sequence of a preset width from the first virtual address as an index.

[0087] The fourth processing module is configured to use the index to address a target group in the translation bypass cache.

[0088] The first storage module is configured to store the merged entry into the target group.

[0089] Therefore, by extracting a configurable bit width sequence from the virtual address as an index, different TLB organization structures are flexibly adapted, and efficient addressing and storage distribution of the merged entry in the TLB are ensured.

[0090] Furthermore, in another possible implementation of this application embodiment, the above-mentioned device for increasing the density of translation bypass cache storage information may further include the following modules: The third acquisition module is used to read the physical address corresponding to the third virtual address from the merge entry in the translation bypass cache when an address translation request for the third virtual address is received. The third virtual address is one of the other virtual addresses in the merge entry besides the first virtual address.

[0091] Therefore, by directly responding to translation requests from other virtual addresses within an entry after storing and merging entries, the mapping prefetching effect is achieved by leveraging address space locality, which greatly reduces the latency of subsequent address translation.

[0092] Furthermore, in another possible implementation of this application embodiment, the above-mentioned device for increasing the density of translation bypass cache storage information may further include the following modules: The fifth processing module is used to configure the proportion of the translation bypass cache that can be used to store merged entries.

[0093] The sixth processing module is used to divide the translation bypass cache into a merged entry storage area according to the capacity ratio of the merged entries. The merged entry storage area is used to perform the merging and storage of merged entries.

[0094] Therefore, by partitioning the TLB storage area and limiting the storage range of merged entries, a flexible area and performance trade-off mechanism is provided in environments with limited hardware resources.

[0095] The device for increasing the density of translation bypass cache storage information provided in this application embodiment can be applied in the foregoing method embodiment. For details, please refer to the description of the above method embodiment, which will not be repeated here.

[0096] Figure 12 This is a schematic diagram of the structure of the electronic device provided in an embodiment of this application. For example... Figure 12 As shown, the electronic device 1200 of this embodiment includes: at least one processor 1210 ( Figure 12 The diagram shows only one processor, a memory 1220, and a computer program 1221 stored in the memory 1220 and executable on the at least one processor 1210. When the processor 1210 executes the computer program 1221, it implements the steps described above in the method embodiment for increasing the storage information density of the translation bypass cache.

[0097] The electronic device 1200 can be a desktop computer, a notebook computer, a palm computer, a cloud server, and the like. The electronic device can include, but is not limited to, a processor 1210, a memory 1220. Those skilled in the art can understand that Figure 12 The electronic device 1200 is only an example and does not limit the electronic device 1200, which can include more or fewer components than shown, or combine some components, or have different components, such as an input / output device, a network access device, and the like.

[0098] The processor 1210 can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, and the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0099] The memory 1220 can be an internal storage unit of the electronic device 1200 in some embodiments, such as a hard disk or a memory of the electronic device 1200. The memory 1220 can also be an external storage device of the electronic device 1200 in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, and the like. Further, the memory 1220 can include both an internal storage unit and an external storage device of the electronic device 1200. The memory 1220 is used to store an operating system, an application program, a boot loader, data, and other programs, such as program codes of the computer program, and the like. The memory 1220 can also be used to temporarily store data that has been output or will be output.

[0100] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be realized in the form of hardware or software. In addition, the specific names of each functional unit and module are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.

[0101] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.

[0102] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0103] In the embodiments provided in the present application, it should be understood that the disclosed devices / equipment and methods can be implemented in other ways. For example, the device / equipment embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0104] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0105] In addition, each of the function units in each of the embodiments of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.

[0106] The integrated module / unit, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer-readable storage medium. The computer program can implement the steps of each method embodiment when executed by a processor. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms. The computer-readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the computer-readable medium can include appropriate contents according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0107] The above-mentioned embodiment methods can also be completed by a computer program product, which, when running on an electronic device, causes the electronic device to perform the steps of each method embodiment.

[0108] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, rather than limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A method for increasing the information density of a translation bypass cache, characterized in that, include: Upon receiving an address translation request for a first virtual address, multiple consecutive address mapping data are obtained based on the first virtual address, wherein the request source corresponding to the first virtual address is an address request unit, and each address mapping data contains a mapping relationship between a virtual address and a physical address; When multiple consecutive address mapping data meet a preset merging condition, the multiple consecutive address mapping data are merged into a merging entry. The merging entry stores the same high-order virtual address portion, the same high-order physical address portion, and stores the low-order address portion and valid bit corresponding to each of the multiple consecutive address mapping data. In the merged entry, the same high-order virtual address portion and the same high-order physical address portion are replaced by an address region buffer identifier, wherein the same high-order virtual address portion, the same high-order physical address portion and the address region buffer identifier are associatedly stored in the address region buffer corresponding to the translation bypass cache; The merged entry is stored in the translation bypass cache, and the target physical address corresponding to the first virtual address is returned to the address request unit according to the merged entry.

2. The method according to claim 1, characterized in that, Before merging the multiple consecutive address mapping data into a single merge entry when the multiple consecutive address mapping data meet the preset merging conditions, the method further includes: Based on the page size corresponding to the address translation request, determine the range of virtual and physical address bits to be checked in the multiple consecutive address mapping data; Based on the multiple consecutive address mapping data, it is determined whether the high-order bits of all virtual addresses are the same within the range to be checked, and it is also determined whether the high-order bits of all physical addresses are the same within the range to be checked. If all virtual addresses have the same high-order bits within the range to be checked, and all physical addresses have the same high-order bits within the range to be checked, then the plurality of consecutive address mapping data are determined to satisfy the merging condition.

3. The method according to claim 1, characterized in that, The method further includes: Upon receiving an address translation request for the second virtual address and querying the translation bypass cache, and finding a match in the merge entry, the address region buffer identifier and the address position portion corresponding to the second virtual address are read from the merge entry. Based on the address region buffer identifier, obtain the high-order part of the physical address corresponding to the address region buffer identifier from the address region buffer; By combining the high-order part of the physical address and the low-order part of the address corresponding to the second virtual address, the physical address corresponding to the second virtual address is obtained.

4. The method according to claim 1, characterized in that, The step of storing the merged entries in the translation bypass cache includes: From the first virtual address, extract a bit sequence of a preset width as an index; Use the index to address the target group in the translation bypass cache; The merged entries are stored in the target group.

5. The method according to claim 1, characterized in that, After storing the merged entry in the translation bypass cache, the method further includes: Upon receiving an address translation request for a third virtual address, the physical address corresponding to the third virtual address is read from the merge entry in the translation bypass cache, wherein the third virtual address is one of the other virtual addresses in the merge entry besides the first virtual address.

6. The method according to claim 1, characterized in that, Before obtaining multiple consecutive address mapping data based on the first virtual address upon receiving an address translation request for the first virtual address, the method further includes: Configure the proportion of the translation bypass cache that can be used to store merged entries; Based on the capacity ratio of the merged entries, the translation bypass cache is divided into a merged entry storage area, wherein the merged entry storage area is used to perform the merging and storage of the merged entries.

7. The method according to claim 1, characterized in that, The address request unit is either an instruction fetch unit or a load-store unit. The instruction fetch unit includes an instruction translation bypass cache area buffer, and the load-store unit includes a data translation bypass cache area buffer. After storing the merged entry in the translation bypass cache and returning the target physical address corresponding to the first virtual address to the address request unit according to the merged entry, the method further includes: When the request source of the first virtual address is an instruction fetching unit, the same high-order part of the virtual address and the same high-order part of the physical address are stored in the instruction translation bypass cache area buffer. When the request source of the first virtual address is a loading storage unit, the same high-order part of the virtual address and the same high-order part of the physical address are stored in the data translation bypass cache area buffer.

8. An apparatus for increasing the information density of a translation bypass cache, characterized in that, include: The acquisition module is used to acquire multiple consecutive address mapping data based on the first virtual address when receiving an address translation request for the first virtual address, wherein the request source corresponding to the first virtual address is an address request unit, and each address mapping data contains a mapping relationship between a virtual address and a physical address. The merging module is used to merge multiple consecutive address mapping data into a single merge entry when the multiple consecutive address mapping data meet a preset merging condition. The merge entry stores the same high-order virtual address portion, the same high-order physical address portion, and the low-order address portion and valid bit corresponding to each of the multiple consecutive address mapping data. The first processing module is used to replace the storage of the same high-order virtual address portion and the same high-order physical address portion in the merged entry with an address region buffer identifier, wherein the same high-order virtual address portion, the same high-order physical address portion and the address region buffer identifier are associatedly stored in the address region buffer corresponding to the translation bypass cache; The processing module is used to store the merged entry in the translation bypass cache, and return the target physical address corresponding to the first virtual address to the address request unit according to the merged entry.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Page table entry merging method and device and electronic equipment

    CN111949572A

  • TLB entry merging method and address translation method

    CN114780452A

  • Address translation method executed by processor and related product

    CN116860665A

  • Address conversion method and device, electronic equipment and storage medium

    CN118568014A

  • Physical address proxy reuse management

    US20220358037A1