Address translation addressing method, apparatus, and electronic device
Patent Information
- Application Number
- CN202610480435.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-10
- Publication Date
- 2026-08-28
AI Technical Summary
[0020] This application predicts the first page granularity identifier corresponding to the address translation request, searches the storage array based on the first page granularity identifier to obtain candidate address translation entries, and performs a consistency check on the first page granularity identifier and the second page granularity identifier in the candidate address translation entries. In response to the check result indicating that the first page granularity identifier and the second page granularity identifier are consistent, the candidate address translation entry is determined to be the target address translation entry obtained through addressing. This application achieves single-path lookup through a prediction mechanism, significantly reducing the dynamic flip power consumption when supporting multi-granularity mapping, solving the energy efficiency bottleneck caused by multi-granularity support, and eliminating addressing conflicts between pages of different granularities when sharing a storage array while maintaining low-latency lookup.
Smart Images

Figure CN122654022A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to the fields of data processing and digital integrated circuit technology. Background Technology
[0002] With the evolution of data centers and artificial intelligence (AI) computing loads, system memory access patterns are showing a significant polarization in demand for "high coverage" and "high precision": on the one hand, ultra-large-scale static data such as AI model weights require hardware support for ultra-large pages of 512MB or even 1GB to improve the coverage of the Translation Lookaside Buffer (TLB) and reduce translation overhead; on the other hand, dynamic fragmented data such as key-value cache (KV Cache) relies on small pages of 4KB or 64KB to ensure memory utilization.
[0003] Therefore, in scenarios that require compatibility with multiple page granularity spans, how to eliminate addressing conflicts between pages of different granularities when sharing a storage array while maintaining low-latency lookups, and how to break down the allocation barriers of hardware resources under different granularity requests, has become one of the important research directions. Summary of the Invention
[0004] This disclosure provides an address translation addressing method, apparatus, and electronic device.
[0005] According to one aspect of this disclosure, an address translation addressing method is provided, comprising:
[0006] Obtain the address translation request and predict the first page granularity identifier corresponding to the address translation request; The storage array is searched based on the first page granularity identifier to obtain candidate address translation entries. The address translation entries include address labels, physical addresses, and a second page granularity identifier used to represent the actual page size. A consistency check is performed on the first page granularity identifier and the second page granularity identifier in the candidate address translation entry. In response to the check result indicating that the first page granularity identifier and the second page granularity identifier are consistent, the candidate address translation entry is determined to be the target address translation entry obtained by addressing.
[0007] In some implementations, it also includes: In response to the verification result indicating that the first page granularity identifier and the second page granularity identifier are inconsistent, the current error waiting is terminated, prediction deviation correction and redundancy control are performed, and a retransmission request is initiated. The retransmission request is used to reprocess the address translation request.
[0008] In some implementations, prediction bias correction and redundancy control are performed, and retransmission requests are initiated, including: Merge redundant page table lookup requests in the page table lookup queue and clear the address translation cache; Correct the granularity identifier of the first page to obtain the granularity identifier of the third page; Initiate a retransmission request based on the granularity identifier on the third page to update the candidate address translation entries.
[0009] In some implementations, the storage array is searched based on a first page granularity identifier to obtain candidate address translation entries, including: Dynamically extract different bit segments from the virtual address as group indexes based on the granularity identifier of the first page; Retrieve the storage array based on the group index to obtain candidate address translation entries.
[0010] In some implementations, different bit segments in the virtual address are dynamically extracted as group indexes based on the first page granularity identifier, including: Determine the page offset based on the granularity identifier of the first page; The group index is the bit segment after removing the page offset from the virtual address.
[0011] In some implementations, different bit segments in the virtual address are dynamically extracted as group indexes based on the first page granularity identifier, including: In response to the first page granularity identifier indicating a super-large page, the high-order bits of the virtual address are extracted as a group index.
[0012] In some implementations, retrieving the storage array based on the group index to obtain candidate address translation entries includes: The storage array is retrieved based on the group index. In response to a missing address translation entry in the group index, a multi-level page table traversal operation is initiated to update the address mapping relationship of the storage array. Retrieve the updated storage array based on the group index to obtain candidate address translation entries.
[0013] In some implementations, the first page granularity identifier corresponding to the predicted address translation request includes: Multi-dimensional feature extraction is performed on the address translation request to obtain the associated feature information of the address translation request. The associated feature information includes at least one of the following: instruction sequence features, register usage features, address space distribution features, or memory access status. Based on the configured prediction algorithm, the associated feature information is predicted to obtain the first page granularity identifier corresponding to the address translation request.
[0014] According to another aspect of this disclosure, an address translation addressing apparatus is provided, comprising: The address translation control module is used to obtain address translation requests and predict the first page granularity identifier corresponding to the address translation request; The index generation and shifting module is used to dynamically extract different bit segments from the virtual address as a group index based on the first page granularity identifier; The address translation control module is also used to retrieve the storage array according to the group index to obtain candidate address translation entries, and to perform consistency verification on the first page granularity identifier and the second page granularity identifier in the candidate address translation entries; in response to the verification result indicating that the first page granularity identifier and the second page granularity identifier are consistent, the candidate address translation entry is determined to be the target address translation entry obtained by addressing; or, in response to the verification result indicating that the first page granularity identifier and the second page granularity identifier are not consistent, the current error waiting is terminated, prediction deviation correction and redundancy control are performed, and a retransmission request is initiated.
[0015] In some embodiments, the address translation addressing device further includes: The page table lookup module is used to initiate a multi-level page table traversal operation in response to a group index miss address translation entry, in order to update the address mapping relationship of the storage array.
[0016] In some implementations, the address translation control module includes: The address translation queue is used to cache address translation requests initiated by the computing cores; The page granularity prediction unit is used to perform multi-dimensional feature extraction on the address translation request, obtain the associated feature information of the address translation request, predict the associated feature information according to the configured prediction algorithm, and obtain the first page granularity identifier corresponding to the address translation request. The status management unit is used to manage the real-time translation status of each address translation request in the pipeline. The translation status includes idle status, hit status, page table lookup waiting status, and hit status after miss. The retransmission and merging unit is used to respond to the verification result indicating that the first page granularity identifier and the second page granularity identifier are inconsistent, merge redundant page table lookup requests in the page table lookup queue, and clear the address translation cache; correct the first page granularity identifier and obtain the third page granularity identifier; and initiate a retransmission request based on the third page granularity identifier to update the candidate address translation entries. The virtual address request interface is used to standardize and encapsulate the virtual address translation information corresponding to the target address translation entry.
[0017] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the address translation addressing method described above.
[0018] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the address translation addressing method described above.
[0019] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described address translation addressing method.
[0020] This application predicts the first page granularity identifier corresponding to the address translation request, searches the storage array based on the first page granularity identifier to obtain candidate address translation entries, and performs a consistency check on the first page granularity identifier and the second page granularity identifier in the candidate address translation entries. In response to the check result indicating that the first page granularity identifier and the second page granularity identifier are consistent, the candidate address translation entry is determined to be the target address translation entry obtained through addressing. This application achieves single-path lookup through a prediction mechanism, significantly reducing the dynamic flip power consumption when supporting multi-granularity mapping, solving the energy efficiency bottleneck caused by multi-granularity support, and eliminating addressing conflicts between pages of different granularities when sharing a storage array while maintaining low-latency lookup.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0022] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a schematic diagram illustrating an exemplary implementation of an address translation addressing method shown in this application; Figure 2 This is a schematic diagram illustrating an exemplary implementation of an address translation addressing method shown in this application; Figure 3 This is a schematic diagram illustrating an exemplary implementation of an address translation addressing method shown in this application; Figure 4 This is a schematic diagram of an address translation addressing system according to an exemplary embodiment of the present disclosure; Figure 5This is a schematic diagram of an address translation addressing apparatus according to an exemplary embodiment of the present disclosure; Figure 6 This is a structural block diagram of an address translation control module according to an exemplary embodiment of the present disclosure; Figure 7 This is a schematic diagram illustrating an exemplary implementation of an address translation addressing method shown in this application; Figure 8 This is a schematic diagram illustrating an exemplary implementation of a retransmission and error correction process after prediction failure as shown in this application; Figure 9 This is a schematic diagram of an exemplary implementation of a cross-group redundancy elimination and PTW merging process after prediction failure, as shown in this application. Figure 10 This is a schematic diagram of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0023] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0024] Artificial Intelligence (AI) is a new technological science that studies and develops theories, methods, technologies, and application systems to simulate, extend, and expand human intelligence. AI is a branch of computer science that attempts to understand the nature of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems.
[0025] Data processing is a fundamental component of systems engineering and automatic control. It permeates all areas of social production and life. The development of data processing technology and the breadth and depth of its applications have profoundly influenced the progress of human society. Data is a form of expression of facts, concepts, or instructions, which can be processed manually or automatically. After being interpreted and given meaning, data becomes information. Data processing involves the collection, storage, retrieval, processing, transformation, and transmission of data. The fundamental purpose of data processing is to extract and derive valuable and meaningful data from large amounts of potentially chaotic and difficult-to-understand data.
[0026] Digital integrated circuits are digital logic circuits or systems made by integrating components and interconnects onto the same semiconductor chip.
[0027] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0028] This application can be applied to the Memory Management Unit (MMU), System Memory Management Unit (SMMU), and Input / Output Memory Management Unit (IOMMU).
[0029] Figure 1 This is a schematic diagram illustrating an exemplary implementation of an address translation addressing method shown in this application, such as... Figure 1 As shown, this address translation addressing method includes the following steps: S101, obtain the address translation request and predict the first page granularity identifier corresponding to the address translation request.
[0030] Accept the address translation request initiated by the computing core.
[0031] For each address translation request, perform page-level prediction, such as analyzing the contextual features of the current memory access request (including but not limited to instruction sequence features, register usage features, or address space distribution features) to obtain the predicted first page-level identifier.
[0032] It should be noted that different page-level prediction logics will result in different prediction accuracies, and low prediction accuracy will lead to corresponding performance losses.
[0033] In this embodiment, before actually retrieving the storage array, the most likely page granularity of the address translation request is pre-locked, thereby guiding the hardware to select a specific addressing path, shifting from "parallel retrieval" to "precise guidance", without needing to drive multiple physical arrays or complex parallel comparison circuits simultaneously.
[0034] S102, the storage array is searched according to the first page granularity identifier to obtain candidate address translation entries.
[0035] The address translation entries include address tags, physical addresses, and a second page granularity identifier used to represent the actual page size.
[0036] In this embodiment of the application, a group index is obtained based on the first page granularity identifier, and then the address mapping relationship of the storage array is retrieved based on the group index. The address translation entry corresponding to the group index is obtained based on the address mapping relationship as a candidate address translation entry.
[0037] S103, perform a consistency check on the first page granularity identifier and the second page granularity identifier in the candidate address translation entry. In response to the check result indicating that the first page granularity identifier and the second page granularity identifier are consistent, determine the candidate address translation entry as the target address translation entry obtained by addressing.
[0038] In the address translation entries of the unified storage array, in addition to the regular address tag and physical address, an additional real-granularity field is defined. This field serves as the hardware standard for address translation and is used for real-time comparison with the predicted identifier during the retrieval phase, providing logical support for subsequent error interception.
[0039] Since the predicted first page granularity identifier may be biased, the hardware cannot yet definitively determine whether the hit is absolutely correct. It needs to wait for the actual page table information to be returned and perform a granularity consistency comparison to determine whether the candidate address translation entry is the target address translation entry obtained from the addressing.
[0040] In response to the verification result indicating that the first page granularity identifier and the second page granularity identifier are consistent, it means that the verification result is successful. This page table translation has found the correct physical page table. Subsequently, the physical address corresponding to the physical page table can be sent out. This is equivalent to the upstream address translation process being completed and successfully sent to the downstream.
[0041] In some implementations, the first page granularity identifier and the second page granularity identifier are the same, indicating that the first page granularity identifier and the second page granularity identifier are consistent.
[0042] In some implementations, the difference between the page granularity indicated by the first page granularity identifier and the second page granularity identifier is less than a preset page granularity threshold, indicating that the first page granularity identifier and the second page granularity identifier are consistent.
[0043] This application predicts the first page granularity identifier corresponding to the address translation request, searches the storage array based on the first page granularity identifier to obtain candidate address translation entries, and performs a consistency check on the first page granularity identifier and the second page granularity identifier in the candidate address translation entries. In response to the check result indicating that the first page granularity identifier and the second page granularity identifier are consistent, the candidate address translation entry is determined to be the target address translation entry obtained through addressing. This application achieves single-path lookup through a prediction mechanism, significantly reducing the dynamic flip power consumption when supporting multi-granularity mapping, solving the energy efficiency bottleneck caused by multi-granularity support, and eliminating addressing conflicts between pages of different granularities when sharing a storage array while maintaining low-latency lookup.
[0044] This application can be applied to the design of memory management units in chips such as high-performance AI accelerators, general-purpose processors (CPUs), graphics processing units (GPUs), and data center-level SoCs. Here, SoC stands for System-on-a-Chip.
[0045] Figure 2 This is a schematic diagram illustrating an exemplary implementation of an address translation addressing method shown in this application, such as... Figure 2 As shown, this address translation addressing method includes the following steps: S201, obtain the address translation request and predict the first page granularity identifier corresponding to the address translation request.
[0046] S202, the storage array is searched according to the first page granularity identifier to obtain candidate address translation entries. The address translation entries include address labels, physical addresses, and a second page granularity identifier used to represent the actual page size.
[0047] S203, perform a consistency check on the first page granularity identifier and the second page granularity identifier in the candidate address translation entry. In response to the check result indicating that the first page granularity identifier and the second page granularity identifier are consistent, determine the candidate address translation entry as the target address translation entry obtained by addressing.
[0048] For a description of steps S201 to S203, please refer to the relevant content in the above embodiments, which will not be repeated here.
[0049] S204, in response to the verification result indicating that the first page granularity identifier and the second page granularity identifier are inconsistent, the current error waiting is terminated, prediction deviation correction and redundancy control are performed, and a retransmission request is initiated.
[0050] The retransmission request is used to reprocess the address translation request.
[0051] The verification result indicates that the first page granularity identifier and the second page granularity identifier are inconsistent, which means that the verification result failed. This page table translation resulted in a false hit and the correct physical page table was not found.
[0052] Redundant page table lookup requests in the page table lookup queue are merged, the address translation cache is cleared, the first page granularity identifier is corrected, the third page granularity identifier is obtained, and a retransmission request is initiated based on the third page granularity identifier to update the candidate address translation entries.
[0053] For example, if a false hit is detected due to a small page request (a small page address translation request) being mispredicted as a large page, the error response is intercepted. Once the pending page table lookup (PTW) returns the true result, requests in the "Hit-without-miss" (HUM) state undergo granular strong validation. If the validation fails, the current error wait is aborted, and a retransmission is forcibly initiated. In the above possible implementations, closed-loop retransmission control can suppress pipelined invalid waits caused by false HUMs, ensuring absolute addressing accuracy.
[0054] For example, if a large page request (a request to translate the address of a large page) misses due to misprediction and enters the wrong set, after the PTW returns the true mapping relationship, the bootstrap hardware writes the mapping entry (including physical address, page size, permission information, etc.) to the correct physical set and cleans up redundant copies in the array caused by misprediction. In the above possible implementations, PTW merging under multiple conflicting sets is achieved through base address mask matching and broadcast cleanup logic, reducing the memory bandwidth pressure caused by repeated memory accesses.
[0055] In this embodiment, the consistency between the predicted granularity and the actual granularity in the stored entries is monitored in real time. When a prediction deviation is detected, a retransmission request is initiated and the system is guided to restore the state.
[0056] This application can suppress false hits and erroneous HUM retransmissions. For scenarios where small page requests are mistakenly predicted as large page requests, it intercepts "false hit" results through granular strong validation. When a suspended large page PTW returns a result and is written to the TLB, granular strong validation is performed on all requests in the HUM state. Once a granularity mismatch is detected (i.e., a "false HUM" occurs), the current erroneous wait is immediately terminated, the request is released, and it is forced to retransmit with the true granularity identifier.
[0057] This application can eliminate index redundancy and mapping bias. For large page requests that enter the wrong set due to misprediction, this mechanism guides the hardware to write the mapping entry to the correct physical set after the page table lookup (PTW) returns, and cleans up potential redundant copies.
[0058] In response to the verification result indicating that the first page granularity identifier and the second page granularity identifier are inconsistent, this application terminates the current error waiting, performs prediction deviation correction and redundancy control, and initiates a retransmission request. Through closed-loop retransmission and merging control, it ensures that the system can quickly converge to the correct path when prediction fails, thus balancing high prediction performance with absolute accuracy of hardware execution.
[0059] Figure 3 This is a schematic diagram illustrating an exemplary implementation of an address translation addressing method shown in this application, such as... Figure 3As shown, this address translation addressing method includes the following steps: S301, Get address translation request.
[0060] For a description of step S301, please refer to the relevant content in the above embodiments, which will not be repeated here.
[0061] S302, perform multi-dimensional feature extraction on the address translation request to obtain the associated feature information of the address translation request.
[0062] The associated feature information includes at least one of the following: instruction sequence features, register usage features, address space distribution features, or memory access status.
[0063] Among the above possible implementation methods, multi-dimensional feature extraction can improve the recognition accuracy of page granularity in different application scenarios during subsequent prediction.
[0064] S303, predict the associated feature information according to the configured prediction algorithm, and obtain the first page granularity identifier corresponding to the address translation request.
[0065] In some embodiments, the associated feature information is mainly the current memory access state. For example, a prediction is made based on a preset prediction algorithm or rule (including but not limited to the context memory access state, etc.), and the predicted first page granularity identifier is output.
[0066] Let's illustrate this with various heterogeneous page granularities including 4KB, 64KB, 2MB, 512MB, and 1GB. For example, if all the page tables found initially are 4KB, then the next address translation request will also be predicted to have a 4KB page table. As another example, if the previous address translation request corresponds to a 2MB page table, then other virtual addresses within the same 1GB range as this 2MB page table will definitely not have 1GB page tables, because 2MB of this 1GB page table has already been found, and there is no complete 1GB page table left; therefore, it can only be 2MB / 4KB in size.
[0067] S304: Dynamically extract different bit segments from the virtual address as group indexes based on the first page granularity identifier.
[0068] Based on different page granularities, virtual page numbers of different bit segments are dynamically extracted from the virtual address to serve as the group index for the unified storage array. Therefore, by dynamically adjusting the extraction position of the index bits, optimal load balancing distribution of page entries of different granularities can be ensured within the physical array.
[0069] In some implementations, the page offset is determined based on a first page granularity identifier. The bit field after removing the page offset from the virtual address is used as a group index.
[0070] In some implementations, in response to the first page granularity identifier indicating an oversized page, the high-order bits of the virtual address are extracted as a group index. In the above possible implementations, by shifting the index bits to the valid page number range, the problem of all-zero low-order bits in the index caused by natural address alignment can be avoided, thus preventing mirror redundancy.
[0071] The implementation of index bit shifting can be achieved in hardware using a simple shift circuit. For example, if the predicted page table size is 1GB, the bits from the lower 30 bits of the address are used as index bits; if the predicted page table size is 2MB, the bits from the lower 21 bits of the address are used as index bits; and if the predicted page table size is 4KB, the bits from the lower 12 bits of the address are used as index bits. Since 1 bit of the address represents 1 byte of address space, a 4KB granularity translates to 12 bits in the address space. The same principle applies to other page table sizes.
[0072] S305, retrieve the storage array based on the group index to obtain candidate address translation entries.
[0073] In some implementations, the storage array is retrieved based on the group index. In response to a mismatch in the address translation entries within the group index, a multi-level page table traversal is initiated to update the address mapping of the storage array. The updated storage array is then retrieved again based on the group index to obtain candidate address translation entries.
[0074] In this application, storage entries are read from the corresponding physical group of the storage array based on the group index for tag comparison. In the embodiments of this application, the storage array of the page table cache adopts a unified physical array structure, supports the mixed storage of storage entries with heterogeneous granularity within a single array, and uses the real granularity identifier within the entry for post-validation.
[0075] When an address translation request (i.e., a memory access request) misses in the memory array, a traversal operation of the multi-level page tables in the physical memory is initiated to obtain the latest address mapping relationship and backfill it into the memory array of the page table cache. This can improve the efficiency and accuracy of address translation addressing.
[0076] The information cached in the storage array is a mapping relationship. For example, the TLB caches "xx virtual address" to "yy physical address". Then, the address translation request received by the SMMU will use its own virtual address as a tag to search in the TLB to see if this virtual address has been cached. If the tag matches, the mapped physical address will be found directly.
[0077] S306, perform a consistency check on the first page granularity identifier and the second page granularity identifier in the candidate address translation entry. In response to the check result indicating that the first page granularity identifier and the second page granularity identifier are consistent, determine the candidate address translation entry as the target address translation entry obtained by addressing.
[0078] For a description of step S306, please refer to the relevant content in the above embodiments, which will not be repeated here.
[0079] The index retrieval fields of a unified set-associative TLB are static and fixed. Since the mapping granularity is unknown during the retrieval phase, the hardware is forced to use the smallest page granularity indexing rule. This inevitably leads to group index stacking in single-entry storage mode for large page mappings, or mirror redundancy in multi-set retrieval mode. This application breaks this temporal dependency through a prediction mechanism, enabling the hardware to dynamically adjust the index retrieval fields based on the prediction results at the moment the retrieval is initiated. This not only ensures that the index logic matches the actual page granularity but also achieves uniform distribution of different large page mappings across the array without the need for mirrored storage. This reduces the probability of group index collisions for ultra-large pages, significantly improving the actual effective capacity of the hardware array compared to traditional solutions.
[0080] This application achieves efficient and unified management of various heterogeneous page granularities, including but not limited to 4KB, 64KB, 2MB, 512MB, and 1GB, in a single array by integrating a page size prediction mechanism, dynamic index shifting logic, and abnormal retransmission and redundancy control mechanism into the hardware addressing pipeline.
[0081] By transforming traditional multi-granularity parallel retrieval into single-path precise guidance through a predictive mechanism, the power consumption of hardware dynamic flipping is significantly reduced when supporting various heterogeneous granularities (such as 4KB, 64KB, 2MB, 512MB, 1GB, etc.). Simultaneously, by utilizing a dynamic index shifting mechanism, page entries of different granularities can be evenly distributed across the storage array, fundamentally solving the set aliasing problem of ultra-large pages and improving the actual effective capacity of the hardware array.
[0082] Figure 4 This is a schematic diagram of an address translation addressing system according to an exemplary embodiment of the present disclosure, such as... Figure 4 As shown, the address translation addressing system 400 of this application can be applied in high-performance computing (HPC), data centers, or large-model inference accelerators. The address translation addressing system 400 includes: Computational core 401 is used to generate memory access requests carrying virtual addresses (VAs). In one possible implementation, computational core 401 can be a central processing unit (CPU), a graphics processing unit (GPU), or a neural network processor (NPU). In another embodiment, computational core 401 can also be any logical module capable of actively initiating virtual address memory access requests, such as a direct memory access (DMA) controller or a specific hardware acceleration unit.
[0083] Address translation and addressing device 402 (hereinafter referred to as SMMU device 402) is coupled between computing core 401 and on-chip network 403. SMMU device 402 is configured to perform page-level prediction and dynamic shift addressing on the received virtual address, thereby achieving fast conversion between virtual address and physical address (PA). It should be noted that, in one embodiment of this application, the SMMU device can be implemented using a distributed architecture. Figure 4 As shown, each computing core is independently configured with a corresponding SMMU device. This point-to-point coupling effectively reduces addressing latency and improves local memory access bandwidth. However, in some other possible embodiments, the physical implementation of the SMMU device is not limited to this. For example, multiple computing cores can be coupled to a unified SMMU device via some arbitration logic or bus. Under this unified architecture, the SMMU device provides centralized address translation services for multiple different computing cores through multi-stream management technology. The above-mentioned choice between distributed or unified implementations falls within the scope of protection of the inventive concept of this application.
[0084] Optionally, a single SMMU can be responsible for address translation for multiple upstream modules, depending on the hardware design, the number of address translation request ports for each upstream module, and so on.
[0085] The Network on Chip (NoC) 403 is coupled between the SMMU device 402 and the physical memory 404. As a system-level communication interconnection architecture, the Network on Chip 403 is responsible for routing memory access requests carrying physical addresses output by the SMMU device 402 to the corresponding memory controllers or subsystems, and supports high-bandwidth data interaction between various hardware modules.
[0086] Physical memory 404 (such as dynamic random access memory DRAM or high bandwidth memory HBM) exchanges signals with SMMU device 402 via on-chip network 403. Physical memory 404 is used to store actual service data content and a multi-level page table structure to support address mapping.
[0087] It is understandable that the above Figure 4The structure of the address translation system 400 shown is merely exemplary. In practical applications, the system 400 may include more or fewer components than illustrated. For example, the system 400 may also include a cache coherence controller, a page table lookup (PTW) unit, or a multi-level cache. Alternatively, the aforementioned components may be combined or arranged in different ways. Figure 4 The description herein does not impose any substantial limitations on the address translation system architecture constituted by the embodiments of this application. The system architecture described in the embodiments of this application is for the purpose of more clearly illustrating the technical solutions of this application and does not constitute a limitation on the technical solutions provided by this application. As the architecture evolves and new business scenarios emerge, the technical solutions provided by this application are also applicable to similar technical problems.
[0088] Figure 5 This is a schematic diagram of an address translation addressing apparatus according to an exemplary embodiment of the present disclosure, such as... Figure 5 The diagram shows the structural block diagram of the address translation and addressing device (SMMU / IOMMU) 500. This device 500, as the core carrier of address translation, is configured to dynamically map virtual addresses to physical addresses based on memory access requests. The address translation and addressing device 500 includes: Address translation control module 510 serves as the logic control center of this device. In one possible implementation, address translation control module 510 is configured to receive memory access requests initiated by the computing core, extract associated feature information (such as context identifiers, instruction pipeline status, etc.) from them, and then output a predicted page granularity identifier. Furthermore, this unit 510 is also responsible for monitoring the complete pipeline status of address translation and managing the retransmission control logic after prediction failure.
[0089] An index generation and shift module 520 is coupled to the output of the address translation control module 510. The index generation and shift module 520 is configured to dynamically generate a group index from the virtual address input from the upstream module of the SMMU based on the predicted page granularity identifier from the address translation control module 510. For example, this module 520 integrates multiplexing and shifting logic, enabling switching of index bit extraction segments between different page granularities (e.g., 4KB, 2MB, 1GB, etc.) to ensure consistency between the addressing path and the predicted granularity.
[0090] It should be noted that the SMMU is the recipient of the address translation request, responsible for converting the address translation request into a physical address request and then sending it downstream of the SMMU.
[0091] The storage array TLB Bank 530 is used to cache address translation entries (PTEs) of different granularities. TLB Bank 530 receives group indexes from module 520 and reads storage entries, i.e., address translation entries, from the corresponding physical groups for tag comparison. In this embodiment, TLB Bank 530 adopts a unified physical array structure, supporting the mixed storage of heterogeneous granularity translation entries within a single array, and using the true granularity identifier within the entry for post-hoc verification.
[0092] Page table lookup (PTW) module 540 is coupled between address translation control module 510 and external storage interface. When a memory access request misses in TLB Bank 530, PTW module 540 initiates a traversal operation of multi-level page tables in physical memory to obtain the latest address mapping relationship and fill it back into TLB Bank 530.
[0093] It is understandable that the above Figure 5 The architecture of the address translation and addressing device 500 shown is only one possible exemplary implementation. In actual engineering deployments, its hardware architecture can have greater complexity and flexibility to meet the needs of different computing platforms.
[0094] In one embodiment, the hierarchical structure of the TLB is not limited to a single-level array, but can employ a multi-level hierarchical design. For example, the address translation addressing device 500 may include: Level 1 micro TLB: Employs a fully associative architecture, with a small capacity but extremely fast access speed, used to handle near-end addressing requests that are highly sensitive to latency; Level 2 distributed TLB: Deployed in specific processing clusters or local bus nodes, used to balance addressing throughput and storage density; Level 3 global TLB: Serves as a globally shared translation cache center, possessing the largest capacity and coverage, responsible for handling mapping needs across modules or large address spaces.
[0095] Furthermore, the interconnection methods between modules, the cache replacement strategy (such as the PLRU algorithm or random replacement used for cache replacement), and the parallelism of PTW can all be adjusted according to the specific architecture design. It should be noted that... Figure 5 The structural block diagrams shown are only for assisting in understanding the core inventive points of this application and do not constitute an exclusive limitation on the physical implementation of the multi-granularity address translation addressing device provided in the embodiments of this application. Any adaptive improvements made based on the inventive concept of this application and for the corresponding architecture are within the protection scope of this application.
[0096] Figure 6This is a structural block diagram of an address translation control module according to an exemplary embodiment of the present disclosure. The address translation control module, labeled 510 in the diagram, serves as the scheduling and coordination core of the entire addressing pipeline, configured to manage the entire lifecycle from virtual address reception to physical address translation completion. The address translation control module 510 internally includes: Address translation queue 511, coupled to the upstream request interface, is configured to cache virtual address access requests (i.e., address translation requests) initiated by the upstream computing core. Address translation queue 511 temporarily stores the received raw requests and distributes them in an orderly manner to downstream address translation modules based on pipeline availability. It is understandable that the storage depth of this address translation queue 511 can be flexibly configured according to specific system design and performance goals to balance memory access throughput and hardware area overhead.
[0097] A page granularity prediction unit 512 is connected to the output of the address translation queue 511. This unit 512 is configured to predict the page granularity for each memory access request. In one embodiment of this application, the page granularity prediction method may include prediction based on the memory access status of the historical context, or the use of hardware prediction algorithms such as address pattern matching. It should be noted that this application does not limit the specific prediction algorithm; it can combine multiple software and hardware prediction methods. In actual engineering implementation, the design of the page granularity prediction unit 512 needs to strike a trade-off between prediction accuracy and hardware implementation complexity to ensure that while improving addressing efficiency, it does not introduce excessive hardware overhead or timing pressure.
[0098] The state management unit 513 is configured to manage the real-time translation state of each virtual address access in the pipeline. For example, the state machine maintained by the state management unit 513 includes, but is not limited to, the following: Idle state: Indicates that the current state entry is not occupied by any address translation request; HIT hit status: This means that the address translation has successfully obtained the correct physical mapping information from the cache array and is waiting to be output to the downstream. PTW wait state: This means that the access was identified as a miss, a hardware page table lookup request has been initiated, and it is in a suspended state waiting for the physical page table to be returned; Hit Under Miss (HUM) waiting state: This indicates that the request triggered the "hit under miss" mechanism under prediction guidance. In this state, due to the possibility of prediction bias, the hardware cannot yet determine whether the hit is absolutely correct. It must wait for the actual page table information to be returned and for granular consistency comparison to be performed before it can jump from this state to HIT or initiate a retransmission.
[0099] The retransmission and merging unit 514, as the deviation correction center of this device, is configured to handle the subsequent compensation scheme when page granularity prediction fails.
[0100] Retransmission logic: If the page granularity prediction results in an incorrect false HUM, the retransmission and merging unit 514 is responsible for correcting the page prediction granularity and driving the pipeline to re-initiate the address translation request. Merging logic: If page granularity prediction results in redundant PTW requests or redundant TLB cache entries, this unit 514 is responsible for performing PTW request merging and TLB cache cleanup to eliminate mirror redundancy and save system bandwidth.
[0101] The virtual address request interface 515 is configured to standardize and encapsulate various types of information required for virtual address translation. In one embodiment, the interface 515 encapsulates necessary parameters, including virtual page number (VPN), access permission information, read / write type information, and predicted page granularity information, into a predefined interface format. This unified encapsulation format facilitates identification and efficient communication between the various sub-modules within the address translation control module 510. It is understood that the bit definition and protocol format of this interface do not constitute a limitation of this application.
[0102] It is understandable that the above Figure 6 The division of the address translation control module 510 shown is merely an illustrative example of one logical function. In actual hardware circuit implementation, the functions of the above modules can be combined, split, or recombined. For example, the status management unit 513 and the retransmission and merging unit 514 can be physically integrated into the same timing control logic. The sub-modules and their connections described in this application are intended to clearly illustrate the inventive concept of this application and do not constitute an exclusive limitation on the specific physical implementation of the hardware.
[0103] Figure 7 This is a schematic diagram illustrating an exemplary implementation of an address translation addressing method shown in this application, such as... Figure 7 As shown, the method mainly includes the following steps: S701, the address translation queue initiates an address translation request. When an address translation request is received from upstream, the address translation queue temporarily stores and distributes the request, extracts the virtual address information to be translated, and officially starts the address translation pipeline.
[0104] Address translation requests initiated by the upstream computing core are used for virtual address access. They are first stored in the address translation queue. This queue will decide whether to release new access based on the current virtual address translation status being processed by the SMMU. If the SMMU hardware is currently at full load, the request will not be sent out for the time being.
[0105] S702, the page granularity prediction unit generates a predicted page granularity according to rules. The virtual address translation information flows to the page granularity prediction unit. This unit makes a prediction based on preset prediction rules (including but not limited to context memory access status, etc.) and outputs a predicted page granularity identifier. This step, by predicting the granularity in advance, provides a control basis for subsequent elimination of group index conflicts.
[0106] S703, the Virtual Address Request Interface generates a TLB query request. The Virtual Address Request Interface receives virtual address translation information from upstream and combines it with the predicted page granularity identifier generated in step S702. This interface module encapsulates the above information into a standardized interface format and generates a TLB query request to be sent to the TLB cache array.
[0107] S704, the TLB Bank generates a group index based on the predicted page granularity and searches the TLB cache. Upon receiving a query request, the TLB Bank does not use a fixed bit segment extraction method, but rather generates a group index guided by the predicted page granularity identifier.
[0108] Path hit (TLB hit): If a hit is determined through tag comparison and granularity consistency check in the corresponding index group, the physical address request is directly output to complete the translation path.
[0109] Path miss (tlb miss): If it is determined to be a miss, the process proceeds to step S705.
[0110] Step S705, PTW page table lookup. When a cache miss occurs, the Page Table Lookup (PTW) module initiates a hardware page table traversal to retrieve the actual page table information from memory. After obtaining the actual information, the device executes granular verification logic: Page table granularity prediction is correct: If the actual granularity matches the predicted granularity, then the address translation and conversion are completed and a physical address request is issued.
[0111] Page table granularity prediction error: If the actual granularity does not match the predicted granularity, it is determined that a prediction error has occurred.
[0112] Abnormal retransmission path (triggered retransmission): When step S705 determines that the page table granularity prediction is incorrect, the system generates a "triggered retransmission" signal. This signal is fed back to the page granularity prediction unit, prompting the pipeline to clean the current error state and guide the address translation request to restart the translation process from step S702 based on the corrected true granularity information.
[0113] It is understandable that the above Figure 7The address translation method flow and the logical relationships between the steps shown are merely one possible illustrative example of an embodiment of this application. In actual hardware logic implementation, pipeline scheduling, or firmware control, the above steps can be combined, split, reorganized, or their execution order adjusted according to specific system performance goals or hardware timing requirements.
[0114] For example, in some highly parallel pipeline designs, the page granularity prediction logic in step S702 can be executed synchronously and in parallel with the request dispatch logic in step S701; or, the interface encapsulation action in step S703 can be implicitly included in the logic implementation of step S702 or step S704, rather than appearing as an independent physical stage. Furthermore, the feedback retransmission path shown in the flowchart can also be optimized according to specific exception handling strategies.
[0115] It should be noted that the above flowchart and its corresponding textual description are intended to clearly demonstrate the inventive concept of "predictive guidance + dynamic addressing + anomaly self-healing" in this application, and do not constitute an exclusive limitation on the actual execution path of the multi-granularity address translation method provided in this application. Any equivalent substitutions or adaptive improvements to the process based on the inventive concept of this application and for a specific computing architecture should fall within the protection scope of this application.
[0116] Figure 8 This is a schematic diagram illustrating an exemplary implementation of a retransmission and error correction process after prediction failure, as shown in this application. Figure 8 As shown, this embodiment uses a predicted page granularity (first page granularity) of 2MB, but the actual page granularity is smaller, as an example for detailed explanation: S801, Initiate Address Translation Request 0 and generate prediction results. The Address Translation Control Module receives the memory access request (i.e., Address Translation Request 0), and the Page Granularity Prediction Unit initially predicts it to be a 2MB page granularity based on contextual features.
[0117] S802, performs index calculation based on prediction granularity. The addressing circuit (which may be coupled to the index generation and shifting module) is guided by the 2MB prediction result and uses the bit field after removing the 2MB offset from the virtual address (VA) as the index of the TLB.
[0118] Regarding the definition of the index bit field: Since the page offset corresponding to a 2MB page is the lower 21 bits (bit[20:0]), the index bit field is extracted starting from the 21st bit of the virtual page number (VPN), denoted as VPN[xx:21]. The bit width (xx-21+1) depends on the number of sets (SETs) N in the TLB array.
[0119] In one embodiment: if the number of TLB groups is N, the addressing circuit extracts the lower log2N bits of the VA field after excluding the 21-bit offset as the physical group index. In another embodiment: to further reduce the collision probability, the addressing circuit can also use a specific hash algorithm to calculate the final group index by operating on the address information of VPN[xx:21] or higher bits.
[0120] S803, Enqueuing and Range Determination of Subsequent Consecutive Address Translation Requests. During the processing of address translation request 0, the pipeline receives consecutive address translation requests 1-N. The system determines that address translation requests 0-N belong to the same 2MB alignment range of a virtual address.
[0121] S804, triggering missing event handling and suspension of spurious HUM state. Address translation request 0 experiences a TLB Miss and initiates a PTW request. At this time, address translation requests 1-N are determined to be in a Hit Under Miss (HUM) state and suspended. Due to potential prediction errors at this point, this HUM state is considered a "spurious HUM".
[0122] S805, Page Table Result Return and Granularity Consistency Verification. The PTW module returns the actual page table entry information. The retransmission control module identifies that the actual page granularity corresponding to the virtual address is not 2MB (e.g., it is actually 64KB, 16KB, or 4KB).
[0123] S806, Prediction Error Detection and Fine-grained Retransmission. Due to a page-level prediction error, the retransmission and merging unit immediately suspends the address translation request 0-N error-pending state. At this time, the system has determined that the batch of requests belongs to the current 2MB page table, but its entry granularity is less than 2MB.
[0124] Retransmission guidance logic: The system guides the batch of requests to correct the prediction granularity identifier and re-initiates the address translation query with a smaller granularity. Multi-granularity determination mechanism: If the system supports multiple intermediate granularities such as 64KB, 16KB, and 4KB, the page granularity prediction unit will further determine the specific granularity used for retransmission based on the deviation characteristics fed back by PTW. It can be understood that the finer the granularity determined by the prediction system, the less likely subsequent secondary retransmissions will be, thus achieving optimal memory access performance.
[0125] It is understandable that the above Figure 8The illustrated process uses a 2MB misprediction as 4KB as a typical example to illustrate the hardware logic of this application when dealing with false HUMs caused by "overestimation". In practical applications, the error correction process provided by this application is also applicable to other granularity combinations (such as mispredicting 1GB as 2MB, or vice versa). It should be noted that the retransmission path design ensures that even if the accuracy of the prediction module is limited by the algorithm complexity, the hardware system can still maintain the absolute correctness of the logic through the post-verification mechanism, thereby ensuring the robustness of the system while pursuing high-performance addressing.
[0126] Figure 9 This is a schematic diagram illustrating an exemplary implementation of a cross-group redundancy elimination and PTW merging process after prediction failure, as shown in this application. Figure 9 As shown, this embodiment is mainly used to solve the problem of bandwidth waste and cache pollution caused by index bit offset when a very large page (e.g., 1GB) is mispredicted as a smaller page (e.g., 2MB). The process specifically includes the following steps: S901, Initiate Address Translation Request 0 and generate prediction results. The Address Translation Control Module receives Address Translation Request 0. The Page Granularity Prediction Unit predicts it as a 2MB page granularity.
[0127] S902 performs index calculation and generates a Miss. The addressing circuit is guided by 2MB prediction and uses the VPN[xx:21] bit field to generate a group index. The retrieval is determined to be a TLB Miss.
[0128] S903, Enqueueing of subsequent consecutive address translation requests. During the PTW processing of address translation request 0, the pipeline receives consecutive address translation requests 1-N. Due to prediction granularity bias, these requests are identified as independent tasks pointing to different 2MB intervals and trigger addressing operations separately.
[0129] S904, multiple Miss requests converge in the PTW queue. Due to multiple Misses occurring within the TLB, address translation requests 0-N all enter the Page Table Lookup (PTW) request queue or the MSHR (Missing Status Register).
[0130] In one embodiment, step S904 can simultaneously perform pre-comparison based on maximum granularity: the retransmission and merging unit monitors the address range in the PTW queue in real time. Even if the actual granularity is not yet known, the hardware logic will still use the maximum page mask supported by the system (such as a 1GB mask) to perform range comparison on address translation requests 0-N. If they are all detected to be within the same naturally aligned 1GB range, the system will temporarily suspend the issuance of subsequent address translation requests (1-N), allowing only address translation request 0 to enter the memory access stage. This step achieves early interception of potentially redundant requests.
[0131] S905, page table result return, granularity verification, and final merging. The physical memory returns the actual page table entry information for address translation request 0, confirming that the actual page granularity is 1GB.
[0132] In one embodiment, step S905 can be implemented as a merge confirmation: the retransmission control module compares the actual 1GB granularity with the address translation requests 1-N pending in the PTW queue. Since the addresses of confirmed address translation requests 1-N are completely covered by the 1GB mapping returned by address translation request 0, the system formally performs the merge operation, directly canceling address translation requests 1-N and any subsequent possible duplicate PTW accesses.
[0133] Technical effect: This mechanism ensures that only a single effective memory access bandwidth loss occurs for the same large page address space, effectively avoiding the "PTW request storm" caused by misprediction.
[0134] S906, Redundancy Elimination and Storage Optimization. The system boots the correct 1GB entry to the correct physical group determined by VPN [xx:30].
[0135] In one embodiment, step S906 can simultaneously perform cross-group cleanup for merging confirmation: since address translation requests 1-N may have previously generated temporary suspended states or invalid entries in different groups of the TLB based on an incorrect 2MB index, the retransmission and merging unit simultaneously initiates a cleanup signal for the 1GB range, deleting all redundant copies from different groups within the array to ensure the uniqueness of the storage space.
[0136] It is understandable that the above Figure 9 The illustrated process aims to demonstrate the self-healing capability of this application in dealing with bandwidth and space waste caused by "underestimation". Through PTW merging logic, this application can effectively reduce invalid access to GPU memory in high-bandwidth scenarios such as loading large model parameters, thereby reducing system congestion.
[0137] It should be pointed out that the above Figure 9 The 1GB and 2MB granularity combinations described are merely examples. The cross-group redundancy elimination and merging logic described in this application is also applicable to any combination scenario where the "predicted granularity is smaller than the actual granularity," such as 4KB / 64KB / 2MB / 1GB. The hardware processing steps shown in this embodiment do not constitute an exclusive limitation on the merging algorithm or cleanup strategy of the technical solution of this application.
[0138] This application also provides a chip, a chip system (such as a SoC), a computer-readable storage medium, and a computer program product. Hardware carrier: The chip or chip system includes a processor and interface circuitry for supporting the processor in executing the multi-granularity address translation method described in the above embodiments. Software carrier: The computer-readable storage medium (such as ROM, RAM, disk, or optical disk) or program product stores computer instructions that, when executed, cause a computer or hardware processor to perform the addressing steps described above, such as page granularity prediction, dynamic index shifting, and retransmission error correction.
[0139] Understandably, the above division of functional units or modules is only a schematic representation of logical functions. In actual hardware implementation, the above functions can be allocated to different physical modules to complete according to performance targets, or multiple units can be integrated into the same physical component (such as integrated into the bus controller or memory management unit).
[0140] The system architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions and do not constitute limitations. With the evolution of computing architectures, the technical solutions provided in this application are also applicable to similar address translation or cache conflict problems. The units described as separate components can be physically separated, located in the same physical location, or even distributed across multiple network nodes.
[0141] Finally, it should be noted that the above descriptions are merely specific embodiments of this application and not limitations on its scope of protection. Any equivalent modifications or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. The scope of protection of this application shall be determined by the scope of the claims.
[0142] The technical solution provided in this application introduces page size prediction and dynamic index shifting mechanisms into the address translation process, which has significant technical effects in improving hardware resource utilization, reducing addressing power consumption, and optimizing memory access latency.
[0143] This application can improve the effective utilization of storage arrays. When dealing with very large pages (such as 1GB or 512MB) with strict address alignment characteristics, the fixed index fetch bit field scheme is prone to causing severe index collisions (Set Aliasing) in specific physical groups (such as Set0). This application uses dynamic index shifting to move the index bits of very large pages to the effective virtual page number range, so that large page entries with different base addresses can be evenly distributed in the physical array. This mechanism can significantly improve the actual effective capacity of the hardware array.
[0144] This application reduces the dynamic power consumption of address translation. Compared to methods such as using multiple physical arrays for parallel lookup (Split TLB) or complex parallel mask TLB to support multi-granularity lookup, this application guides the hardware to perform single-path addressing through a prediction mechanism. In scenarios supporting mixed access of various heterogeneous granularities (such as 4KB, 64KB, 2MB, 1GB, etc.), this application effectively reduces unnecessary memory bit line flips, thereby reducing the dynamic power consumption of the addressing pipeline.
[0145] This application optimizes system memory access bandwidth and latency. The PTW merging logic implemented through retransmission and merging units effectively identifies and merges redundant missing requests targeting the same physical large page range. In high-bandwidth scenarios such as loading large model weights, this mechanism reduces repeatedly initiated memory lookup tasks, lowers bus congestion, and optimizes the system's average memory access latency while ensuring addressing accuracy.
[0146] This application enhances the scalability and robustness of the system. The addressing framework described in this application does not rely on specific physical array partitioning and has good page-level scalability. Simultaneously, the closed-loop retransmission control logic ensures that even when prediction errors occur, the system can quickly converge to the correct path through a hardware self-healing mechanism, balancing high-performance addressing with rigorous logic execution.
[0147] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0148] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0149] like Figure 10As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded into random access memory (RAM) 1003 from storage unit 1008. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.
[0150] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0151] The computing unit 1001 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as address translation addressing methods. For example, in some embodiments, the address translation addressing method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the address translation addressing method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to perform an address translation addressing method by any other suitable means (e.g., by means of firmware).
[0152] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0153] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0154] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0155] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0156] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0157] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0158] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0159] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An address translation addressing method, wherein, include: Obtain the address translation request and predict the first page granularity identifier corresponding to the address translation request; The storage array is searched according to the first page granularity identifier to obtain candidate address translation entries. The address translation entries include address labels, physical addresses, and a second page granularity identifier used to represent the actual page size. A consistency check is performed on the first page granularity identifier and the second page granularity identifier in the candidate address translation entry. In response to the check result indicating that the first page granularity identifier and the second page granularity identifier are consistent, the candidate address translation entry is determined to be the target address translation entry obtained by addressing.
2. The method according to claim 1, wherein, Also includes: In response to the verification result indicating that the first page granularity identifier and the second page granularity identifier are inconsistent, the current error waiting is terminated, prediction deviation correction and redundancy control are performed, and a retransmission request is initiated. The retransmission request is used to reprocess the address translation request.
3. The method according to claim 2, wherein, The execution of prediction deviation correction and redundancy control, and the initiation of retransmission requests, includes: Merge redundant page table lookup requests in the page table lookup queue and clear the address translation cache; The first page granularity identifier is corrected to obtain the third page granularity identifier; A retransmission request is initiated based on the third page granularity identifier to update the candidate address translation entry.
4. The method according to any one of claims 1-3, wherein, The step of retrieving candidate address translation entries from the storage array based on the first page granularity identifier includes: Dynamically extract different bit segments from the virtual address as group indexes based on the first page granularity identifier; The storage array is retrieved based on the group index to obtain candidate address translation entries.
5. The method according to claim 4, wherein, The step of dynamically extracting different bit segments from the virtual address as a group index based on the first page granularity identifier includes: The page offset is determined based on the first page granularity identifier; The group index is the bit segment obtained by removing the page offset from the virtual address.
6. The method according to claim 4, wherein, The step of dynamically extracting different bit segments from the virtual address as a group index based on the first page granularity identifier includes: In response to the first page granularity identifier indicating a super-large page, the high-order bits of the virtual address are extracted as the group index.
7. The method according to claim 4, wherein, The step of retrieving the storage array based on the group index to obtain candidate address translation entries includes: The storage array is retrieved according to the group index. In response to the group index not hitting the address translation entry, a multi-level page table traversal operation is initiated to update the address mapping relationship of the storage array. The updated storage array is retrieved again based on the group index to obtain the candidate address translation entries.
8. The method according to any one of claims 1-3, wherein, The prediction of the first page granularity identifier corresponding to the address translation request includes: Multi-dimensional feature extraction is performed on the address translation request to obtain the associated feature information of the address translation request. The associated feature information includes at least one of instruction sequence features, register usage features, address space distribution features, or memory access status. The associated feature information is predicted according to the configured prediction algorithm to obtain the first page granularity identifier corresponding to the address translation request.
9. An address translation addressing device, wherein, include: The address translation control module is used to obtain address translation requests and predict the first page granularity identifier corresponding to the address translation request; The index generation and shifting module is used to dynamically extract different bit segments in the virtual address as group indexes based on the first page granularity identifier. The address translation control module is further configured to retrieve the storage array according to the group index to obtain candidate address translation entries, and to perform consistency verification on the first page granularity identifier and the second page granularity identifier in the candidate address translation entries; In response to the verification result indicating that the first page granularity identifier and the second page granularity identifier are consistent, the candidate address translation entry is determined to be the target address translation entry obtained by addressing; or, in response to the verification result indicating that the first page granularity identifier and the second page granularity identifier are not consistent, the current error waiting is terminated, prediction deviation correction and redundancy control are performed, and a retransmission request is initiated.
10. The apparatus according to claim 9, wherein, The address translation and addressing device further includes: The page table lookup module is used to initiate a multi-level page table traversal operation in response to a missing address translation entry in the group index, so as to update the address mapping relationship of the storage array.
11. The apparatus according to claim 9, wherein, The address translation control module includes: An address translation queue is used to cache address translation requests initiated by the computing core; The page granularity prediction unit is used to perform multi-dimensional feature extraction on the address translation request, obtain the associated feature information of the address translation request, predict the associated feature information according to the configured prediction algorithm, and obtain the first page granularity identifier corresponding to the address translation request. The status management unit is used to manage the real-time translation status of each address translation request in the pipeline. The translation status includes idle status, hit status, page table lookup waiting status, and hit status after miss. The retransmission and merging unit is used to respond to the verification result indicating that the first page granularity identifier and the second page granularity identifier are inconsistent, merge redundant page table lookup requests in the page table lookup queue, and clear the address translation cache; correct the first page granularity identifier and obtain the third page granularity identifier; and initiate a retransmission request based on the third page granularity identifier to update the candidate address translation entry. The virtual address request interface is used to standardize and encapsulate the virtual address translation information corresponding to the target address translation entry.
12. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.
14. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-8.