Embedded Information System and Method for Memory Management

By adopting load control circuits and cache control circuits in embedded information systems and using variable size loading units and metadata management, the problems of latency and bandwidth limitations in the memory hierarchy system are solved, and more efficient memory management and performance optimization are achieved.

CN113111020BActive Publication Date: 2025-06-24NXP USA INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110015704.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-01-10
Filing Date
2021-01-06
Publication Date
2025-06-24
Estimated Expiration
2041-01-06

AI Technical Summary

Technical Problem

In embedded information systems, memory hierarchical systems face latency and bandwidth limitations, especially when large random access memory (RAM) are used, it is difficult to effectively manage memory to optimize performance.

Method used

An embedded information system is adopted to organize instructions and constant data using loading control circuits and cache control circuits using variable-sized loading units (LUs) and specify the properties of LUs in the metadata to optimize the use of internal memory and realize cache management.

Benefits of technology

Through this approach, it is possible to effectively reduce latency and bandwidth limitations, optimize memory contents, and improve system performance, especially when the bandwidth and latency of external NVMs are limited.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113111020B_ABST
    Figure CN113111020B_ABST
Patent Text Reader

Abstract

An embedded information system includes: a load control circuit coupled to an external memory, the external memory containing instructions and constant data associated with application code of a software application; at least one processor that executes at least one application code; an internal memory, a first part of which is a main system memory and a second part of which is a cache for storing instructions and constant data for executing at least one application code from the external memory. The load control circuit loads LUs associated with at least one application code from the external memory into the internal memory in the granularity of a single LU. The cache control circuit manages the second part based on metadata corresponding to the LUs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The field of the present invention relates to an embedded information system and a method for memory management therein. The field of the present invention is applicable to, but not limited to, mechanisms for memory management in a memory-constrained environment for software execution in an embedded information system such as an in-vehicle system. Background Art

[0002] Computer systems generally benefit from a hierarchical memory design. For example, in such a design, (at least partial) copies of memory contents can be stored (i.e., cached) at different levels within the memory hierarchy. Generally, the hardware supporting the different memory levels has different capacities, costs, and access times. Generally speaking, faster and smaller memory circuits are typically closer to the processor core or other processing elements within the system and act as caches. Slower memories may be larger but are also relatively slower compared to those acting as caches.

[0003] It is well known that some levels of this memory hierarchy can be located on the same semiconductor device as the processor core or other master device, while other levels can be located on another semiconductor device. For example, when implementing a pure logic device or a device containing non-volatile memory requires different technologies, the corresponding memory hierarchy also allows leveraging the technology differences between semiconductor devices. A semiconductor device implementing both computing logic (e.g., a processor element) and non-volatile memory (NVM) must bear the cost burden of both technologies. In contrast, implementing the two functions in two separate semiconductor devices can allow for a more cost-optimized implementation.

[0004] In such a memory hierarchy configuration, the main memory within the computing logic semiconductor device can act as a cache to load instructions and constant data from another semiconductor device implementing / providing external NVM. For performance reasons, these instructions can then be executed by reading the cache copy. For cost reasons, the main memory may be smaller than the size of the embedded NVM, which is also beneficial in cases where the NVM stores more than one copy of an application; for example, supporting the over-the-air download of a new version of a software application while executing the 'actual' version of the application. Further reducing the available main memory to, for example, significantly less than the size of the software application being executed can provide additional cost benefits, which can make the appropriate size of such a memory subsystem an ideal design goal.

[0005] However, compared with traditional processor caches (commonly referred to as 'Level 1' and 'Level 2' caches), using the main memory as an instruction cache in such a memory hierarchy system has different characteristics, and in addition to those traditional caches, it is also possible to implement using the main memory as an instruction cache. The first difference of such an instruction cache is its relative size. Although traditional caches are generally smaller or much smaller than the main memory, the size of the corresponding cache is typically about 75%-25% of the application image size. When the actual used part of the application is small, this size can be further reduced.

[0006] Currently, typical users of embedded information systems used in vehicles are requesting application images up to 16 megabytes in size, requiring an instruction cache in the range of 4 to 12 megabytes. These values may increase in the next few years. It is worth noting that the size of such an instruction cache exceeds the normal memory requirements for data storage of an application that may have the same size. The second equally important difference is the latency and bandwidth characteristics of these external NVMs, which not only require a relatively large overhead for a single transfer, but also provide a very limited bandwidth for loading instructions.

[0007] Considering these limitations, it is beneficial to appropriately limit the amount of data to be loaded (i.e., the size of the loading unit (LU)) to ensure that only the required information is loaded; otherwise, it may have a negative impact on the bandwidth. Since the amount of instructions used by the software may vary greatly, a variable-sized LU must be used to avoid loading unnecessary data. On the other hand, these LUs should be large enough to avoid any significant impact of transaction overhead. Both of these parameters indicate that the size of the LU is larger than the cache line of a traditional processor Level 1 or Level 2 cache, but also small enough to require such an instruction cache to support hundreds, thousands, or even one or two orders of magnitude larger numbers of LUs. Additionally, in the case of cache misses, the typical on-demand requests used by traditional Level 1 or Level 2 caches may cause excessive access latency; this means that other cache management mechanisms to avoid on-demand loading should also be studied.

[0008] In the case where the cache size is relatively large compared to the application image, a potential solution is to preferentially store more valuable instructions, which results in only rarely needed instructions being loaded. This not only reduces the amount of loading operations required (beneficial to the bandwidth), but also limits the possibility of cache misses when rarely needed instructions can be properly identified.

[0009] Therefore, a memory hierarchy system is needed in which the limitations of latency and bandwidth can be reduced or alleviated, especially when using large random access memories (RAMs) for operations. SUMMARY OF THE INVENTION

[0010] The present invention provides an embedded information system and a method for memory management therein.

[0011] According to one aspect of the present invention, there is provided an embedded information system (200), comprising:

[0012] Loading control circuits (168, 230), which can be coupled to an external memory (170, 216), the external memory (170, 216) containing instructions and constant data associated with at least one application code (228) of a software application,

[0013] At least one processor (250......258), the at least one processor (250......258) being coupled to at least one interconnect (264) and configured to execute the at least one application code (228);

[0014] An internal memory (266), the internal memory (266) being coupled to the at least one interconnect (264) and configured to store data for the software application as a main system memory in a first part of the internal memory (266), and being configured as a cache to store the instructions and the constant data for executing the at least one application code (228) from the external memory (170, 216) in a second part of the internal memory (266);

[0015] Cache control circuits (150, 260, 350), the cache control circuits (150, 260, 350) being coupled to the at least one interconnect (264) and additionally coupled to or including the loading control circuits (168, 230), wherein the instructions and the constant data are organized by a load unit LU of variable size, and wherein at least one attribute of the LU is specified within metadata (128, 220, 328);

[0016] wherein the loading control circuits (168, 230) are configured to load the LU associated with the at least one application code (228) from the external memory (170, 216) into the internal memory (128, 266, 328) in the granularity of a single LU; and

[0017] wherein the cache control circuits (150, 260, 350) manage the second part of the internal memory (266) configured as a cache based on the metadata corresponding to the LU by being configured to perform the following operations:

[0018] Observing (645) at least a portion of the execution of the at least one application code (228) by detecting at least one of the following: the LU being executed, a change from one LU to another within the internal memory (266);

[0019] Loading metadata information corresponding to the LU instance from the external memory (170, 216) or the internal memory (266),

[0020] Designating the next LU to be loaded into the second portion of the internal memory (266) by the load control circuit (168, 230); and

[0021] When there is not enough space to load the next LU, designating the next LU to be evicted from the second portion of the internal memory (266).

[0022] According to one or more embodiments, a first request to load the next LU into the internal memory (128, 266, 328) and a second request to evict a LU from the internal memory (128, 266, 328) are different operations that occur independently by the cache control circuit (150, 260, 350) based on at least one of the following: at different times, associated with different LU instances.

[0023] According to one or more embodiments, the first request to load the first LU instance corresponds to a first weight for selecting the first LU, and the second request to evict the second LU instance corresponds to a second weight for selecting the second LU.

[0024] According to one or more embodiments, the first weight and the second weight are independently calculated using different weight calculation formulas that use metadata information constituted by at least one of the first weight and the second weight.

[0025] According to one or more embodiments, the cache control circuit (150, 260, 350) includes a first dedicated hardware container LSH and a second dedicated hardware container SEC for storing metadata associated with a single LU instance, where: the metadata in the LSH is used to calculate and manage the first weight for the first request to load the next LU from the external memory (170, 216) into the internal memory (266), and the metadata in the SEC is used to calculate and manage the second weight for the second request to evict a LU from the internal memory (266).

[0026] According to one or more embodiments, the LU metadata within the first hardware container LSH and the second hardware container SEC includes the same general raw data for different LU instances, and the general raw data is used in different ways when calculating the first weight and the second weight of the same LU instance.

[0027] According to one or more embodiments, the calculation of at least one of the first weight and the second weight includes: the sum of intermediate weight results generated from selected components of the weight, where at least one of these components is derived from the raw data provided by the LU metadata after applying a sign and multiplying the selected component by a relevant factor.

[0028] According to one or more embodiments, the cache control circuit (150, 260, 350) is configured to record at least one of the following using a recorded observation of the execution of at least a portion of the application code (288): a previously used LU, a next LU, or a change from the previously used LU to the next LU.

[0029] According to one or more embodiments, the metadata is augmented with the recorded observation of the execution of at least a portion of the application code (228) and additional internal information (164) that can be used in the cache control circuit (150, 260, 350), where the additional internal information (164) includes at least another component derived from or augmented with at least one of the following: an application state reported to the cache control circuit (150, 260, 350), an internal state of the cache control circuit (150, 260, 350), a state of the LU associated with the first weight or the second weight (228).

[0030] According to one or more embodiments, the cache control circuit (150, 260, 350) includes a third dedicated hardware container HBE for metadata associated with a previously executed LU; and a fourth dedicated hardware container ALE for metadata associated with a currently executed LU, where a determined change between LU instances moves the metadata of the currently executed LU from the ALE to the HBE.

[0031] According to one or more embodiments, the cache control circuit (150, 260, 350) is configured to read metadata from memory into the LSH, SEC, and ALE dedicated hardware containers, and the cache control circuit (150, 260, 350) is configured to write at least the metadata associated with the recorded observation information from the HBE dedicated hardware container to memory.

[0032] According to one or more embodiments, at least one attribute of the metadata for the LU instance includes the start address of the LU instance in the internal memory and at least one of the following: the size of the LU instance, the end address of the LU instance.

[0033] According to a second aspect of the present invention, there is provided a method (600) for memory management in an embedded information system (200), the embedded information system (200) including at least one processor (250......258), at least one interconnect (264), an internal memory (266), and a cache control circuit (150, 260, 350) coupled to or including a load control circuit (168, 230), wherein the method includes:

[0034] Connecting the load control circuit (168, 230) to an external memory (170, 216) containing instructions and constant data associated with at least one application code (228);

[0035] Organizing the instructions and the constant data of the at least one application code (228) by a load unit LU of variable size, wherein at least one attribute of the LU is specified within metadata (128, 220, 328);

[0036] Configuring the internal memory (266) to store data for the software application as a main system memory in a first part of the internal memory (266), and being configured as a cache to store the instructions and the constant data for executing the at least one application code (228) loaded from the external memory (170, 216);

[0037] Loading (640) at least a portion of the at least one application code (228) from the external memory (170, 216) into the internal memory (266) in a single LU granularity;

[0038] Executing (650) at least a portion of the at least one application code (228) associated with at least one LU instance located within the internal memory (266), and

[0039] Managing a second part of the internal memory (266) configured as a cache based on the metadata (128, 220, 328) corresponding to the LU by:

[0040] Observing (645) at least a portion of the execution of the at least one application code (228) by detecting at least one of the following: the LU being executed, a change from one LU to another within the internal memory (266);

[0041] Loading (620) metadata information corresponding to the LU instance from the external memory (170, 216) or the internal memory (266);

[0042] Designating the next LU to be loaded into the second portion of the internal memory (266); and

[0043] When there is not enough space to load the next LU, designating the next LU to be evicted from the second portion of the internal memory (266).

[0044] According to one or more embodiments, the method further includes processing a first request to load the next LU into the internal memory (128, 266, 328), and processing a second request to evict an LU from the internal memory (128, 266, 328), wherein the processing of the first request and the second request occurs as different operations independently of at least one of the following: at different times, associated with different LU instances.

[0045] According to one or more embodiments, the method further includes writing (690, 212) the result of the execution observation into metadata.

[0046] These and other aspects of the invention will be apparent from the embodiments described hereinafter and will be elucidated with reference to the embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Additional details, aspects, and embodiments of the invention will be described by way of example only with reference to the drawings. In the drawings, like reference numerals are used to identify like or functionally similar elements. The elements in the figures are shown for simplicity and clarity and are not necessarily drawn to scale.

[0048] Figure 1 An overview example diagram of a method of using a memory as a cache to load a set of instructions and constant data from an external memory using a variable-sized load unit according to some example embodiments of the invention is shown.

[0049] Figure 2 An example diagram of an instruction cache is shown according to some example embodiments of the invention, the instruction cache using an internal main memory as a storage device for software instructions to be executed within an information processing system.

[0050] Figure 3 Illustrates an example method of using a memory as a cache to load a set of instructions and constant data from an external memory using a variable-sized load unit, where the external memory uses beneficial cache management supported by execution observation.

[0051] Figure 4 Illustrates the benefits of using different criteria for cache management according to some example embodiments of the present invention.

[0052] Figure 5 Illustrates, according to some example embodiments of the present invention, Figure 1 , Figure 2 or Figure 3 An example timing diagram of elements used in the system schema of, and the use of some of these elements during a load unit switch.

[0053] Figure 6 Illustrates an example flowchart of an observation operation performed by a hardware controller circuit according to some example embodiments of the present invention. Detailed Description

[0054] To address the limitations, drawbacks, and restrictions of classical methods for using a memory as a cache to preload and / or load a set of instructions and constant data from an external memory, a system with a large memory is described. The large memory implements a cache structure for instructions and constant data in a memory (e.g., random access memory (RAM)), notably using a variable-sized load unit (which is different from known processor cache arrangements) and supporting separate load requests and cache eviction operations for cache operations. In an example of the present invention, a cache controller circuit (preferably, in hardware) is provided, which is configured to optimize the content of the large memory by using a combination of separate load requests and cache memory eviction operations, and is configured to observe the software being executed, such that the operations can run independently, at different times, but can be managed from the same data set. In this way, the cache can be quickly preloaded with data from the external memory, and the cache controller circuit can identify the most valuable data elements to be saved in the memory. Thus, in this way, the control of loading corresponding information (instructions and constant data) can utilize the information that already exists.

[0055] In some examples, the cache controller circuit can also be configured to record relevant information of the observed software being executed determined when one requested LU switches to another LU. According to at least one example embodiment, the information collected can also be used by the cache management to control its operations.

[0056] The data collected and other existing information (also known as metadata since it is data that describes data) are additionally referred to as raw data. The existing metadata may be generated by another process that is ultimately external and timely separated (such as a tool flow). In some examples, this metadata can be combined with the observed collected data. The corresponding information can be delivered in the form of a data structure that provides relevant information for each LU or set of LU instances, where the term LU "instance" refers to a single loading unit. In some examples, updating the content of this data structure with the collected information permits the permanent recording of this data.

[0057] In some examples, a first element that processes a load request is separated from a second element that processes a cache eviction operation for cache management. In some examples, the first element performs a first calculation of a first weight that independently controls the load request operation using the raw data, and the second element performs a second calculation of a second weight that independently controls the cache eviction operation using the raw data. In some examples, the first calculation of the first weight (load request) and the second calculation of the second weight (cache eviction) can be based on the same raw data provided by the data structure, thereby using separate, specific programmable formulas that can specify different selections and interpretations from this common raw data. For this purpose, the corresponding metadata can be loaded from the data structure into internal processing elements in order to fulfill its use.

[0058] Examples of the present invention present an embedded information system that includes an internal memory configured to store data for a software application as a main system memory in a first portion of the internal memory and configured as a cache to store instructions and constant data for executing at least one application code from an external memory in a second portion of the internal memory. At least one processor and the internal memory are both coupled to an interconnect. In some examples, the interconnect may additionally be coupled to a cache control circuit and a load control unit (among other peripheral devices). The load control unit may be coupled to an external memory containing instructions and constant data associated with at least one application code. In some alternative examples, the cache control circuit may include the load control unit. In some examples, the cache control circuit or the load control circuit may also be directly coupled to the internal memory. The load control circuit is configured to load LUs associated with at least one application code from the external memory into the internal memory in the granularity of a single LU. Instructions and constant data of at least one application are organized by load units (LUs) of different sizes, and at least one attribute of the LU is specified within metadata. Examples of LU attributes stored within the metadata are: the start address and / or size or end address of this LU (substantially defining the address range to be loaded). Other beneficial or possibly required attributes include a checksum (e.g., created by cyclic redundancy check) that can be used to verify correct loading, or other information that can be used to define the loading sequence of multiple LUs. The load control unit is configured to load LUs associated with at least one application code from the external memory into the internal memory in the granularity of a single LU. The cache control circuit manages the second portion of the internal memory configured as a cache based on the metadata corresponding to the LUs by being configured to perform the following operations: observing the execution of at least a portion of at least one application code by detecting at least one of the following: the LU being executed, a change from one LU to another within the internal memory; loading metadata information corresponding to the LU instance from the external memory or the internal memory; designating the next LU to be loaded by the load control unit, and when there is not enough space to load the next LU, designating the next LU to be evicted from the internal memory.

[0059] In this manner, a cache control circuit configured to perform the following operations advantageously improves (and in some instances optimizes) the actual cache content to handle the bandwidth and latency limitations of an external memory system: designating the next LU to be loaded by a load control unit, designating the next LU to be evicted from the main system memory when there is not enough space to load the next LU, and observing the execution of at least one application through the system at the granularity of a single LU, advantageously improving (and in some cases optimizing) the actual cache content to handle the bandwidth and latency limitations of an external memory system. Such observation is only possible through a hardware-based scheme that also provides (i) the concurrency of the required operations independent of the application execution, and (ii) the performance required to reduce / minimize any impact on the application execution. In an example of the present invention, it is further contemplated that the hardware uses additional implementations (e.g., through high-priority "urgent" load operations) to react quickly to missing LUs, which reduces / minimizes any impact of the corresponding stall situations on the application execution. Such observation also helps to record the original data associated with such events simultaneously to support problem debugging or subsequent runtime optimization. In some examples, the observation can be an access during the execution of the application code.

[0060] Now referring to Figure 1 , an overview of a system 100 that uses memory as a cache to load or preload a set of instructions and constant data from an external memory is shown in accordance with some example embodiments of the present invention. The system 100 identifies a tool-based flow 120 for identifying the uses of LUs and the beneficial uses achieved by such an instruction cache implementation through the tool-based flow 120.

[0061] A classical software development environment 110 generates corresponding application code, which is then analyzed 122 by the tool-based flow 120. For example, only the tool-based flow 120 interacts with the software development environment 110 to control or update the specification of a load unit (LU).

[0062] The loading unit executed by the tool-based flow 120 generates metadata 126 for building application code within a set of LU instances by using the analysis results of the application code. This LU information (metadata 126) is then stored in a database called the LU table 128. The LU table 128 provides metadata about LUs of variable size (i.e., data about data) in the form of a replicated data structure, where the content of a single data structure specifies a single LU instance and holds other relevant information. The initial content of the LU table 128 can be generated from static information, for example, by analyzing the application code, especially the set of included functions and their callers / callees and other relationships. Grouping one or more functions into LUs according to these relationships aims to generate LU instances of acceptable size. Organizing the software application into LU instances defines the granularity of the loading process, which requires storing at least the start address and size (or end address) of the LU in the metadata of the LU. In some examples, the correlation between LU instances can be used to additionally specify the loading sequence. According to some examples of the present invention, additional attributes can be specified for each LU. According to some examples of the present invention, data collected by the cache controller circuit or acquired in other ways such as by profiling functions can be used to further refine the content of the LU table 128.

[0063] Since the application software of the application executes at the granularity of functions, the output of the tool-based flow 120 is beneficially a specification of LU instances that combine one (or more) software functions, and the software functions themselves are assumed to be atomic elements by this LU generation step / operation. Thereafter, any LU can be conceived as an address range that combines one or more software functions.

[0064] The hardware (HW) controller device 150 that manages such instruction caches in the context of executing application software implements an advanced use of this LU metadata. For ease of explanation and not to confuse the concepts described herein, other elements that are obviously also used for this purpose are not shown in the figures, especially the internal memory for storing the loading unit or the processing element running the application software. Additionally, it is worth noting that generating the LU table 128 information by the LU generation tool 124 can be separated from using this information by the HW controller device 150 in terms of time and space.

[0065] The HW controller device 150 employs a control data management unit 154, which reads the metadata 152 from the LU table 128 and provides this read metadata 162 to a cache management circuit 166 arranged to control instruction cache management. Generally, the cache management circuit 166 will use additional information, for example, data 164 dynamically generated for this purpose. Examples of such dynamic data include:

[0066] i) The fill state of the main memory dedicated to the instruction cache,

[0067] ii) Information about in - progress or unfinished load operations,

[0068] iii) Information about the application program state that can be provided via programming of the status register by software,

[0069] iv) Information about the state of the processing system, or

[0070] v) Status information provided by elements inside the HW controller device 150 or by an external memory 170, etc.

[0071] The main function covered by the cache management circuit 166 is the loading of the LU. For this purpose, the HW controller device 150 will typically include or employ a load control unit 168 configured to perform this function. For this purpose, the load control unit 168 is connected to the external NVM memory 170.

[0072] According to an example of the present invention, such processing steps / operations can advantageously be observed by an observation circuit 160, which is configured to identify events or attributes of interest for using this instruction cache in the context of application program execution. Since the instruction cache uses the LU for its management, these events are typically related to the LU. It is conceivable that some examples of the information collected include, for example, the number of times the LU is used, optionally further qualified with the previous cache state of the LU, the execution time of the LU, the time when the LU changes, or the corresponding latency. Since one purpose of the observation circuit 160 is to identify potential optimizations related to these LUs, such information related to the LUs can be beneficially collected. For this purpose, the observation circuit 160 can advantageously be coupled to the control data management unit 154 so as to be able to access the corresponding information. In some examples of the present invention, the information collected by the observation circuit 160 can then be reported back 156; that is, reported to the application software via a status register (not shown in the figure) or through a specific reporting or observation channel dedicated to this purpose. In some examples, it is conceivable that existing debugging facilities (e.g., trace generation facilities) can also be used to implement this reporting - back 156 mechanism. However, in this example, it should be noted that this reporting - back 156 mechanism will prevent these facilities and functions from being used for their original purposes during a temporary period. The main candidates for the use of this information are tool - based flows 120, which can use this information for further improvements.

[0073] In some examples, it can be envisioned that the HW controller device 150 used by the instruction cache as described above needs to consider cache management factors that are very different from those previously considered in known instruction cache installations. For example, known L1 / L2 caches typically have 64 - 256 cache sets, which is an unacceptably small number for managing the amount of LU instances for an entire application image. Also, the data managed by a single cache set in known instruction cache installations has a fixed length and is very small, typically within 4 or 8 times the data path consisting of 16 to 64 words. Due to the bandwidth limitations of the previously described target external non - volatile memory, these two properties, namely the small size of the cache data sets and the limitation of only supporting a single fixed cache line size, are unacceptable for the example embodiments. Additionally, the absolute number of LUs (from hundreds of LUs to thousands of LUs, or one or two orders of magnitude more) that must be supported to implement an instruction cache that can hold 25% - 75% of an entire application image is too large to fit the LU table into the HW controller device 150. Since the management of each LU requires multiple pieces of information (at least the LU address range typically encoded by a start address and size or end address, which can be combined with other information such as a checksum and load sequence information stored in the metadata), this is a non - negligible amount of data. Moreover, this amount can vary greatly, which invalidates any beneficial use of the embedded memory for this purpose. Simply put, any selected size that is valid for one application will be wrong for another; because there can be too many variations when building a piece of software. And, the metadata associated with a single LU can be relatively large, which makes it necessary to minimize the hardware elements required to store this information. Therefore, to address this problem and the previous ones, examples of the present invention advantageously use the main memory to store LU table information. Multiple circuits or components within the HW controller device 150 may require metadata; for example, the cache management circuit 166 for managing cache operations requires metadata, and the observation circuit 160 also requires metadata. Any circuit or component that requires metadata must be able to access the LU currently being managed by that circuit or component, and the LUs are not necessarily the same for all processing elements, circuits, or components.

[0074] In some examples according to the present invention, using LUs of variable sizes, such as allowing the LU size to range from 1k to 64k (or in some instances 1k to 256k), may cause confusion. Supporting LU sizes smaller than 1k is not beneficial because the bandwidth overhead for loading smaller LU instances will be very large. For example, it is possible that one (large) LU has to replace multiple (smaller) LUs. On the other hand, when a large LU is replaced earlier leaving enough space, a small LU can be loaded without replacing another LU. Therefore, to address this potential problem of using LUs of variable sizes, examples of the present invention propose separating the corresponding operations such that the load request is different from the cache eviction operation used to provide available space for storing LUs during the load operation.

[0075] Compared with the management of traditional level-1 or level-2 processor caches, the above series of confusions make the management of such an "instruction cache" a unique problem. On the other hand, the main characteristic of the described instruction cache is a system-wide cache within an embedded information system running a known software application, rather than a cache specific to a processor or a processing cluster (such as an L1 / L2 cache). Additionally, there are corresponding differences in its behavior and management with respect to its size relative to the main memory and the external memory.

[0076] It is worth noting that in some aspects of these examples, any similarity with the management of caches within a disk controller is different because the disk controller completely lacks the execution context for the loaded information (data and instructions). Additionally, the performance characteristics of these caches typically correspond to the physical properties of the read process (the bandwidth provided by the read head and the mechanical time for moving the read head), and thus have nothing to do with the execution of an application (or a part thereof). These characteristics are very different from the design goals (supporting the execution of known applications) and bandwidth requirements of the described cache, which is related to the access bandwidth of one or more embedded processors on one side and the very limited access bandwidth provided by the target type of external NVM on the embedded information system on the other side. This difference in methods and implementations also applies to solid-state disks, which are also NVM-based but operate under completely different types of system bandwidths and latencies.

[0077] Now refer to Figure 2, shows an example of an embedded information processing system 200 with an instruction cache according to some example embodiments of the present invention, where the instruction cache uses an internal main memory as a storage device for software instructions executed by the system. The information processing system 200 implements related processing elements within a first integrated circuit 285, while the software instructions to be executed by the information processing system 200 and the associated constant data 228 are stored in an external non-volatile memory 216 within a second integrated circuit 222.

[0078] Generally, an application image always includes the instructions that make up the software part of this application, as well as the associated constant data, which is either used as constant read-only data or for initializing variable values (in both cases, the corresponding values are defined during the software development process). In the latter case (sometimes referred to as variable initialization), the constant read-only memory data is copied to the volatile memory that stores the application data when the application is started. Compared with these read-only initialization values, the constant read-only data can be directly accessed, equivalent to reading software instructions. Both types of read-only constant data are emitted by the compiler into the application code image and should be stored in non-volatile memory. Therefore, this type of read-only (constant) data must be managed equivalently to software instructions, especially since it is also part of the application image.

[0079] In some examples of the present invention, the first integrated circuit includes processing structures, such as a processor core or other bus masters (e.g., a direct memory access (DMA) controller or a coprocessor), which are depicted in an exemplary manner as bus masters 250 to 258. Any number of (one or more) bus masters 250 - 258 and any type of bus master (i.e., a processor core, a coprocessor, or a peripheral device that can also act as a bus master) can be implemented. These processing structures are connected to an interconnect 264 (which can be a single interconnect or a set of specific interconnects) via a plurality of bus interfaces (not shown in the figure). The interconnect 264 connects these bus masters to bus controllers that respond to access requests, such as an internal memory 266 or peripherals represented in an exemplary manner by peripheral - P0 252 and peripheral - P1 254, which are typically connected to the interconnect via a specific peripheral bridge 256. And, in some examples, a direct connection from the peripheral to the interconnect 264 may be possible. In many cases, the internal memory 266 (which implements, for example, the main system memory for storing application data) is composed of multiple memory blocks, which can be organized in random access memory (RAM) banks, shown in an exemplary manner here as RAM banks 268 to 269.

[0080] The first part of the instruction cache implemented within the information processing system 200 uses the internal memory 266 as an intermediate memory for the instructions and constant data of the application program read from the external memory 216; in addition to the traditional use of this memory, it is also used to store the data being processed by this application program. For cost reasons, the size of the internal memory is desired to be smaller than the size of the application program, and the available constant data 228 is provided by the application program image in the second semiconductor circuit 222. Therefore, only a part of the application program and the constant data 228 stored within the second semiconductor circuit 222 may be available for execution and read access within the cache storage device, and the cache memory is implemented by a part of the main memory 266 dedicated for this purpose. A typical ratio of this size would be 25% - 75% of the size of the application program and the constant data 228 in the external memory 216 stored within the second semiconductor circuit 222, which exceeds the memory required to store the application program data. Alternatively, in one example embodiment, the maximum supported application program size of 16 megabytes may be provided within the external NVM memory, and an additional 4 megabytes are required to store its application program data. For this example, assuming a ratio of 50%, in addition to the 4 megabytes of constant data 228 (such that at least 12 megabytes of main memory need to be implemented), the corresponding size requirement of the main memory would be 50% of 16 megabytes (i.e., 8 megabytes).

[0081] The second part of the instruction cache implemented within the information processing system 200 uses a hardware controller circuit, such as Figure 1Preferred embodiments of the HW controller 150. The hardware controller is operable to perform at least three separate functions, such as observation, data loading, and data management, and / or includes circuitry for performing said functions. In some examples, this functionality may be implemented within a single circuit. In one exemplary embodiment, a first portion of the hardware controller functionality is implemented within the load control circuit 230, which is connected to the external NVM memory 216 and is configured to access application code and constant data 228 stored in the external NVM memory 216. A second portion of the hardware controller functionality is the observation function, depicted here by a set of observation circuits 262, where one of these observation circuits is observing the processing of the bus masters 250-258, particularly observing the accesses corresponding to software execution instructions and the related accesses to the application code and constant data 228. For this purpose, the observation circuit 262 may be connected to the interface 251 that connects the bus masters 250-258 to the interconnect 264. It is contemplated that in other examples, other connectors may also be used when other connectors permit identifying the actual requests and their sources. A third portion of the hardware controller functionality is the cache control circuit 260, which uses the observation information provided by the observation circuit 262 and controls the load control circuit 230 to load data from the external memory 216. Generally, similar to the peripheral devices P0 252, P1 254, such a cache control circuit 260 interfaces with software via a set of programmable registers. Like these peripheral devices, the cache control circuit 260 is directly or via a bridge 256 connected to the interconnect 264 (where Figure 2 this potential connection is not shown in order to avoid obscuring the description).

[0082] According to some exemplary embodiments of the present invention, the hardware controller reads metadata information from the LU table 220 in order to determine the load unit it uses to manage the instruction cache (contents). In a preferred embodiment, this metadata is located within the internal memory 266 of the data structure, which contains metadata information about each identified LU instance. In another exemplary embodiment, the metadata information for each identified LU instance may also be stored in the external memory 216 in the form of read-only information and loaded directly from this memory. In another embodiment, the metadata information may be provided in the external memory 216 and loaded into the internal memory 266 before or during the start of the application loading. To be able to access this metadata, the hardware controller as well as Figure 2In the example, the cache control circuit 260 uses one or more additional connectors 270 to the interconnect 264, and the one or more additional connectors 270 at least permit access to the content of the internal memory 266. Additionally, the load control circuit 230 requires a similar ability to access the internal memory 266 in order to be able to store the application code (instructions and constant data) 228 read by the load control circuit 230 from the external memory 216 into the internal memory 266. For this purpose, in some examples, the load control circuit 230 may use or reuse one of the additional connectors 270 of the cache control circuit 260 (which must be connected to this connector in any case), or in some examples, the load control circuit 230 may alternatively use its own dedicated (optional) connector 226, which at least permits write access to the internal memory 266 (highlighted in a separate dashed box in the figure for clarity only).

[0083] According to some example embodiments of the present invention, the cache control circuit 260 accesses the metadata 210 provided in the LU table 220 located within the internal memory 266 via a set of dedicated metadata containers (or elements) 224. Each instance of these metadata containers 224 is operable to receive the metadata associated with a single LU instance; and is used by a specific functionality within the cache control circuit 260. In this way, the independent use of these metadata containers 224 within the cache control circuit 260 permits beneficial concurrency of the corresponding data, which makes the processing faster and avoids any potential stalls or locking situations that might otherwise be caused by multiple functional requests accessing the metadata 210 information of the LU instance currently being processed. The nature of hardware is that it can only support a fixed number of storage locations, which makes it very important to identify the minimum number of such containers to provide an optimal number of metadata containers 224 that support the cache management functionality; the optimal number of metadata containers 224 is the minimum set of containers that permits the maximum concurrency of the cache management operations. One or more of these containers may also be used to save or record the raw data collected during the execution observation, and this raw data can be written back 212 to the LU metadata table 220, which generates a permanent copy of this information within the LU metadata 210.

[0084] Now referring to Figure 3 , an example method is shown for using a memory as a cache to load a collection of instructions and constant data from an external memory using a variable-sized load unit, with beneficial cache management supported by the execution observation of the cache management 300 according to some example embodiments of the present invention. And, this example uses a tool-based generation process to identify those load units from the application software. Similar to Figure 1 , Figure 3Identify the advanced use of LU metadata by the hardware (HW) controller circuit 350, which is a preferred embodiment of the controller that manages such instruction caches. To avoid confusing the diagrams and the description, other elements required to execute or such instruction caches are not shown, such as the main memory for storing the load unit or the processing element that runs the application software. Additionally, elements or processing steps equivalent to those in Figure 1 are identified in this description by the corresponding numbers of the elements or processing steps in the previous figure. It is also worth noting that in this exemplary embodiment, the information for generating the LU table 328 can be separated from the use of this information in a timely manner and locally by the HW controller circuit 350. This involves the fact that the tool flow for generating the LU table content can be executed by a computer system different from the information system 200 that executes the application generated through the software development process 110; and the execution of the tool flow is typically part of or after software development, which is different from the time when the information system 200 executes the application.

[0085] In a manner similar to that previously described Figure 1 the classical software development environment 110 generates the corresponding application code, which is then analyzed 322 by the tool flow 320. The main operations of this process are performed by the LU generation software tool 324 within the tool flow 320, which generates metadata 326 regarding the construction of the application code within a set of LU instances. Subsequently, this LU metadata is stored in a database, such as the LU table 328. Since the application executes the software of the application at the granularity of functions, the output of the tool flow 320 is useful LUs that combine one or more software functions, which are themselves assumed to be atomic elements through this LU generation step.

[0086] The initial content of the LU table 328 is in the same way as Figure 1Generated in a manner similar to the content of the LU table 128 described therein, which contains at least equivalent information that can be additionally augmented with additional information. According to some examples of the present invention, the metadata generated by the LU generation tool 324 can optionally be additionally augmented with LU attributes 312 that can be provided by or generated from such information by software developers or users based on, for example, knowledge about the software or its behavior. In some examples, for the data content of one or more applications, it is contemplated that the attributes that can be defined by the user may include at least one of the following: code details (such as security-related code, shared code, startup code, critical functionality, maintenance code, etc.) and aspects of code usage (usage frequency, urgency of use, code values, replacement requirements). Additionally, other information can be recorded in the LU metadata generated by observing the execution of the application code, or in the LU metadata generated by the LU generation tool flow based on the information collected by observing this execution. The following paragraphs detail examples corresponding to this data and its use through some exemplary embodiments.

[0087] Figure 3 The left side reflects the cache controller hardware, while Figure 3 the right side reflects the associated tool flow, both of which are aspects of the processing of the application code provided by the application software development process 110. The tool flow identifies the LU instance structure of this application code, while the cache controller hardware handles the loading and storage of these LU instances in the internal main memory. The only feedback to the software development process 314 is through the linker control file 316 in order to affect the LU association and decoding location in the memory. Thus, Figure 3 the left side identifies the advanced use of LU metadata through some preferred exemplary embodiments of the HW controller circuit 350 that manages such instruction caches in the context of executing the application software. Here, the HW controller circuit 350 employs a control data management circuit 354 that reads the metadata 352 from the LU table 328 and provides this data to any part of the HW controller circuit 350 that needs this information. For example, one functional unit that needs this metadata 352 is the cache management circuit 366, which is configured to control the instruction cache management. For this purpose, the cache management circuit 366 can directly use the metadata 352 information without any additional processing.

[0088] In some examples, especially when the information / metadata 352 may be beneficial for determining arbitration or selection criteria, an additional processing operation is performed by the weight calculation circuit 380. The cache management circuit 366 and the weight calculation circuit 380 can use other information for processing in a similar manner to Figure 1 the dynamically generated data 164 of

[0089] In some examples in accordance with the present invention, there is a beneficial interaction between the loading of metadata 352, the use of the loaded metadata 352 by the weight calculation circuit 380, and the use of the weight calculation results. In this context, in an exemplary embodiment, the control data management circuit 354 loads only the information of interest into a specific metadata container element provided for this purpose. Thus, in some examples, there is a single container element available for any particular cache management operation, which is then used to calculate the weights specific to the corresponding operation. In some examples, the weight calculation formula for each specific weight can be controlled by software, and the software can specify the usage (for example, by selecting whether to include the constituent usage enable bits), and can specify the sign and the associated factor for each constituent used in the weight calculation (for example, by selecting the constituent multiplication factors defined within a range such as *256...*1 and *1 / 2...*1 / 256). A constituent in this context is one of the calculation factors of the weight calculation factors. Additionally, in some examples, as Figure 1 described in, it is possible to combine some of these inputs with the dynamic data 164. In some examples, some of the calculations can be LU-specific, or cache state-specific, or application-specific (for example, task-related, startup code, or software in an emergency).

[0090] In some examples, thereafter, the actual calculation of the weight value can be the sum of intermediate results, where each intermediate result is generated from the selected constituents after applying the sign and multiplying it by the associated factor. This enables the calculation of LU-specific weights to take into account important aspects of the application, the instruction cache, and its usage. Additionally, in particular, it permits the reuse of the same original input with completely different meanings in different aspects of cache management, which is an important factor in reducing the amount of original data required to control cache management operations. This method is very beneficial for reducing the amount of storage required for this data: a) within the HW controller 350, which can be an important cost factor in a hardware implementation; and b) in the LU table 328, where the amount of data required affects the amount of main memory required for this purpose. These factors play an important role when considering the amount of LU instances managed by the preferred exemplary embodiment. Additionally, in some examples, the ability to control weight calculation using formulas that can be controlled by software provides flexibility to support the different requirements that different application settings may exhibit to be supported by the same hardware implementation.

[0091] In some examples, one function supervised by the cache management circuit 366 is to load LUs. For this purpose, the HW controller 350 typically includes or uses (when one is available in addition to the HW controller 350) a load control unit 168 that performs this function. For this purpose, the load control unit 168 is connected to the external NVM memory 170.

[0092] In some examples, another function of the cache management circuit 366 is to remove LUs from the cache to make room for the next load operation, sometimes referred to as cache eviction. In some examples, both LU processing operations require weight calculations: (i) loading request weights to enable arbitration between multiple LU instances that may be (pre)loaded; and (ii) determining the weight of the LU for cache eviction, e.g., the goal is to select the LU with the smallest value as the next replacement eviction candidate. Clearly, the load request and load eviction instances are different (since only LU instances already in the cache can be evicted, and only LU instances not in the cache need to be loaded), so it is inevitable to use different metadata 352. Therefore, examples of the present invention propose using different containers (referred to herein as LSH and SEC, respectively) for the corresponding metadata. This allows these two operations to be separated and performed independently and simultaneously by different parts of the weight calculation circuit 380 and the cache management circuit 366.

[0093] Other examples of functions that require this metadata 352 are the LU observation circuit 360, which identifies events or attributes of interest in the context of application execution using this instruction cache. For this purpose, the observation circuit 360 is advantageously coupled to the control data management circuit 360 to be able to access the corresponding metadata information. Since the instruction cache is configured to be managed using LUs, these events are typically related to the load unit. In some examples, the collected metadata information may include, for example, the number of times an LU is used, the execution time of the LU, the time the LU was changed, or the corresponding latency, optionally further qualified with the previous cache state of the LU. Since one intention of this observation activity is to identify potential optimizations related to these LUs, it is beneficial to collect such metadata information related to the LUs.

[0094] Another beneficial use of the collected metadata information can be the use of the weight calculation circuit 380. Thus, in one exemplary embodiment, the observation circuit 360 is able to write this metadata information into the metadata container of the corresponding LU instance 385 managed by the control data management circuit 354. The control data management circuit 354 of one exemplary embodiment is additionally able to subsequently write the collected information 356 into the LU table 328. Writing the collected information 356 into the LU table 328 in this context can have several benefits, such as:

[0095] i) It provides a reporting channel that does not require an additional observation channel,

[0096] ii) allows this data to be used by the executed application software, and

[0097] iii) continuously records the observed data for later use.

[0098] In some examples, the observation can be an access during the execution of the application code. In some examples, a beneficial later use of the recorded observation information (which thus becomes part of the LU metadata) can be the use in weight calculation. In this way, these calculations can be loop-optimized based on the observation results. Examples of metadata reflecting performance data (such as raw general data) collected during the execution of the application in the LU table 328 can be one or more of the following: the amount of cache miss events associated with a specific cache state of the corresponding LU, the associated latency (such as minimum and / or maximum and / or average) associated with a specific cache state of the corresponding LU, the execution runtime-related information of the corresponding LU (such as minimum and / or maximum and / or average execution time), etc.

[0099] In some examples, another beneficial use of the recorded observation information can be the use of the LU generation tool 324 by reading back 330 from the LU table 328. In this way, reading back 330 the recorded observation information from the LU table 328 can occur, for example, during the development phase in a software development environment, and during the development phase in a real-world environment after some actual use of the application. Since this information is now part of the LU metadata, this information can be collected for multiple software executions and thus used to combine information of interest in different scenarios that occur only in different software executions. In some examples, it can additionally be used to advantageously identify differences in the information collected for those different scenarios. In some examples, another beneficial use can be the use of the LU generation tool 324 to update or optimize other metadata information based on the information extracted from this data, which can now show a more comprehensive view of the potential behavior of the instruction cache. Last but not least, the LU generation tool 324 can additionally use the information extracted from this data to provide optimization feedback 314 to the software development environment 110; for example, by updating or enhancing the linker control file 316 used to include improvements based on the results obtained or acquired from this observation information, thereby generating a modified location and / or LU of the application function or a set of application functions. According to some examples of the present invention, data collected by the cache controller circuit 350 or acquired by other means such as by analyzing functions can be used to further refine the metadata content of the LU table 328 for a single LU instance in this table.

[0100] In some examples, it is contemplated that any previous definition of the LU metadata can be enhanced based on data collected during one application execution run (or multiple runs, such as during an iterative process). In some examples, it is contemplated that the information collected can consist of raw information; for example, the amount of LU misses, the latency involved, the execution time, the sorting information, etc. In some examples, it is contemplated that such collected information may involve complex calculations [e.g., execution ratio: = (total code run time) / code size], and may also involve data from multiple runs; the corresponding calculations can be performed by the LU generation tool flow based on the raw data collected. In some examples, it is contemplated that specific attributes can be used to further improve the information collected, such as based on development observations, application knowledge, or user input.

[0101] The above usage options for the collected observation information identify multiple different usage scenarios, each of which can be optimized and refined on its own:

[0102] i) Immediately use the collected information during runtime through weight calculation;

[0103] ii) Combine the information collected across multiple executions for optimization via the tool flow, and new metadata can also be generated based on the evaluation of the corresponding results;

[0104] iii) Use the metadata updated based on such findings for subsequent use in the instruction cache;

[0105] iv) Use the observations or subsequent information obtained from its evaluation to control possible further improvements in the software development environment; and

[0106] v) Use the observations or subsequent information obtained from its evaluation to identify LU attributes based on code knowledge.

[0107] Thus, it can be envisioned that any of the above processing operations can use the collected data in different ways and from different perspectives. Advantageously, having an intensive recording function that only records the original data can minimize the associated processing steps and minimize the storage requirements for this information. The ability to account for different uses and different viewpoints via programmable weight calculation within the HW controller 350 makes the use of the same data applicable and useful in all those scenarios, which results in a significant size reduction and further reduces the accompanying hardware required to support this function. In some examples, it can also be envisioned that any reduction in this metadata and the number of container elements will directly affect the associated hardware cost. Using common raw metadata can further reduce the storage required for the LU table, which is also an important benefit.

[0108] Thus, in Figures 1-3 each of which describes an example of the present invention, where Figure 2 describes an exemplary complete embedded system, while Figure 1 and Figure 3 mainly describe the cache controller circuit employed by such a system that interacts with the tool flow to provide relevant metadata information, where the tool flow and its interactions can be performed by different systems during the software development phase. Additionally, Figure 3 also shows the reception of observation information for optimization and its impact on the support environment. It can be envisioned that in some alternative examples, the functionality and processing of the cache controller can occur in a system different from the metadata preparation and post-processing performed by the support environment (e.g., the tool flow) as described, and / or occur at different times in some instances (e.g., after the execution of the application and before the response). However, this part of the processing is very important for understanding the beneficial use of metadata across (multiple) runs and as feedback for LU generation and application software generation. For completeness, it can be envisioned that Figures 1-3 each of which only shows a part of the possible functionality relevant to the purpose of the corresponding interaction described, i.e., in Figure 1 andFigure 3 is the interaction with the software development environment and tool flow, and in Figure 2 is the interaction with other components in the system implementing the corresponding instruction cache. Therefore, to avoid confusion in the description of the present invention, other standard components and circuits that can be used in such a memory system are not shown.

[0109] Now refer to Figure 4 , according to some exemplary embodiments of the present invention, the illustrated graph 400 shows the benefits of using different criteria for cache management. According to some examples of the present invention, different weight values can have different objectives, especially when there are different arbitration or selection criteria. According to some examples, the objective of a load request operation (e.g., determining the most urgently needed LU so as to minimize the time of the load operation) and the objective of a cache eviction operation (e.g., determining the LU with the lowest value that can be replaced, thus optimizing the value of the instruction cache content) are ideal examples of such different criteria. However, it is conceivable that any one of these criteria may need to be calculated on the same LU instance (when a load is required, or it has been effectively determined as a candidate for replacement), but not necessarily simultaneously (it is unlikely that an LU needs to be loaded into the cache and evicted from the cache at the same time). It is conceivable that in some examples, some of these different criteria may just use the same basic information in different ways; for example, load request arbitration may prefer smaller LU instances (because they can be loaded faster), while cache eviction may prefer larger LU instances for replacement (because they provide more available space). The common basic raw data is the size of the LU instance, which can be preferentially used for both weight calculations, thus reducing the amount of data to be stored. In addition, some attributes stored in the LU metadata may be related to the load request (e.g., load urgency hint), but not related to cache eviction, and vice versa (e.g., data value hint). Therefore, restricting the LU metadata to a subset of the attributes that are relevant to both of these functions will be restricted.

[0110] Figure 3The weight calculation circuit 380 is configured to generate weight values from its inputs and then use them as arbitration or selection criteria. For this purpose, using weights as normalization values allows for a common definition of corresponding criteria across multiple applications. When it is possible to encode different aspects of using this data in different ways, using common raw data for weight calculation can reduce the amount of metadata required. In some examples, this is achieved through a programmable calculation formula specific to weights, which allows for different signs and independent quantifiers to be assigned to each component of the calculation. In this case, the term "component" refers to and includes any single input to the calculation, which, as previously mentioned, can come from LU metadata or be provided by the internal information of the cache controller circuit. Examples of weight calculations using common raw data performed by the weight calculation circuit 380 and using different controls (e.g., using only two of a larger set of possible input values) include the following.

[0111] The first example of the load request weight calculation formula uses two common raw inputs, <LU size> and <state#1 delay>, which reflect the delay experienced in the case of a severe (state#1) cache miss. The goal of this load request weight calculation formula is to prefer smaller LU instances that have experienced a greater delay in previous cases; where the latter component is, for example, 4 times more important than the size aspect. Potential factors include:

[0112] (i) The sign of <LU size>: minus sign (preferably taking the smaller value);

[0113] (ii) The factor of <LU size>: ×2;

[0114] (iii) The sign of <state#1 delay>: plus sign (preferably taking the larger delay); and

[0115] (iv) The factor of <state#1 delay>: ×16 (this factor is 4 times more important).

[0116] Based on this set of factor examples, the example (first) request weight calculation can be:

[0117] Request_weight := sign(LU_size)*factor(LU_size)*<LU_size> + sign(state#1delay)*factor(state#1 delay)*<state#1 delay> [1]

[0118] The second example of the cache eviction weight calculation formula again uses two common raw inputs, which are also <lusize>Sum factor <startup>, which are only applicable when the application software indicates that the startup phase has been completed, as reflected by some dynamic data provided by software programmable registers. The goal of this cache eviction weight calculation formula is to prefer LU instances that are only used during the startup phase (but only after this phase is completed), otherwise eviction will preferably take larger LU instances. Potential factors include:

[0119] (i) The sign of <LU size>: plus sign (preferably taking the larger value);

[0120] (ii) The factor of <LU size>: ×4;

[0121] (iii) <startup>Symbol: ;

[0122] (iv) <startup>Value: basically 1 (when the internal status indicates that the startup is completed), otherwise 0;

[0123] (v) <startup>Factor: ×64 (This factor is 16 times more important).

[0124] Based on this factor set example, the example (second) eviction weight calculation can be:

[0125] Eviction_weight: = sign(LU_size)*factor(LU_size)* <lusize>+factor(startup)* <startup>[2]

[0126] In the above example, both weights include a common component <LU_size> and a second specific component, namely the cache eviction weight of <startup>An indication flag, corresponding to <state#1_delay> in the case of loading a request weight. Other examples of preferred embodiments may use a more general raw input; for example, a preferred embodiment calculates each weight based on approximately 30 components that can be individually controlled, for example, by a calculation formula.

[0127] Now refer to Figure 5 , which shows an example timing diagram 500 of the components employed in the system diagram of Figure 1 , Figure 2 or Figure 3 according to some example embodiments of the present invention, and the use of these components during a load unit switch. Figure 5 The use of the containers SEC 560, LSH 570, ALE 580, and HBE 590 is also identified along the timeline 505.

[0128] In this example, the cache replacement operation continuously selects new candidates for replacement. The relevant metadata container SEC 560 is periodically loaded with the metadata of the selected LU, and in some examples, SEC 560 is used to calculate the weight of this replacement (at 510, 511, 512,... 519). In this example, the prediction mechanism irregularly selects new LUs to be predicted (identified by events (A) 520, 521, 522, 523). Whenever a new LU is selected (in this example, any one of the above events (A)), the relevant metadata container LSH 570 is loaded with the LU metadata. In some examples, an observation identification is performed by executing an event (LU switch event (B) 540, 541, 542, 543, 544) that hits an address in another LU. The contents of the containers ALE 580 and HBE 590 change in a specific order after such an event, and this specific order will be detailed for one of these events in the extended timing diagram 550. The extended timing diagram 550 identifies the relevant processing of two subsequent switches (e.g., LU switch events (B) 540, 541) and the use of the containers ALE 580 and HBE 590. The relevant modifications to the contents of these two containers ALE 580 and HBE 590 occur in these LU switch events 530, 531, 532, 533, 534.

[0129] More specifically, as shown below, the extended timing diagram 550 identifies the related use of the metadata containers ALE and HBE. After detecting the LU swap event (B) that occurs at 540 whenever the address accessed by the observed bus master is within another LU, the current content of the ALE element 580 is immediately moved to the HBE element 590 at time 550. In some examples, this movement can occur almost immediately; for example, within a single cycle. Immediately following the movement of the current content of the ALE element 580 to the HBE element 590, the metadata of the next executing LU is loaded into the ALE element 580 at 560 (the ALE element 580 is free after moving its previous content to the HBE). This facilitates the rapid loading of the corresponding LU metadata of the next executing LU, which may be urgently required for cache management operations. In some examples, performance data associated with the new LU, such as any latency associated with LU switching, can now be immediately stored in the ALE container 580. Now, any remaining performance data to be collected that is related to the previous LU and corresponds to the LU switch 540 (such as the execution time of the previous LU) can be stored in the HBE container 590 without causing any interference to the operations related to the next executing LU. Once this data collection is complete, the content of the HBE metadata container 590 can be written to the system RAM at 570. Using the second metadata container HBE has several benefits:

[0130] i) It permits the immediate loading of the LU metadata of the next LU (which is urgently required),

[0131] ii) It avoids waiting for the contained ALE metadata to be freed by writing back the collected performance data,

[0132] iii) It permits the independent collection of performance data corresponding to the previous LU and the next LU, and

[0133] iv) It permits the delayed write-back of the collected performance data corresponding to the previous LU.

[0134] Specifically, the last (iv) benefit is useful because both i) loading the LU metadata of the next LU and ii) writing back the performance information collected for the previous LU must access a common resource, namely the LU table, and cannot be performed simultaneously when this resource is located in the system RAM.

[0135] Since this operation of writing to the system RAM at 570 must occur before the next LU switch at 541, the use of a second container (i.e., HBE container 590 in this example) permits a beneficial separation of all the required operations listed above in the event of an LU switch. In this case, the use of two metadata containers for LU switch observations permits maximum concurrency to be achieved with minimal required hardware work (e.g., using two containers instead of a single container) through the above-described processing operations. In this way, the separation advantageously avoids any internal stalling / locking of any associated operations that would otherwise be required to avoid information loss.

[0136] Reference now Figure 6 , showing a method of controlling a circuit (e.g., a controller) by a hardware controller according to some example embodiments of the present invention. Figure 3 The hardware controller circuit 350) performs an observation operation, i.e., an example flowchart 600 of LU switching observation. Figure 5 Subsequently, any weight calculation operations (e.g., any cache eviction weight calculations associated with operations {510, ..., 519, ...}, and any cache eviction weight calculations associated with operations {510, ..., 519, ...}) Figure 5 Any load request weight calculations associated with the operations {520, ..., 523, ...} shown) can be performed independently and simultaneously with the execution observation ( Figure 6 In this example, LU switching observations associated with a single bus master are performed (e.g., triggered by a corresponding need for such weight calculations, e.g., due to identification of load predictions or eviction requirements). Figure 2 The processor core 250 is depicted executing application software (e.g., by Figure 2 The observation element 262 of the hardware controller circuit 350 depicted is Figure 3 The hardware controller circuit 350 performs LU switching observation).

[0137] In an example embodiment employing multiple bus masters (e.g., Figure 2 In the example embodiment depicted in FIG. 1 , there are multiple such LU switching observations performed simultaneously, which are performed by Figure 6 605. This potential diversity and concurrency of operations is one of the many reasons for the need for arbitration, which is supported by using weights as arbitration criteria for cache controller load requests and cache eviction operations.

[0138] After the reset operation at 610, a first processing sequence 607 is executed, in which at 620 the LU metadata of the new (currently used) LU (i.e., an LU instance with an address range owned by the application software to be executed by the bus master next) is loaded from the LU table located in the internal memory into the ALE element. Possibly simultaneously with this operation, at 625 the corresponding original data is recorded, which is part of the new LU metadata required by the associated bus master. In some examples, the recording at 625 can also be performed after loading the LU metadata.

[0139] At 630, it is urgently necessary for the LU metadata to be processed by the cache controller to control cache activities related to the newly used LU. Examples of the corresponding information are, for example, the start address and size of this LU required for the address mapping function. Such address mapping is sometimes a function required by such an instruction cache in order to allow the application software loaded from the external memory to be executed at any location within the internal memory. In some examples of the present invention, it is conceivable that after using the LU metadata at 630, the corresponding processing operation is performed by the cache controller (e.g., HW controller circuit 350).

[0140] After performing the required processing (e.g., using the metadata) operation at 630, when the loaded LU is already available (cached) in the internal memory, the application software can use the loaded LU; otherwise, before the application software can use the instructions and constant data related to the newly used LU, these instructions and constant data must be loaded. Another example of an operation that requires LU metadata is an emergency load operation required when the newly used LU has not been loaded. Therefore, in some examples, an optional "perform LU load" operation can be adopted at 640. When the instructions and constant data related to the newly used LU have been loaded into the internal memory, the corresponding bus master can execute the application code contained in / associated with this LU at 650; then, at 645, the observation part of the cache control circuit observes this execution. In the case of an LU that has already been cached, this operation can be entered immediately without having to perform the LU load operation at 640.

[0141] While the application code is being executed (starting), at 655, the cache controller records the raw data associated with the start of the execution of the associated bus master. This recording is performed in any case, whether the LU has to be loaded or the LU is already cached. In some examples, one purpose of this data can be to record information related to cache management, e.g., the corresponding latency within the performance data. In some examples, different latencies may be recorded when the corresponding LU has to be loaded (e.g., at 640) or when the LU is already in the cache, so that only the cache management information has to be updated.

[0142] It is beneficial that the cache controller can record the raw data at 625 and 655 while the associated processing operations are performed at 620 and the corresponding 650, because this allows for the correct observation of the latency associated with 645. Additionally, it is beneficial because it enables those associated operations to be performed as soon as possible, otherwise the execution of the cache or the application may be affected due to the need to record the collected data. Thus, in some examples, recording the raw data associated with some cache management operations or the effects of those cache management operations (such as loading the LU) occurs substantially simultaneously with those associated operations, such as loading a new LU at 640 or starting the execution of the application at 650.

[0143] At this time, the execution of the application code contained within the newly used LU is performed by the associated bus master / processor core, which is observed by the cache controller circuit at 645 until a change to the new LU is detected at 660. In this case, the processing 607 (shown as 620 to 660) that has been described and the metadata related to the next LU loaded into the ALE element at 620 are repeated.

[0144] After 660, when a LU change is detected, a second, particularly independent processing operation 665 is started. Advantageously, the second independent processing operation 665 is performed independently and simultaneously with the previously described processing 607. The only relevance of these additional processing operations 665 to the previous processing operation 607 is the starting point after the first change of the new LU 660 is detected, and this processing operation needs to be completed before another second change of the new LU 660 is detected (which may be the result of repeating steps 620 to 660 of operation 607).

[0145] In some examples of the preferred embodiment, detecting an LU change can trigger two activities: (i) moving the LU metadata contained in the ALE element to the HBE element at 670; and (ii) loading new content into the ALE container at 620. Although the latter is used to repeat the previously described processing loop, moving the metadata to the second container will enable independent and parallel processing of this data, which will now be described.

[0146] At 680, performance data corresponding to the LU change (and related to the previously executed LU) is recorded. At 690, the recorded LU metadata is written to the RAM. In this way, the write operation continuously records the information within the LU table. In some examples, the write operation can be performed later, but it must occur before another new LU change is detected at 660.

[0147] As previously described, the two execution sequences 607 (including steps 620...660) and 665 (composed of steps 670...690) are executed simultaneously, and the general detection of the LU change 660 is used as the only synchronization element. And, by providing a second metadata container (HBE), rather than using a single metadata container, this concurrent execution can be achieved. Thus, providing the second metadata container (HBE) enables beneficial separation and reordering of operations that would otherwise have to be performed sequentially and in a less favorable order.

[0148] Some examples of the improvements provided by this separation can be understood by looking at the operations required without adopting these concepts. For example, the collected raw data must be written at 690 before new data can be loaded into the ALE element at 620 in order to avoid overwriting this data. By adopting the concepts described herein, the LU metadata related to the newly used LU can be loaded more urgently before the collected observation information is written. Additionally, for example, the LU metadata contained in the LU table cannot be read and written simultaneously. In contrast, by adopting the concepts described herein, delaying the write permits separation of the two operations and permits writing this information later when this data does not need to be loaded urgently. As another example, moving the LU metadata from the ALE to the HBE container (reading data) can be performed simultaneously with the new use of the ALE container at 620 and 625 (while this data is being written). Unless one of these operations itself requires a long processing time (which is likely for loading the LU metadata in 620), any one of these operations can be performed simultaneously within a single cycle.

[0149] Accordingly, examples of the present invention provide one or more of the following features not disclosed in known memory systems. First, the cache examples described herein employ a criterion where the load requests (request weights) and cache evictions (eviction weights) of LUs use different calculation formulas for these two weights. Additionally, the cache examples described herein identify the use of metadata to specify a large number of load units used by the instruction cache. Second, these examples describe beneficial support for the instruction cache, which consists of a relatively large available internal memory, that is, when the application image is stored in external memory, the instruction cache accounts for 25% - 75% of the application image. Additionally, for the purpose of caching instructions, some of the examples describe the use of internal memory. Third, some examples describe the use of internal memory that can be operated by additional instruction execution or memory access observation. In some examples, in cases where the memory system suffers from severe bandwidth limitations, for example, in combination with a relative latency difference greater than the relative latency difference experienced between the internal memory type (usually within one order of magnitude of a clock cycle) and the external memory (greater than a single clock magnitude), the cache examples described herein can be beneficially used. In this way, the cache examples described herein provide a mechanism to keep the most valuable data elements, thereby avoiding frequent reloading (note that known processor caches do not have such bandwidth limitations). Additionally, in the cache examples described herein, preloading must meet critical timing constraints, which means it cannot be traded off with the LUs being requested. Fourth, the cache examples described herein do not use metadata about load units as a basis for determining the weight criterion; rather, some examples use it as metadata for the (request weight and eviction weight) calculation formulas. Fifth, it should be noted that the processor cache does not use metadata. Sixth, the cache examples described herein benefit from general raw data with different uses, thereby significantly reducing this data, which is important. Seventh, the cache examples described herein use a minimum amount of internal storage within the cache controller hardware to load metadata, such as metadata container elements to support the observation and weight calculation (two weights) for each observed bus master. Eighth, the cache examples described herein identify the use of observations in cache management and related code execution aspects, and in some examples record the corresponding performance data. This cache example feature can provide multiple benefits. For example, the related code execution aspects can be used as a debugging aid, allowing optimization at runtime and after / across multiple runs, and allowing identification of problematic LU elements, which can then be specially processed.This cache example feature may also affect the software execution of applications supported by the cache memory; this is due to the following relevant performance data: the number of execution cycles in the load unit and the stall cycles observed due to cache misses. Additionally, the combination of these two observations provides key information for controlling the cache. Ninth, the cache examples described herein support collecting the observed performance data during multiple runs. Therefore, the cache examples described herein can be executed during development, where typically there are multiple verification / test runs covering different scenarios, and all runs combined cover a complete set of intended use cases. Tenth, the cache examples described herein can use the collected performance data and / or internal information to influence weight calculation. The benefit of this feature is that it allows for a higher degree of dynamic control over cache usage, enabling it to react to changing system states (startup, normal, emergency), special situations (a large number of cache misses, severe latency), or allow for LU-specific behavior (simply saving LUs that always or often cause problems).

[0150] Those skilled in the art will understand that the integration level of the hardware controller or tool flow or control data management circuit or components may depend on the implementation in some instances. Obviously, the various components within the hardware controller or tool flow or control data management circuit can be implemented in discrete or integrated component form, and thus the final structure is application-specific or a design choice.

[0151] Since the illustrated embodiments of the present invention can be implemented to a large extent using electronic components and circuits well known to those skilled in the art, details will not be explained below to a greater extent than considered necessary for an understanding and appreciation of the underlying concepts of the present invention and to avoid obscuring or distracting from the teachings of the present invention.

[0152] In the foregoing specification, the present invention has been described with reference to specific examples of embodiments of the invention. However, it will be apparent that various modifications and changes can be made herein without departing from the scope of the invention as set forth in the appended claims, and the claims are not limited to the specific examples described above.

[0153] The connections discussed herein can be any type of connection suitable for transmitting signals from or to corresponding nodes, units, or devices, for example, via an intermediate device. Thus, unless otherwise implied or stated, the connections can be, for example, direct connections or indirect connections. Connections can be shown or described as a single connection, multiple connections, unidirectional connections, or bidirectional connections. However, different embodiments can vary the implementation of the connections. For example, separate unidirectional connections can be used instead of a bidirectional connection, and vice versa. Also, multiple connections can be replaced by a single connection that transmits multiple signals in a continuous manner or in a time-division multiplexing manner. Similarly, a single connection carrying multiple signals can be divided into various different connections carrying subsets of those signals. Thus, there are many alternatives for transmitting signals.

[0154] Those skilled in the art will recognize that the architectures depicted herein are merely exemplary, and that many other architectures can actually be implemented to achieve the same functionality. Any arrangement of components that achieves the same functionality is effectively 'associated' so as to achieve the desired functionality. Thus, any two components that are combined herein to achieve a particular functionality can be considered to be 'associated' with each other so as to achieve the desired functionality, regardless of the architecture or intermediate components. Similarly, any two components so associated can also be considered to be "operably connected" or "operably coupled" to each other to achieve the desired functionality.

[0155] In addition, those skilled in the art should recognize that the boundaries between the operations described above are merely illustrative. Multiple operations can be combined into a single operation, a single operation can be distributed over additional operations, and the execution of operations can at least partially overlap in time. Moreover, alternative embodiments can include multiple instances of a particular operation, and the order of operations can be changed in various other embodiments. And for example, in one embodiment, the illustrated examples can be implemented as circuitry located on a single integrated circuit or within the same device. Alternatively, the circuit and / or component examples can be implemented as any number of separate integrated circuits or separate devices interconnected in a suitable manner. And for example, the examples described herein or portions thereof can be implemented as a software or code representation of a physical circuit system or, for example, in any suitable type of hardware description language, can be transformed into a logical representation of a physical circuit system.

[0156] Moreover, the present invention is not limited to physical devices or units implemented in non-programmable hardware, but can also be applied to programmable devices or units capable of performing the desired anti-spyware countermeasures by operating according to appropriate program codes. For example, microcomputers, personal computers, notebooks, personal digital assistants, video games, automobiles, and other embedded information systems, which are collectively referred to as 'computer systems' in this application. However, it is contemplated that other modifications, variations, and alternative solutions are also possible. Therefore, the specification and the drawings should be regarded as illustrative rather than restrictive.

[0157] In the claims, any reference signs placed in parentheses shall not be construed as limiting the claim. The word 'comprising' does not exclude the presence of other elements or steps than those listed in a claim. Further, as used herein, the term "a" or "an" is defined as one or more than one. Also, the use of introductory phrases such as 'at least one' and 'one or more' in the claims should not be construed as implying that another claim element introduced by the indefinite article "a" or "an" limits any particular claim containing such introduced claim element to inventions containing only one such element, even when the same claim includes the introductory phrase 'one or more' or 'at least one' and the indefinite article "a" or "an". This also applies to the use of the definite article. Unless otherwise stated, terms such as 'first' and'second' are used arbitrarily to distinguish such elements described by such terms. Therefore, these terms are not necessarily intended to indicate a temporal or other precedence of such elements. The mere fact that certain measures are recited in mutually different claim items does not indicate that a combination of these measures cannot be used to advantage.< / startup> < / startup> < / lusize> < / startup> < / startup> < / startup> < / startup> < / lusize>

Claims

1. An embedded information system (200), characterized in that, Comprising: A load control circuit (168, 230) capable of being coupled to an external memory (170, 216), the external memory (170, 216) containing instructions and constant data associated with at least one application code (228) of a software application; At least one processor (250...258), the at least one processor (250...258) being coupled to at least one interconnect (264) and configured to execute the at least one application code (228); An internal memory (266), the internal memory (266) being coupled to the at least one interconnect (264) and configured to store, as a main system memory, data for the software application in a first portion of the internal memory (266), and being configured as a cache to store the instructions and the constant data for executing the at least one application code (228) from the external memory (170, 216) in a second portion of the internal memory (266); A cache control circuit (150, 260, 350), the cache control circuit (150, 260, 350) being coupled to the at least one interconnect (264) and additionally coupled to or including the load control circuit (168, 230), wherein the instructions and the constant data are organized by a load unit LU of variable size, and wherein at least one attribute of the LU is specified within metadata (128, 220, 328); Wherein the load control circuit (168, 230) is configured to load the LU associated with the at least one application code (228) from the external memory (170, 216) into the internal memory (128, 266, 328) in the granularity of a single LU; and Wherein the cache control circuit (150, 260, 350) manages the second portion of the internal memory (266) configured as a cache based on the metadata corresponding to the LU by being configured to perform the following operations: Observing (645) the execution of at least a portion of the at least one application code (228) by detecting at least one of the following: the LU being executed, the Change within the internal memory (266) from one LU to another LU; Loading metadata information corresponding to the LU instance from the external memory (170, 216) or the internal memory (266); Designating the next LU to be loaded by the load control circuit (168, 230) into the second portion of the internal memory (266); and When there is not enough space to load the next LU, designating the next LU to be evicted from the second portion of the internal memory (266).

2. The embedded information system according to claim 1, characterized in that, The first request to load the next LU into the internal memory (128, 266, 328) and the second request to evict a LU from the internal memory (128, 266, 328) are different operations that occur independently by the cache control circuit (150, 260, 350) based on at least one of the following: at different times, associated with different LU instances.

3. The embedded information system (200) according to claim 2, characterized in that, The first request to load the first LU instance corresponds to a first weight for selecting the first LU, and the second request to evict the second LU instance corresponds to a second weight for selecting the second LU.

4. The embedded information system (200) according to claim 3, characterized in that, The first weight and the second weight are independently calculated using different weight calculation formulas that use metadata information constituted by at least one of the first weight and the second weight.

5. The embedded information system (200) according to claim 3 or claim 4, characterized in that, The cache control circuit (150, 260, 350) includes a first dedicated hardware container LSH and a second dedicated hardware container SEC for storing metadata associated with a single LU instance, where: The metadata in the LSH is used to calculate and manage the first weight for the first request to load the next LU from the external memory (170, 216) into the internal memory (266), and The metadata in the SEC is used to calculate and manage the second weight for the second request to evict a LU from the internal memory (266).

6. The embedded information system (200) according to claim 5, wherein The LU metadata in the first dedicated hardware container LSH and the second dedicated hardware container SEC includes the same general raw data for different LU instances, and the general raw data is used in different ways when calculating the first weight and the second weight for the same LU instance.

7. The embedded information system (200) according to claim 5, wherein The calculation of at least one of the first weight and the second weight includes: the sum of intermediate weight results generated by selected constitutions of the weight, where at least one of these constitutions is derived from the raw data provided by the LU metadata after applying a sign and multiplying the selected constitution by a relevant factor.

8. The embedded information system (200) according to claim 3, wherein The cache control circuit (150, 260, 350) is configured to record at least one of the following using a record observation of the execution of at least a portion of the application code (288): a previously used LU, the next LU, or a change from the previously used LU to the next LU.

9. The embedded information system (200) according to claim 8, characterized in that, The metadata is augmented with the record observation of the execution of at least a portion of the application code (228) and additional internal information (164) that can be used in the cache control circuit (150, 260, 350), and the additional internal information (164) includes at least another constitution derived from or augmented with at least one of the following: the application program state reported to the cache control circuit (150, 260, 350), the internal state of the cache control circuit (150, 260, 350), the state of the LU associated with the first weight or the second weight (228).

10. A method (600) for memory management in an embedded information system (200), characterized in that, The embedded information system (200) includes at least one processor (250...258), at least one interconnect (264), an internal memory (266), and cache control circuits (150, 260, 350) coupled to or including load control circuits (168, 230), wherein the method includes: Connecting the load control circuits (168, 230) to an external memory (170, 216) containing instructions and constant data associated with at least one application code (228) of a software application; Organizing the instructions and the constant data of the at least one application code (228) by a load unit LU of variable size, wherein at least one attribute of the LU is specified within metadata (128, 220, 328); Configuring the internal memory (266) to store data for the software application as a main system memory in a first part of the internal memory (266), and configured as a cache to store the instructions and the constant data for executing the at least one application code (228) loaded from the external memory (170, 216) in a second part of the internal memory (266); Loading (640) at least a portion of the at least one application code (228) from the external memory (170, 216) into the internal memory (266) in a single LU granularity; Executing (650) at least a portion of the at least one application code (228) associated with at least one LU instance located within the internal memory (266), and Managing the second part of the internal memory (266) configured as a cache based on the metadata (128, 220, 328) corresponding to the LU by: Observing (645) the execution of at least a portion of the at least one application code (228) by detecting at least one of the following: an LU being executed, a change within the internal memory (266) from one LU to another; Loading (620) metadata information corresponding to the LU instance from the external memory (170, 216) or the internal memory (266); Designating the next LU to be loaded into the second part of the internal memory (266); and When there is not enough space to load the next LU, designating the next LU to be evicted from the second part of the internal memory (266).

Citation Information

Patent Citations

  • Method and apparatus for optimizing code execution using annotated trace information having performance indicator and counter information

    US20050155026A1

  • Methods to efficiently implement coarse granularity cache eviction

    US9892044B1