Multi-level shared Cache modeling and static analysis method
Through the user-configured multi-level shared Cache model and static analysis method, the problem of simple and inability to adapt to diversity in the existing multi-core processor WCET analysis tool is solved, and flexible modeling and accurate analysis of modern multi-core processor Cache architecture is realized, improving the accuracy of WCET analysis.
Patent Information
- Application Number
- CN202411712747.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-06-06
AI Technical Summary
The Cache model of existing multi-core processor WCET analysis tools is relatively simple and cannot adapt to the diversity of modern multi-core processors, especially the inability to effectively capture the Cache interference behavior between cores, resulting in WCET analysis being too pessimistic.
A multi-level shared Cache modeling and static analysis method is proposed. By configuring the parameters of the cache series, number of partitions, size, connection degree and other parameters of the cache, a configurable multi-level shared Cache model is established, and static Cache analysis is carried out in combination with abstract interpretation technology, independent analysis and inter-core interference analysis are analyzed to accurately simulate the exclusive and sharing of the Cache partition by the kernel.
This method can be applied to modern domestic multi-core processor Cache structures with different forms, solves the problem of diversity in Cache architecture, improves the accuracy and flexibility of WCET analysis, and reduces the pessimism of Cache interference analysis.
Smart Images

Figure CN120104548A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of real-time system analysis, and in particular to a multi-level shared Cache modeling and static analysis method. Background Art
[0002] With the rapid development of semiconductor technology, multi-core processors have been widely used in real-time control systems such as aviation, automobiles, and industrial automation. In order to balance the speed difference between the core and the main memory while reducing hardware costs, modern multi-core processors often use a multi-level cache storage architecture. However, a core's access to an exclusive cache will inevitably delay the execution of another core's task. The replacement of a core to a shared cache may cause the instructions or data required by another core to be replaced, thereby delaying the execution of the core's task. For a specific shared cache structure, how to capture the maximum interference behavior between cores and obtain the worst execution time of the task has always been the focus of research in the field of WCET analysis of multi-core processors.
[0003] It is usually assumed that all storage blocks mapped to the same cache line in a multi-core processor share the cache, but this will inevitably interfere with each other and cannot be used. The multi-level shared cache structures adopted by different processor manufacturers are different, making it difficult to cross-test and use them. In addition, the cache model adopted by the existing multi-core processor WCET analysis tools is relatively simple, and most of them are dual-core processors. They mainly maintain their own first-level caches through two processor cores and jointly manage the second-level caches, but obviously cannot meet the use of multi-core processors. Therefore, a multi-level shared cache modeling and static analysis method is urgently needed to be applicable to the modern domestic multi-core processor cache structures of various forms, promote the development of multi-core processor WCET analysis technology, solve the diversity problem of existing processor cache architectures, and solve the technical problems that the currently constructed cache model is relatively simple and not flexible enough and can only adapt to the fixed cache structure of two-core processors. Summary of the invention
[0004] In view of this, the present invention proposes a multi-level shared Cache modeling and static analysis method, in which the user determines the number of Cache levels, the number of partitions in each level of Cache, the size of a single partition, the degree of connectivity, the block size, the mapping method, etc. Each Cache partition can correspond to one or more cores, and can simulate the exclusive and sharing of the Cache partitions by the cores in the processor, so as to solve the technical problems of the diversity of the Cache architecture of existing domestic processors and the fact that the currently constructed Cache model is relatively simple and not flexible enough and can only adapt to the fixed Cache structure of a two-core processor. Based on the abstractly interpreted static Cache analysis model, independent cache analysis and inter-core cache interference analysis are performed to solve the problem that the current multi-core processor shared Cache interference analysis is too pessimistic.
[0005] In order to achieve the above technical objectives, the specific technical solutions adopted by the present invention are:
[0006] A configurable multi-level shared cache modeling and static analysis method includes the following steps:
[0007] S1. Establish a Cache architecture model;
[0008] S2. Establish a static cache analysis model based on abstract interpretation;
[0009] S3, updating and merging the two models in step S2 in step S1 through MUST analysis and MAY analysis;
[0010] S4, confirming the status of the instruction output in step S2 through MUST analysis and MAY analysis;
[0011] S5, perform inter-core shared cache interference analysis on the cache simulation analysis results of the multi-core processor; in step 5, in the case of different cores of the multi-core processor, determine the task timing relationship. If the tasks are assigned to different cores for execution, there is a timing relationship between the tasks. When task task1 on core 1 requires task task0 on core 0 to start execution under certain conditions, that is, task1 is executed after task0, so task1 and task0 cannot have a conflict on shared resources, otherwise there is a conflict;
[0012]
[0013] In the case of a multi-core processor with the same core, the task timing relationship is judged. If two tasks are assigned to the same core for execution, then there must be a sequence relationship between the two tasks, and it is impossible for them to compete for shared resources, and there will be no competition for shared resources between them.
[0014]
[0015] In the case of multi-core processors with multi-level caches with diversified allocation, the task timing relationship is judged. If two tasks do not share the same cache group in the i-level cache, the tasks cannot cause cache conflicts. On the contrary, if they share the same cache group in the i-level cache, cache conflicts will occur.
[0016]
[0017] Furthermore, in step S1, the architecture model of each cache partition is established as follows:
[0018] S i,j =(Nset i,j ,Bsize i,j ,Assoc i,j ,Repl i,j )
[0019] The structure of the jth partition of the i-th level cache is mainly composed of Nset i,j , Bsize i,j 、Assoc i,j and Repl i,j Four parameters determine, Nset i,j Indicates the number of groups owned by the cache partition, Bsize i,j Indicates the block size of the partition when performing cache replacement, Assoc i,j Represents the connectivity of the partition, that is, the number of blocks in each group, Repl i,j Indicates the strategy adopted by the cache partition for cache replacement, using LRU or FIFO strategy. The capacity of the partition can be expressed as:
[0020] Capacity i,j =Nset i,j ×Bsize i,j ×Assoc i,j
[0021] The overall cache architecture model is established as follows:
[0022] S={S i,j |i <m and j<n i}
[0023] Where m represents the number of cache levels of the processor, and represents the number of partitions of the i-th level cache. In modern processor architectures, the first-level cache is private to each core. By setting ni =1 to achieve.
[0024] Furthermore, in step S1, the mapping mode between the cache and the main memory is determined. There are three mapping modes: direct mapping, fully associative mapping and set associative mapping. Both fully associative mapping and direct mapping are regarded as two special forms of set associative mapping. Direct mapping and fully associative mapping are set and implemented.
[0025] Fully connected mapping: Nset i,j =1, the connectivity is equal to the number of cache blocks within the partition;
[0026] Direct Mapping: Assoc i,j =1, the number of groups is equal to the number of cache blocks inside the partition.
[0027] Furthermore, in step S2, in the static cache analysis process based on abstract interpretation, the analysis process is divided into independent cache analysis and inter-core cache interference analysis. The independent analysis performs multi-level analysis on the instruction sequence executed by each core and determines the instruction state. The inter-core interference analysis determines the possible interference between instructions executed by different cores, and adjusts the instruction state according to the number of interferences. After the cache state classification, the instruction or memory block state can be divided into five types: AH, AM, NC, PS, and AX. The description of each state and the analysis method adopted are as follows:
[0028] The AH state means that the access will definitely hit, so the MUST analysis method is used. The AM state means that the access will definitely not hit, so the MAY analysis method is used. The PS state means that the first access will not hit but the subsequent accesses can hit, so the PERSISTENCE analysis method is used. The NC state means that it cannot be determined, so the other analysis methods are used.
[0029] Furthermore, in step S2, the cache analysis model based on the abstract interpretation technology is constructed as follows:
[0030] CA M =(CS,C R ,AC s ,AC U ,AC J )
[0031] Among them: CS: Cache structure configuration, including cache capacity, block size, set associativity, and address mapping method;
[0032] C R : Cache replacement strategy, such as LRU replacement strategy, FIFO replacement strategy and random replacement strategy;
[0033] AC s: Abstract Cache state set, that is, the abstract state set of memory blocks stored in the Cache under all program nodes;
[0034] AC U : Abstract Cache update function;
[0035] AC J : Abstract Cache merge function;
[0036] After defining the cache analysis model based on the specific cache structure, replacement strategy, abstract state, update function and merge function, combined with the instruction sequence and initial state of the test program, the cache state analysis model of the test program is constructed as follows:
[0037] CA result =(S MA ,S init ,CA M )
[0038] in:
[0039] S MA : The program execution sequence mainly includes the division information of the program basic blocks and the instruction sequence within each basic block, and the basic blocks are arranged in order according to the program structure;
[0040] S init : Initial abstract cache state. The incoming edge of each CFG graph node corresponds to an entry state set in all cache groups. The state set is initially set to empty.
[0041] CA M :Static cache analysis model, determined based on specific cache structure, replacement strategy, abstract state, update function and merge function.
[0042] Furthermore, in step S3, the update function and the merge function are slightly different for different analysis methods, and the cache update and merge under different analysis methods are performed based on the LRU replacement strategy.
[0043] Furthermore, in step S3, the cache update is performed in the following three cases:
[0044] S301: The instruction to be executed in the node does not exist in the entry state Acs_in. In this case, no matter what type of analysis is performed, the update function is the same, and the function definition is as follows:
[0045] UPDATE(S set ,m)={l 1 →m}∪{l i →S set (li-1 )|i=2,3...A}
[0046] Where: m: the memory block where the instruction to be executed in the node is located, indicating the cache group to which the memory block m is mapped;
[0047] A: Associativity, which indicates the number of cache lines contained in each group of the set-associative mapping;
[0048] S302: The instruction to be executed in the node exists in the entry state of the node. At this time, the update function of MUST analysis is defined as follows:
[0049] MUST U (S set ,m)={l 1 →m}∪{l i →S set (l i-1 )|i=2,3...h-1}
[0050] ∪{l h →S set (l h-1 )∪Differ(m,S set (l h ))},if m∈S set (l h )
[0051] Where: l 1 →m: means placing storage block m in the first position of the set group;
[0052] l i →S set (l i-1 ): means placing the storage block at the position in the original state of the set group to the position;
[0053] Differ(m,S set (l h ): indicates the storage blocks different from m in the set group of the Cache;
[0054] S303, the instruction to be executed in the node exists in the entry state of the node. At this time, the update function of MAY analysis is defined as follows:
[0055] MAY U (S set ,m)={l 1 →m}∪{l i →S set (l i-1 )|i=2,3...h}
[0056] ∪{lh+1 →S set (l h )∪Differ(m,S set (l h ))},if m∈S set (l h )
[0057] Where:
[0058] l 1 →m: means placing storage block m in the first position of the set group;
[0059] l i →S set (l i-1 ): means placing the storage block at the position in the original state of the set group to the position;
[0060] Differ(m,S set (l h ): indicates the storage blocks different from m in the set group of the Cache;
[0061] Cache merging is performed in the following two situations:
[0062] The merge function under MUST analysis is defined as follows:
[0063]
[0064] The merge function for MAY analysis is defined as follows:
[0065]
[0066] Where:
[0067] S 1 , S 2 : Indicates the cache status corresponding to the two incoming edges entering a CFG node, which contains all cache group information and the storage blocks corresponding to different cache lines in each group;
[0068] MUST J (S 1 ,S 2 ): indicates the result of merging two abstract cache states;
[0069] l a , l b : Represents cache line information in two abstract cache states;
[0070] addr(l a )、addr(lb ): Indicates the memory addresses corresponding to the storage blocks temporarily stored in the two cache lines. Equal addresses indicate that the two are the same memory block.
[0071] Furthermore, in step S4, after a given node enters a state and instruction sequence, the instruction state can be finally determined by analyzing the offset shift generated by the preceding instruction in the same node.
[0072] Furthermore, in step S4, MUST analysis will determine the instructions that can be hit, and count the number of predecessor instructions from the first instruction of the node to the current instruction to be executed and the instruction to be analyzed inst that are mapped to the same cache group, that is, the offset. If the sum of age and shift of the instruction to be analyzed inst in the node entry state is greater than the associativity, the instruction may be evicted before execution, and the hit situation is NC. If age+shift is less than the associativity, it means that in the worst case, the instruction to be executed cannot be evicted from the cache before execution, and the hit situation is AH.
[0073] MAY analysis will determine the instructions that cannot be hit. If the instruction to be executed in the node does not exist in the node's cache entry state, the hit status of the instruction is AM. If the instruction to be executed on the node exists in the node's cache entry state, the number of predecessor instructions mapped to the same cache group as the instruction to be analyzed inst is counted from the first instruction of the node to the current instruction to be executed, that is, the offset. If the sum of the age and shift of the instruction to be analyzed inst in the node entry state is greater than the connectivity, the instruction will be evicted before execution, and the hit status is AM. If age+shift is less than the connectivity, it means that the instruction to be executed may still exist in the cache before execution, and the hit status is NC.
[0074] Furthermore, in step S5, when tasks running on a multi-core processor share a cache, the replacement behavior of tasks running on a certain core will cause cache blocks needed by other cores to be replaced. When memory blocks cannot be hit in the private cache but can be hit in the shared cache, these memory blocks are analyzed. We call the set of these memory blocks hit_set (Li), where i represents the cache level.
[0075] For the memory block MB belonging to hit_set(Li), traverse other tasks and count the number of memory blocks in other tasks that are the same as map(addr(MB)), which we call the number of conflicts nconflicts. Combine the age of the memory block in the group with the number of conflicts nconflicts to detect whether the memory block exceeds the degree of association. If it exceeds the degree of association, convert the AH state to NC. The formula is defined as follows:
[0076]
[0077] By adopting the above technical solution, the present invention can also bring the following beneficial effects:
[0078] 1. The multi-level shared cache modeling and static analysis method mentioned in the present invention constructs a configurable multi-level shared cache model by abstractly modeling the cache architecture of modern multi-core processors. The model is suitable for the cache structures of modern domestic multi-core processors of various forms, and solves the problem of the diversity of the cache architecture of modern multi-core processors.
[0079] 2. The present invention mentions a multi-level shared cache modeling and static analysis method, and then establishes a static cache analysis model based on abstract interpretation, which is used for independent cache analysis and inter-core cache interference analysis. The state of each instruction or storage block at the corresponding execution point is obtained through static cache analysis, and the time delay caused by different states is combined as the basis for the schedulable analysis of real-time system tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0081] Figure 1 A flowchart of a multi-level shared cache modeling and static analysis method mentioned in the present invention;
[0082] Figure 2 It is a multi-level shared cache architecture in this embodiment;
[0083] Figure 3 It is the static cache analysis framework in this embodiment;
[0084] Figure 4 This is a flowchart of Cache status merging in this embodiment;
[0085] Figure 5This is a flow chart for determining the MUST analysis instruction status in this embodiment. DETAILED DESCRIPTION
[0086] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0087] The following describes the embodiments of the present invention through specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.
[0088] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present invention, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this device and / or practice this method.
[0089] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention. The drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0090] Additionally, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, it will be understood by those skilled in the art that the aspects described may be practiced without these specific details.
[0091] Example 1
[0092] In one embodiment of the present invention, see Figure 1This embodiment provides a multi-level shared cache modeling and static analysis method, including the following steps: first, establish a cache architecture model, then establish a static cache analysis model based on abstract interpretation, update and merge the cache analysis model, then confirm the instruction status, and finally perform inter-core shared cache interference analysis.
[0093] Reference Figures 2 to 5 The specific steps of the configurable multi-level shared Cache modeling and static analysis method of the present invention are as follows:
[0094] 1. Establishment of Cache Architecture Model
[0095] The architecture model of each cache partition is as follows Figure 2 As shown, the establishment is as follows:
[0096] S i,j =(Nset i,j ,Bsize i,j ,Assoc i,j ,Repl i,j )
[0097] The structure of the jth partition of the i-th level cache is mainly composed of Nset i,j , Bsize i,j 、Assoc i,j and Repl i,j Four parameters determine, Nset i,j Indicates the number of groups owned by the cache partition, Bsize i,j Indicates the block size of the partition when performing cache replacement, Assoc i,j Represents the connectivity of the partition, that is, the number of blocks in each group, Repl i,j Indicates the strategy used by the cache partition for cache replacement, such as LRU or FIFO strategy. The capacity of the partition can be expressed as:
[0098] Capacity i,j =Nset i,j ×Bsize i,j ×Assoc i,j
[0099] The overall cache architecture model is established as follows:
[0100] S={S i,j |i <m and j<n i}
[0101] Where m represents the number of cache levels of the processor, n iIndicates the number of partitions owned by the i-th level cache. Since the first-level cache is private to each core in modern processor architectures, it can be set by n i =1 to achieve.
[0102] Determine the mapping method between cache and main memory. There are three mapping methods: direct mapping, fully associative mapping, and set associative mapping. Fully associative mapping and direct mapping can be regarded as two special forms of set associative mapping. By setting Nset i,j and Assoc i,j Implement direct mapping and fully connected mapping.
[0103] Fully connected mapping: Nset i,j =1, the connectivity is equal to the number of cache blocks inside the partition.
[0104] Direct Mapping: Assoc i,j =1, the number of groups is equal to the number of cache blocks inside the partition.
[0105] like Figure 2 As shown, the multi-level shared cache architecture can be simply summarized as follows: L1 cache consists of several partitions, each partition structure consists of Nset i,j , Bsize i,j 、Assoc i,j and Repl i,j Four parameters, Nset i,j Indicates the number of groups owned by the cache partition, Bsize i,j Indicates the block size of the partition when performing cache replacement, Assoc i,j Represents the connectivity of the partition, that is, the number of blocks in each group, Repl i,j Indicates the strategy adopted by the cache partition for cache replacement, such as LRU or FIFO strategy. Recursively, each level of cache partition can be constructed using these four parameters.
[0106] 2. Establishment of static cache analysis model based on abstract interpretation
[0107] In the static cache analysis process based on abstract interpretation, the analysis process is divided into independent cache analysis and inter-core cache interference analysis. Independent analysis performs multi-level analysis on the instruction sequence executed by each core and determines the instruction state. Inter-core interference analysis determines the possible interference between instructions executed by different cores and adjusts the instruction state according to the number of interferences. After cache state classification, the instruction or memory block state can be divided into five types: AH, AM, NC, PS, and AX. The description of each state and the analysis method used are as follows.
[0108] The AH state means that the access will definitely hit, so the MUST analysis method is used. The AM state means that the access will definitely not hit, so the MAY analysis method is used. The PS state means that the first access will not hit, but the subsequent accesses can hit, so the PERSISTENCE analysis method is used. The NC state means that it cannot be determined, so the other analysis methods are used.
[0109] The cache analysis model based on abstract interpretation technology is constructed as follows:
[0110] CA M =(CS,C R ,AC s ,AC U ,AC J )
[0111] in:
[0112] CS: Cache structure configuration, including cache capacity, block size, set associativity, and address mapping method;
[0113] C R : Cache replacement strategy, such as LRU replacement strategy, FIFO replacement strategy and random replacement strategy;
[0114] AC s : Abstract Cache state set, that is, the abstract state set of memory blocks stored in the Cache under all program nodes;
[0115] AC U : Abstract Cache update function;
[0116] AC J : Abstract Cache merge function.
[0117] After defining the cache analysis model based on the specific cache structure, replacement strategy, abstract state, update function and merge function, combined with the instruction sequence and initial state of the test program, the cache state analysis model of the test program is constructed as follows:
[0118] CA result =(S MA ,S init ,CA M )
[0119] in:
[0120] S MA : The program execution sequence mainly includes the division information of the program basic blocks and the instruction sequence within each basic block, and the basic blocks are arranged in order according to the program structure;
[0121] Sinit : Initial abstract cache state. The incoming edge of each CFG graph node corresponds to an entry state set in all cache groups. The state set is initially set to empty.
[0122] CA M : Static cache analysis model, determined based on specific cache structure, replacement strategy, abstract state, update function and merge function.
[0123] like Figure 3 As shown in the figure, the static cache analysis framework based on abstract interpretation can be simply summarized as follows: core 0 and core 1 share a group in L2 and L3 cache, core 2 and core 3 are independently grouped in L2 cache, and share a group in L3 cache. First, L1 cache analysis is performed independently on core 0 and core 1, and the shared L2 Bank1 cache is analyzed through filters. After the independent cache analysis of core 0 and core 1 is completed, the analysis results are filtered again through filters to perform L2 Bank1 cache interference analysis of the two cores. The interference analysis results are filtered by filters, and then L3 Bank1 cache analysis is performed separately and independently. After the analysis is completed, filter filtering is performed, and the filtering results are subjected to L3 Bank1 cache interference analysis. Similarly, the cache analysis process of tasks on core 2 and core 3 can be obtained, and finally the status is confirmed based on the comprehensive analysis results.
[0124] 3. Cache update and merge
[0125] In the control flow graph of the program, each node may have multiple incoming edges. Due to different execution contexts, the cache state corresponding to each incoming edge is different. By merging the states of multiple incoming edges into one state through the merging method, the analysis results can be guaranteed to be safe while reducing path searches. For different analysis methods (MUST analysis, MAY analysis), the update function and merge function are slightly different. This step will perform cache updates and merges under different analysis methods based on the LRU replacement strategy.
[0126] Cache updates are divided into the following three situations:
[0127] 1) The instruction to be executed in the node does not exist in the entry state Acs_in. In this case, regardless of the type of analysis, the update function is the same, and the function definition is as follows:
[0128] UPDATE(S set ,m)={l 1 →m}∪{l i →S set (l i-1)|i=2,3...A}
[0129] Where:
[0130] m: the memory block where the instruction to be executed in the node is located, S set =MAP(m) indicates the cache group to which the memory block m is mapped;
[0131] A: Associativity, which indicates the number of cache lines contained in each group of the set-associative mapping.
[0132] 2) The instructions to be executed in the node exist in the node's entry state. At this time, the update function of MUST analysis is defined as follows:
[0133] MUST U (S set ,m)={l 1 →m}∪{l i →S set (l i-1 )|i=2,3...h-1}
[0134] ∪{l h →S set (l h-1 )∪Differ(m,S set (l h ))},if m∈S set (l h )
[0135] Where:
[0136] l 1 →m: means placing storage block m in the first position of the set group;
[0137] l i →S set (l i-1 ): means to change the original state of the set group to l i The storage block at location l is placed i-1 Location;
[0138] Differ(m,S set (l h ): indicates a storage block different from m in the set group of the Cache.
[0139] 3) The instructions to be executed in the node exist in the node's entry state. At this time, the update function of MAY analysis is defined as follows:
[0140] MAY U (S set ,m)={l 1 →m}∪{li →S set (l i-1 )|i=2,3...h}
[0141] ∪{l h+1 →S set (l h )∪Differ(m,S set (l h ))},if m∈S set (l h )
[0142] Where:
[0143] l 1 →m: means placing storage block m in the first position of the set group;
[0144] l i →S set (l i-1 ): means placing the storage block at the position in the original state of the set group to the position;
[0145] Differ(m,S set (l h ): indicates a storage block different from m in the set group of the Cache.
[0146] The merge function under MUST analysis is defined as follows:
[0147]
[0148] The merge function for MAY analysis is defined as follows:
[0149]
[0150] Where:
[0151] S 1 , S 2 : Indicates the cache status corresponding to the two incoming edges entering a CFG node, which contains all cache group information and the storage blocks corresponding to different cache lines in each group;
[0152] MUST J (S 1 ,S 2 ): indicates the result of merging two abstract cache states;
[0153] l a , l b : Represents cache line information in two abstract cache states;
[0154] addr(l a )、addr(l b ): Indicates the memory addresses corresponding to the storage blocks temporarily stored in the two cache lines. Equal addresses indicate that the two are the same memory block.
[0155] like Figure 4 As shown in the figure, the cache state merging process can be simply summarized as follows: first, the state of all incoming edges of the node is merged to determine whether the node has a predecessor node. If so, each set in the cache is merged separately, and then the entry state of the set(i) of the current node is merged with the exit state of the set(i) of the predecessor node. After the merging is completed, it is determined whether the merging of all groups is completed. If not, the state merging is repeated for each set in the cache. If it is completed, the state merging is ended and the result is used as the entry state of the node. If the node has no predecessor node, the state merging is ended directly and the result is used as the entry state of the node.
[0156] Step 4: Command status confirmation method
[0157] A program control flow graph node generally contains multiple instructions. Even if an instruction exists in the node's entry state, it is possible that the instruction cannot be hit during execution due to the replacement of the previous instruction before execution. Therefore, after a given node entry state and instruction sequence, the offset shift generated by the previous instruction in the same node must be analyzed to finally determine the instruction state.
[0158] MUST analysis will determine the instructions that can be hit, and count the number of predecessor instructions from the first instruction of the node to the current instruction to be executed and the instruction to be analyzed inst that are mapped to the same cache group, that is, the offset. If the sum of age and shift of the instruction to be analyzed inst in the node entry state is greater than the associativity, the instruction may be evicted before execution, and the hit case is NC. If age+shift is less than the associativity, it means that in the worst case, the instruction to be executed cannot be evicted from the cache before execution, and the hit case is AH.
[0159] MAY analysis will determine the instructions that cannot be hit. If the instruction to be executed in the node does not exist in the node's cache state, the hit status of the instruction is AM. If the instruction to be executed in the node exists in the node's cache state, the number of predecessor instructions mapped to the same cache group as the instruction to be analyzed inst is counted from the first instruction of the node to the current instruction to be executed, that is, the offset. If the sum of the age and shift of the instruction to be analyzed inst in the node entry state is greater than the association degree, the instruction will be evicted before execution, and the hit status is AM. If age+shift is less than the association degree, it means that the instruction to be executed may still exist in the cache before execution, and the hit status is NC.
[0160] like Figure 5 As shown, the MUST analysis instruction status determination process can be simply summarized as follows: first initialize the instruction status space, then determine whether all instruction statuses in the node have been fully marked, if not all marking is required, calculate the memory block where the instruction is located, determine whether the memory block exists in the entry state of the node, if so, calculate the offset caused by the predecessor instruction of the instruction, and after the calculation is completed, determine whether the instruction is likely to be replaced before execution, if it cannot be replaced, mark the instruction status as AH, if it can be replaced, mark the instruction status as NC, then return to the beginning to determine whether all instruction statuses in the node have been fully marked, if the memory block does not exist in the entry state of the node, mark the instruction status as NC, if it is determined that all instruction statuses in the node have been fully marked, the process ends.
[0161] Step 5: Inter-core shared cache interference analysis method
[0162] When tasks running on a multi-core processor share a cache, the replacement behavior of tasks running on a core may cause cache blocks needed by other cores to be replaced. Therefore, when a memory block cannot be hit in the private cache but can be hit in the shared cache, these memory blocks are analyzed. We call the set of these memory blocks hit_set (li), where i represents the cache level.
[0163] For the memory block MB, traverse other tasks and count the number of memory blocks in other tasks that are the same as map(addr(MB)), which we call the number of conflicts nconflicts. Combine the age of the memory block in the group with the number of conflicts nconflicts to detect whether the memory block exceeds the degree of association. If it exceeds the degree of association, convert the AH state to NC. The formula is defined as follows:
[0164]
[0165] When performing interference analysis between tasks of a multi-core processor, it is not necessary to perform conflict analysis on all tasks. The present invention reduces the scope of interference analysis in the following three ways:
[0166] There is often a timing relationship between tasks in different cores of a multi-core processor. For example, task 1 on core 1 requires task 0 on core 0 to start executing under certain conditions, that is, task 1 is executed after task 0. Therefore, it is impossible for task 1 and task 0 to conflict on shared resources. Therefore, the timing relationship between tasks needs to be considered when considering shared resource conflicts.
[0167]
[0168] If two tasks are assigned to the same core for execution, then there must be a sequence relationship between the two tasks, and it is impossible for them to compete for shared resources, and there will be no competition for shared resources between them.
[0169]
[0170] If two tasks do not share the same cache group in the i-th level cache, then the task cannot cause cache conflict.
[0171]
[0172] In summary, the present invention constructs a configurable multi-level shared Cache model by abstractly modeling the Cache architecture of modern multi-core processors. The user determines the number of Cache levels, the number of partitions in each level of Cache, the size of a single partition, the degree of connectivity, the block size, the mapping method, etc. Each Cache partition can correspond to one or more cores, and can simulate the exclusive and shared Cache partitions of the cores in the processor. It is suitable for the Cache structures of modern domestic multi-core processors of various forms, and solves the problem of the diversity of Cache architectures of modern multi-core processors.
[0173] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A configurable multi-level shared cache modeling and static analysis method, characterized in that: The following steps are involved: S1. Establish a Cache architecture model; S2. Establish a static cache analysis model based on abstract interpretation; S3, updating and merging the two models in step S2 in step S1 through MUST analysis and MAY analysis; S4, confirming the status of the instruction output in step S2 through MUST analysis and MAY analysis; S5. Perform inter-core shared cache interference analysis based on the cache simulation analysis results of the multi-core processor; In step 5, in the case of different cores of the multi-core processor, the task timing relationship is judged. If the tasks are assigned to different cores for execution, there is a timing relationship between the tasks. When task 1 on core 1 requires task 0 on core 0 to start execution under certain conditions, that is, task 1 is executed after task 0, task 1 and task 0 cannot conflict on shared resources, otherwise there is a conflict. In the case of a multi-core processor with the same core, the task timing relationship is judged. If two tasks are assigned to the same core for execution, then there must be a sequence relationship between the two tasks, and it is impossible for them to compete for shared resources, and there will be no competition for shared resources between them. In the case of diversified allocation of multi-level cache in multi-core processors, a judgment is made on the task timing relationship. If two tasks do not share the same cache group in the i-level cache, the tasks cannot cause cache conflicts. Conversely, if they share the same cache group in the i-level cache, cache conflicts will occur.
2. A multi-level shared cache modeling and static analysis method as claimed in claim 1, characterized in that: In step S1, the architecture model of each cache partition is established as follows: S i,j =(Nset i,j ,Bsize i,j ,Assoc i,j ,Repl i,j ) The structure of the jth partition of the i-th level cache is mainly composed of Nset i,j , Bsize i,j 、Assoc i,j and Repl i,j Four parameters determine, Nset i,j Indicates the number of groups owned by the cache partition, Bsize i,j Indicates the block size of the partition when performing cache replacement, Assoc i,j Represents the connectivity of the partition, that is, the number of blocks in each group, Repl i,j Indicates the strategy adopted by the cache partition for cache replacement, using LRU or FIFO strategy. The capacity of the partition can be expressed as: Capacity i,j =Nset i,j ×Bsize i,j ×Assoc i,j The overall cache architecture model is established as follows: S={S i,j |i<m and j<n i } Where m represents the number of cache levels of the processor, and represents the number of partitions of the i-th level cache. In modern processor architectures, the first-level cache is private to each core. By setting n i =1 to achieve.
3. A multi-level shared cache modeling and static analysis method as claimed in claim 2, characterized in that: In the step S1, the mapping mode between the cache and the main memory is determined. There are three mapping modes: direct mapping, fully associative mapping and set associative mapping. Both fully associative mapping and direct mapping are regarded as two special forms of set associative mapping. Direct mapping and fully associative mapping are set and implemented. Fully connected mapping: Nset i,j =1, the connectivity is equal to the number of cache blocks within the partition; Direct Mapping: Assoc i,j =1, the number of groups is equal to the number of cache blocks inside the partition.
4. A multi-level shared cache modeling and static analysis method as claimed in claim 3, characterized in that: In step S2, in the static cache analysis process based on abstract interpretation, the analysis process is divided into independent cache analysis and inter-core cache interference analysis. The independent analysis performs multi-level analysis on the instruction sequence executed by each core and determines the instruction state. The inter-core interference analysis determines the possible interference between instructions executed by different cores, and adjusts the instruction state according to the number of interferences. After the cache state classification, the instruction or memory block state can be divided into five types: AH, AM, NC, PS, and AX. The description of each state and the analysis method adopted are as follows: The AH state means that the access will definitely hit, so the MUST analysis method is used. The AM state means that the access will definitely not hit, so the MAY analysis method is used. The PS state means that the first access will not hit but the subsequent accesses can hit, so the PERSISTENCE analysis method is used. The NC state means that it cannot be determined, so the other analysis methods are used.
5. A multi-level shared cache modeling and static analysis method as claimed in claim 4, characterized in that: In step S2, the cache analysis model based on the abstract interpretation technology is constructed as follows: CA M =(CS,C R , AND s , AND U , AND J ) Among them: CS: Cache structure configuration, including cache capacity, block size, set associativity, and address mapping method; C R : Cache replacement strategy, such as LRU replacement strategy, FIFO replacement strategy and random replacement strategy; AC s : Abstract Cache state set, that is, the abstract state set of memory blocks stored in the Cache under all program nodes; AC U : Abstract Cache update function; AC J : Abstract Cache merge function; After defining the cache analysis model based on the specific cache structure, replacement strategy, abstract state, update function and merge function, combined with the instruction sequence and initial state of the test program, the cache state analysis model of the test program is constructed as follows: THAT result =(S MA ,S init ,THAT M ) in: S MA : The program execution sequence mainly includes the division information of the program basic blocks and the instruction sequence within each basic block, and the basic blocks are arranged in order according to the program structure; S init : Initial abstract cache state. The incoming edge of each CFG graph node corresponds to an entry state set in all cache groups. The state set is initially set to empty. CA M : Static cache analysis model, determined based on specific cache structure, replacement strategy, abstract state, update function and merge function.
6. A multi-level shared cache modeling and static analysis method as claimed in claim 5, characterized in that: In step S3, the update function and the merge function are slightly different for different analysis methods, and the cache update and merge under different analysis methods are performed based on the LRU replacement strategy.
7. A multi-level shared cache modeling and static analysis method as claimed in claim 6, characterized in that: In step S3, cache update is performed in the following three situations: S301: The instruction to be executed in the node does not exist in the entry state Acs_in. In this case, no matter what type of analysis is performed, the update function is the same, and the function definition is as follows: UPDATE(S set ,m)={l1→m}∪{l i →S set (l i-1 )|i=2,3...A} Where: m: the memory block where the instruction to be executed in the node is located, indicating the cache group to which the memory block m is mapped; A: Associativity, which indicates the number of cache lines contained in each group of the set-associative mapping; S302: The instruction to be executed in the node exists in the entry state of the node. At this time, the update function of MUST analysis is defined as follows: MUST U (S set ,m)={l1→m}∪{l i →S set (l i-1 )|i=2,3...h-1} ∪{l h →S set (l h-1 )∪Differ(m,S set (l h ))},if m∈S set (l h ) Where: l1→m: means placing storage block m in the first position of the set group; l i →S set (l i-1 ): means placing the storage block at the position in the original state of the set group to the position; Differ(m,S set (l h ): indicates the storage blocks different from m in the set group of the Cache; S303, the instruction to be executed in the node exists in the entry state of the node. At this time, the update function of MAY analysis is defined as follows: MAY U (S set ,m)={l1→m}∪{l i →S set (l i-1 )|i=2,3...h} ∪{l h+1 →S set (l h )∪Differ(m,S set (l h ))},if m∈S set (l h ) Where: l1→m: means placing the storage block m in the first position of the set group; l i →S set (l i-1 ): means placing the storage block at the position in the original state of the set group to the position; Differ(m,S set (l h ): indicates the storage blocks different from m in the set group of the Cache; Cache merging is performed in the following two situations: The merge function under MUST analysis is defined as follows: The merge function for MAY analysis is defined as follows: Where: S1, S2: indicates the cache status corresponding to the two incoming edges entering a CFG node, which contains all cache group information and the storage blocks corresponding to different cache lines in each group; MUST J (S1, S2): represents the result of merging two abstract cache states; l a , l b : Represents cache line information in two abstract cache states; addr(l a )、addr(l b ): Indicates the memory addresses corresponding to the storage blocks temporarily stored in the two cache lines. Equal addresses indicate that the two are the same memory block.
8. A multi-level shared cache modeling and static analysis method as claimed in claim 7, characterized in that: In step S4, after a given node enters a state and instruction sequence, the instruction state can be finally determined by analyzing the offset shift generated by the preceding instruction in the same node.
9. A multi-level shared cache modeling and static analysis method as claimed in claim 8, characterized in that: In step S4, MUST analysis will determine the instructions that can be hit, and count the number of previous instructions from the first instruction of the node to the current instruction to be executed and the instruction to be analyzed inst mapped to the same cache group, that is, the offset. If the sum of age and shift of the instruction to be analyzed inst in the node entry state is greater than the associativity, the instruction may be evicted before execution, and the hit situation is NC. If age+shift is less than the associativity, it means that in the worst case, the instruction to be executed cannot be evicted from the cache before execution, and the hit situation is AH. MAY analysis will determine the instructions that cannot be hit. If the instruction to be executed in the node does not exist in the node's cache state, the hit status of the instruction is AM. If the instruction to be executed in the node exists in the node's cache state, the number of predecessor instructions mapped to the same cache group as the instruction to be analyzed inst is counted from the first instruction of the node to the current instruction to be executed, that is, the offset. If the sum of the age and shift of the instruction to be analyzed inst in the node entry state is greater than the association degree, the instruction will be evicted before execution, and the hit status is AM. If age+shift is less than the association degree, it means that the instruction to be executed may still exist in the cache before execution, and the hit status is NC.
10. A multi-level shared cache modeling and static analysis method as claimed in claim 9, characterized in that: In step S5, when tasks running on a multi-core processor share a cache, the replacement behavior of tasks running on a certain core will cause cache blocks needed by other cores to be replaced. When memory blocks cannot be hit in the private cache but can be hit in the shared cache, these memory blocks are analyzed. We call the set of these memory blocks hit_Set (Li), where i represents the level of the cache. For the memory block MB belonging to hit_Set(Li), traverse other tasks and count the number of memory blocks in other tasks that are the same as map(addr(MB)), which we call the number of conflicts nconflicts. Combine the age of the memory block in the group with the number of conflicts nconflicts to detect whether the memory block exceeds the degree of association. If it exceeds the degree of association, convert the AH state to NC. The formula is defined as follows: