Processor, display card, computer equipment and constant reading method

By implementing a cache sharing mechanism between processing cores in a multi-core processor, the problem of high latency in reading constant data is solved, reading efficiency is improved, transmission latency is reduced, and cache utilization is improved.

CN120596427APending Publication Date: 2025-09-05MOORE THREADS TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511093569.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In multi-core processors, the read latency of constant data is high, resulting in low control flow efficiency of the processing core. Although the existing technology reduces the read latency by equipping each processing core with an independent constant cache, the read latency from the last-level cache or main memory is high when the cache is missed.

Method used

Implement a cache sharing mechanism between processing cores, allowing the first processing core to obtain constant data from other processing cores when a cache miss occurs, making requests and responses through the on-chip network, avoiding direct reading from the last-level cache or main memory.

Benefits of technology

It improves the reading efficiency of the processing core in the case of constant data miss, reduces transmission latency, improves cache utilization and reduces redundant storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596427A_ABST
    Figure CN120596427A_ABST
Patent Text Reader

Abstract

The invention discloses a processor, a graphics card, computer equipment and a constant reading method, and belongs to the field of computer system structures. The processor comprises at least two processing cores; a first processing core in the at least two processing cores is used for acquiring a first constant cached in a second processing core under the condition that the first constant is not hit in a first cache; wherein the first cache is the cache of the first processing core. According to the processor, sharing of constants in caches of at least two processing cores is achieved, the first processing core is supported to directly request the first constant from the second processing core, and compared with a mode of requesting the first constant from a last-stage cache or a main memory in a traditional mode, the reading time delay is lower, and the reading efficiency is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer architecture, and in particular to a processor, a graphics card, a computer device, and a method for reading constants. Background Art

[0002] In multi-core processors, constant data (such as instruction data, program constants, and program control flow) is frequently read by each processing core. Furthermore, these constants are often used in control flow-related processes within each processing core, such as as conditions in conditional programs. High read latency for these constants can affect the control flow of the processing cores.

[0003] In order to reduce the read latency of constant data, related technologies often equip each processing core with an independent constant cache to reduce the read latency of constant data. Although this method can reduce the read latency of constant data to a certain extent, when a cache miss occurs in the constant cache of the processing core, that is, a constant data does not hit in the constant cache, the processing core needs to read the constant data from the last-level cache or main memory.

[0004] Since the last-level cache and main memory are usually located at the back end of the on-chip network, the process of reading constant data from the last-level cache or main memory has a high transmission delay, making the processing core less efficient in reading cache data in the event of a cache miss. Summary of the Invention

[0005] The present application provides a processor, a graphics card, a computer device, and a method for reading constants, and the technical solution is as follows.

[0006] According to one aspect of the present application, a processor is provided, the processor comprising at least two processing cores; The first processing core of the at least two processing cores is configured to obtain the first constant cached in the second processing core if the first constant does not hit in the first cache; The first cache is a cache of the first processing core.

[0007] According to one aspect of the present application, a graphics card is provided, comprising the above-mentioned processor.

[0008] According to one aspect of the present application, a computer device is provided, comprising the above-mentioned processor.

[0009] According to one aspect of the present application, a method for reading a constant is provided, the method being executed by a processor including at least two processing cores; the method comprising: A first processing core among the at least two processing cores obtains the first constant cached in a second processing core when the first constant does not hit in the first cache; The first cache is a cache of the first processing core.

[0010] The beneficial effects brought about by the technical solution provided in this application include at least the following.

[0011] This allows the first processing core to request the first constant from another processing core if the first constant is missed. Compared to the traditional method of requesting the first constant from the last-level cache or main memory, the cache in the processing core typically has higher read and write speeds than the last-level cache and main memory. This means that the second processing core can retrieve the first constant from the second cache faster than from the last-level cache or main memory. In other words, the second processing core's read latency is lower than the last-level cache or main memory's read latency. This indirectly speeds up the first constant's access through the second processing core, improving the first processing core's read efficiency. Furthermore, due to processor architecture, the transmission path (or transmission latency) between the first processing core and the last-level cache or main memory is often longer than the transmission path (or transmission latency) between the first processing core and the second processing core. This reduces the transmission latency between the first processing core's read request for the first constant from the second processing core and the second processing core's return of the read response to the first processing core. This speeds up the second processing core's access to the first constant and improves the first processing core's read efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0013] Figure 1 A schematic structural diagram of a computer device provided by an exemplary embodiment of the present application is shown; Figure 2 A schematic diagram of the structure of a processor provided by an exemplary embodiment of the present application is shown; Figure 3 A schematic diagram showing a topological structure of at least two processing cores provided by an exemplary embodiment of the present application is shown; Figure 4 A schematic structural diagram of a processor provided by another exemplary embodiment of the present application is shown; Figure 5A schematic structural diagram of a processor provided by another exemplary embodiment of the present application is shown; Figure 6 A flowchart of a method for reading a constant provided by an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION

[0014] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0015] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0016] The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0017] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the information such as the settings operations involved in this application is obtained with full authorization.

[0018] It should be understood that although the terms first, second, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, a first parameter may also be referred to as a second parameter, and similarly, a second parameter may also be referred to as a first parameter without departing from the scope of this disclosure. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0019] First, the relevant terms involved in this application are introduced.

[0020] Cache: A small, high-speed memory, part of a storage system, that stores instructions and data frequently used by programs. It can also be called cache memory. Processors typically include multiple levels of cache, each with a different capacity. Smaller caches typically have higher read and write efficiency because they need to process less data.

[0021] The cache is divided into multiple levels, such as the first level (L1), second level (L2), and third level (L3). The cache capacity of each level gradually increases, while the speed gradually decreases. The last level cache (LLC) is the last level cache of the CPU core, such as the L3 cache. When the processing core needs to access data, if the data does not hit in the L1 cache or L2 cache, it will continue to search the LLC. If the data still does not hit in the LLC, the processing core needs to read the required data from the main memory.

[0022] Caches can be divided into instruction caches and data caches based on the different information stored therein. Instruction caches are caches used to store instructions, and data caches are caches used to store data. The present embodiment of the application mainly uses data caches as an example, but the present embodiment of the application can also be used to support other caches that store data, and the present embodiment of the application is not limited to this.

[0023] Network-on-Chip (NoC): A network-based communication subsystem within an integrated circuit. It is typically used for data transmission and communication between different modules in a system-on-chip (SoC). NoC technology connects processor cores, memory, and various peripherals through routers, forming a highly parallel communication architecture that effectively improves data transmission efficiency and communication bandwidth. Compared to traditional shared bus approaches, NoC technology offers greater scalability and performance, making it particularly suitable for multi-core systems.

[0024] An instruction is a command that instructs a computer to perform a certain operation and is the smallest functional unit of computer operation. An instruction is a statement in machine language, or a set of meaningful binary codes. The collection of all instructions for a computer constitutes its instruction set, also known as its instruction set.

[0025] Instruction format: A human-readable representation of an instruction. An instruction typically consists of an opcode and operands. The opcode describes the type of operation to be performed, while the operands provide the data or addresses required to execute the instruction. While the opcode is essential, an instruction may contain no operands, one operand, or two operands.

[0026] In multi-core processors, constant data (such as instruction data, program constants, and program control flow) is frequently read by various components within the processing core. This data is often closely tied to the program's control flow and is latency-sensitive. Therefore, in modern processor cache hierarchies, each processing core is often equipped with a separate constant cache to reduce the latency of reading constant data.

[0027] However, constant data is actually shared across multiple processing cores. Each processing core independently caches the same constant data, resulting in storage redundancy and reducing overall cache utilization. Furthermore, the last-level cache shared by all cores is often located behind the on-chip network. When a cache miss occurs in a processing core, the missed constant data must be read from the last-level cache or main memory, resulting in high transmission and read latency.

[0028] For example, Figure 1 As shown, at least two processing cores in the processor 100 are connected via an on-chip network 200, and the processor 100 is connected to the last-level cache 300 and the main memory 400 via the on-chip network 200. Each of the at least two processing cores includes a cache, which can be a general cache or a constant cache. If the cache is a constant cache, each processing core will usually also include a general cache, which is in the Figure 1 In the related art, if the first constant of the processing core 1 does not hit the cache of the processing core 1, that is, a cache miss occurs, the processing core 1 needs to request the first constant from the last level cache 300, and the transmission path 10 is as follows: Figure 1 As shown in Figure 1, the last-level cache 300 has a large capacity, so the read and write speed of the last-level cache 300 is slow. On the other hand, the last-level cache 300 is located at the back end of the on-chip network 200, and the transmission delay of the read request from the processing core 1 to the last-level cache 300 is relatively high.

[0029] In order to solve this problem, the processor shown in the embodiment of the present application is as follows Figure 2 As shown, it supports cache sharing between at least two processing cores to reduce the latency of constant data. The details are as follows.

[0030] Figure 2 FIG. 1 is a schematic diagram showing the structure of a processor provided by an exemplary embodiment of the present application. The processor 100 includes at least two processing cores.

[0031] Optionally, the processor 100 shown in the embodiments of the present application includes at least two processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 100 can be implemented in at least one hardware form of a DSP (Digital Signal Processing), an FPGA (Field Programmable Gate Array), or a PLA (Programmable Logic Array). The processor 100 can be a CPU (Central Processing Unit). In some embodiments, the processor 100 can be a GPU (Graphics Processing Unit). In other embodiments, the processor 100 can also be an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning. The embodiments of the present application are described using the example of the processor 100 being a GPU.

[0032] Optionally, each of the at least two processing cores may communicate with each other via a bus, or via a bus such as Figure 2 The on-chip network 200 shown here can achieve mutual communication, and mutual communication can also be achieved through traces on the processor circuit board, etc. It should be noted that the above only shows some examples of the connection between at least two processing cores, and does not constitute a limitation on the connection between at least two processing cores in the embodiments of the present application.

[0033] The first processing core of the at least two processing cores is configured to obtain the first constant cached in the second processing core when the first constant does not hit in the first cache.

[0034] The first cache is a cache of the first processing core.

[0035] Optionally, when the first constant does not hit in the first cache, the first processing core of the at least two processing cores requests the first constant from the second cache, where the second cache is the cache of the second processing core.

[0036] Optionally, the first cache and the second cache may be dedicated constant caches, i.e., the data cached in the first cache and the second cache is read-only data. In this case, the first cache may be referred to as a first constant cache, and the second cache may similarly be referred to as a second constant cache. Alternatively, the first cache and the second cache may be general caches, i.e., they are used to cache both constants and variables. In other words, the data cached in the first cache and the second cache may be both read-only data and data that supports write operations (i.e., writable data). For example, the first cache and the second cache may be the first-level caches in the first processing core and the second processing core, respectively.

[0037] Optionally, the first cache is a private cache of the first processing core, that is, the first cache is a cache exclusively used by the first processing core. The first cache is usually only accessible to the first processing core, or in other words, data in the first cache cannot be directly queried by other processing cores.

[0038] Optionally, for the first processing core among the at least two processing cores, when the first constant needs to be used, the first cache is first accessed. If the first constant hits in the first cache, no subsequent steps are required. If the first constant does not hit in the first cache, a second processing core is selected from the at least two processing cores, and a read request carrying the address of the first constant and the identifier of the first processing core is sent to the second processing core.

[0039] Optionally, the second processing core is one processing core or multiple processing cores, that is, the first processing core can identify one or more processing cores of the at least two processing cores as the second processing core.

[0040] Optionally, the second processing core accesses the second cache based on the read request sent by the first processing core; if the first constant hits the second cache, the first constant is sent to the first processing core according to the identifier of the first processing core indicated in the read request; if the first constant does not hit the second cache, miss information is returned to the first processing core according to the identifier of the first processing core indicated in the read request.

[0041] Optionally, the first constant is data frequently read by the at least two processing cores, or in other words, the at least two processing cores read the first constant at a frequency greater than a frequency threshold. That is, when the at least two processing cores read the constant at a frequency greater than the frequency threshold, the method provided in the embodiments of the present application can be used to improve the efficiency of the at least two processing cores in reading the constant.

[0042] It should be noted that the "requests" and "responses" shown in the embodiments of the present application, whether they are read requests, query requests, read responses, query responses, etc., usually exist in the form of instructions in the processor, but the embodiments of the present application are not limited to this.

[0043] In summary, the processor provided by the embodiment of the present application supports the first processing core to request the first constant from other processing cores when the first constant is not hit. Compared with the traditional method of requesting the first constant from the last-level cache or main memory, on the one hand, the read and write speed of the cache in the processing core is usually higher than the last-level cache and main memory, that is, the speed at which the second processing core searches for the first constant from the second cache is faster than the speed at which the last-level cache or main memory searches for the first constant, or in other words, the read latency of the second processing core is lower than the read latency of the last-level cache or main memory, which indirectly makes it faster to request the first constant through the second processing core, thereby improving the reading efficiency of the first processing core for the constant. On the other hand, due to the architectural problem of the processor, the transmission path (or transmission delay) between the first processing core and the last-level cache or main memory is often longer than the transmission path (or transmission delay) between the first processing core and the second processing core. For example, the transmission path between the first processing core and the last-level cache is longer than the transmission path (or transmission delay) between the first processing core and the second processing core. Figure 1 As shown in the path 10 in FIG, the transmission path between the first processing core and the second processing core is as follows Figure 2 As shown in path 20, the transmission delay of the first processing core obtaining the read request of the first constant from the second processing core and the second processing core returning the read response of the first constant to the first processing core is also lower than that of the traditional method, so that the speed of obtaining the first constant through the second processing core request is faster, and the reading efficiency of the constant by the first processing core is improved.

[0044] Furthermore, the processors shown in the embodiments of the present application implement cache sharing between at least two processing cores. By sharing the cache, each processor does not need to ensure that constants are cached in its own cache. Instead, it only needs to ensure that a processing core with the constant cache exists among at least two processing cores. For example, when a first processing core needs a first constant, it reads the first constant from a second processing core, thereby implementing private cache sharing between at least two processing cores. Because the private cache sharing mechanism is primarily for constants, which are read-only, cache inconsistencies are less likely to occur. This ensures the reliability of constant data while avoiding redundancy of the first constant in at least two processing cores, thereby improving cache utilization in at least two processing cores.

[0045] 1. Determination of the second processing core.

[0046] For the first processing core, it is preferred to request the first constant from the second processing core to improve reading efficiency. However, if the first constant is still not hit in the second processing core, the first processing core may need to select the second processing core again or use the traditional method to read the first constant from the last-level cache or main memory. Therefore, additional design is required for the selection of the second processing core. The embodiments of the present application provide two methods for determining the second processing core.

[0047] Determination method one: determining the adjacent processing core as the second processing core.

[0048] Determination method two: setting a shared directory unit, and having the shared directory unit determine the processing core having the first constant cached therein as the second processing core.

[0049] Next, the two methods for determining the second processing core will be introduced one by one. It should be understood that the order of introduction does not represent the superiority or inferiority of the determination methods.

[0050] Determination method one: determining the adjacent processing core as the second processing core.

[0051] Optionally, the first processing core, the processor, or the shared directory unit is configured to, if the first constant does not hit the first cache, determine all or part of the adjacent cores as the second processing core, where the adjacent cores are processing cores that are physically or logically adjacent to the first processing core, and request the first constant from the second processing core. In other words, the second processing core is all or part of the adjacent cores.

[0052] It should be noted that the unit that determines the adjacent processing core as the second processing core may be the first processing core, the processor, the shared directory unit, etc. shown above, or may be other control modules in the processor, other processing cores (e.g., processing cores other than the first processing core), etc. The embodiments of the present application do not limit the unit that determines the adjacent processing core as the second processing core.

[0053] Optionally, for each of the at least two processing cores, there is at least one adjacent core, wherein the adjacent core is a processing core that is physically or logically adjacent to the first processing core.

[0054] Optionally, physical proximity refers to the presence of a direct adjacent trace on the circuit board of the processor 100. That is, when the first processing core and the second processing core have a direct adjacent trace on the circuit board, the first processing core and the second processing core can be considered physically adjacent.

[0055] Alternatively, logical proximity refers to a proximity relationship indicated by a driver, firmware, or other means. For example, for a register corresponding to a neighboring core in a first processing core, a driver may indicate that one or more processing cores are neighboring cores of the first processing core. The one or more processing cores may be indicated by physical identification or logical identification.

[0056] Optionally, when the first constant does not hit the first cache, the first processing core determines some of the adjacent cores as the second processing core. The some of the adjacent cores are processing cores selected by the first processing core. The some of the adjacent cores may be randomly selected by the first processing core or may be selected by the first processing core according to a selection condition.

[0057] Optionally, the adjacent core is a processing core that is logically adjacent to the first processing core. The adjacent core of the first processing core supports dynamic adjustment of the processor based on the status of at least two processing cores. Specifically, the adjacent processing core is a processing core that meets a first condition. The first condition includes at least one of the following: a load below a first threshold; a core that is ranked in the top n positions in ascending order of load, where n is a positive integer; a distance from the first processing core below a second threshold; or a core that is ranked in the top k positions in ascending order of distance from the first processing core, where k is a positive integer.

[0058] Exemplarily, a processor or a shared directory unit is used to determine a processing core that meets a first condition as an adjacent core of the first processing core; wherein the first condition includes at least one of the following: a load lower than a first threshold; being in the top n positions in ascending order of load, where n is a positive integer; a distance from the first processing core is lower than a second threshold; being in the top k positions in ascending order of distance from the first processing core, where k is a positive integer.

[0059] Optionally, the number of neighboring cores supported and stored by the first processing core is a first number, and the number of processing cores that meet the first condition is a second number. If the first number is greater than or equal to the second number, all processing cores that meet the first condition are identified as neighboring cores. If the first number is less than the second number, some processing cores that meet the first condition are selected as neighboring cores, such as selecting the first number of processing cores in ascending order of load, or selecting the first number of processing cores in ascending order of distance, etc.

[0060] Exemplarily, the first threshold is set by a developer, or is set based on expert experience, or is dynamically determined based on the loads of the at least two processing cores. For example, the first threshold is an average value, a median value, etc. of the loads of the at least two processing cores. Assume there are five processing cores. During a first time period, the load of processing core 1 is 50%, the load of processing core 2 is 20%, the load of processing core 3 is 40%, the load of processing core 4 is 80%, and the load of processing core 5 is 20%. The first threshold is (50% + 20% + 40% + 80% + 20%) / 5 = 42%. The processor triggers neighboring core adjustment at a fixed interval. If the first processing core is processing core 3, and its original neighbors are processing cores 2 and 4, since the load of processing core 2 is 20%, which is lower than the first threshold of 42%, processing core 2 is still determined to be a neighbor of processing core 3. However, since the load of processing core 4 is 80%, which is higher than the first threshold, the processor selects another processing core from the at least two processing cores whose load is lower than the first threshold and closer to processing core 3 as another neighbor of processing core 3. For example, processing core 5 is determined to be another neighbor of processing core 3.

[0061] Optionally, the first processing core is configured to, when the first constant does not hit in the first cache, identify an adjacent core that meets the first condition as the second processing core.

[0062] For example, assume there are five processing cores. During a first time period, the load of processing core 1 is 50%, the load of processing core 2 is 20%, the load of processing core 3 is 40%, the load of processing core 4 is 80%, and the load of processing core 5 is 20%. In this case, the first threshold is (50% + 20% + 40% + 80% + 20%) / 5 = 42%. If the first processing core is processing core 3, and its corresponding adjacent cores are processing cores 2 and 4, since the load of processing core 2 is 20%, which is lower than the first threshold of 42%, processing core 2 can be determined as the second processing core. However, since the load of processing core 4 is 80%, which is higher than the first threshold of 42%, processing core 4 is not determined as the second processing core. After determining the second processing core, processing core 3 (i.e., the first processing core) requests the first constant from processing core 2 (i.e., the second processing core). After receiving a read request carrying the address of the first constant, processing core 2 searches for the first constant in the cache of processing core 2. If the query is successful, the query returns the first constant to processing core 3. If the query is unsuccessful, the query returns unsuccessful information to processing core 3. At this time, processing core 3 can use the following "confirmation method 2: setting a shared directory unit, and having the shared directory unit determine the processing core with the first constant cached as the second processing core" to reconfirm the second processing core, or directly request the first constant from the last-level cache or main memory. It can also use the above two methods in parallel, that is, use the following "confirmation method 2: setting a shared directory unit, and having the shared directory unit determine the processing core with the first constant cached as the second processing core" to reconfirm the second processing core, and then request the first constant from the last-level cache or main memory. Ultimately, the result returned first by the shared directory unit, the last-level cache, or the main memory shall prevail.

[0063] Exemplarily, the second threshold is set by a developer, or is set based on expert experience, or is dynamically determined based on the loads of the at least two processing cores. For example, the second threshold is an average value, a median value, etc. of the loads of the at least two processing cores.

[0064] For example, the distance between a processing core and a first processing core is usually an integer. For example, if a processing core is directly adjacent to a first processing core, the distance between the first processing core and the first processing core can be considered to be 1. If there is one processing core between the first processing core and the first processing core, the distance between the first processing core and the first processing core can be considered to be 2. Figure 3 As shown, assuming that at least two processing cores are five processing cores, and a ring structure is used between the five processing cores, for processing core 1, the processing cores at a distance of 1 from processing core 1 include processing core 2 and processing core 5, and the processing cores at a distance of 2 from processing core 1 include processing core 3 and processing core 4.

[0065] In summary, the processor provided in the embodiment of the present application illustrates a method in which a first processing core autonomously determines an adjacent core as a second processing core. With this method, the first processing core, on the one hand, does not need to request a first constant from all of the at least two processing cores in the processor except the first processing core. Instead, it selects all or some of the adjacent cores of the first processing core to request the first constant, thus avoiding congestion of read requests between processing cores.

[0066] On the other hand, the transmission delay between adjacent cores, especially physically adjacent processing cores, and the first processing core is low. In particular, compared with the conventional method, the transmission delay between the first processing core and its adjacent processing cores is low, which can effectively improve the efficiency of the first processing core in reading the first constant.

[0067] Furthermore, the processor dynamically adjusts the neighboring cores of each processing core based on the first condition. If load is used to adjust neighboring cores, the first processing core can always request the first constant from the processing core with the lower load, thus avoiding increasing the processing burden on the processing core. If distance from the first processing core is used to determine neighboring cores, a more greedy criterion can be used, i.e., the first processing core always selects the processing core with the lowest transmission latency as the second processing core, thereby minimizing the transmission latency of the first processing core requesting the first constant.

[0068] Determination method two: setting a shared directory unit, and having the shared directory unit determine the processing core having the first constant cached therein as the second processing core.

[0069] Alternatively, as Figure 4 As shown, the processor 100 includes a shared directory unit 500, and the shared directory unit 500 is connected to at least two processing cores; or Figure 5 As shown, the shared directory unit 500 is independent of the processor 100 , and the processor 100 is connected to the shared directory unit 500 .

[0070] Optionally, the first processing core is configured to send a query request to the shared directory unit 500 when the first constant does not hit the first cache, the query request including the address of the first constant, and the shared directory unit 500 is configured to indicate the cache status of at least one constant in at least two processing cores. The shared directory unit 500 is configured to determine, based on the query request, that the processing core that caches the first constant is the second processing core. The shared directory unit 500 is configured to request the first constant from the second processing core, or to send the identifier of the second processing core to the first processing core so that the first processing core requests the first constant from the second processing core.

[0071] Optionally, the shared directory unit 500 is used to indicate the cache status of at least one constant in at least two processing cores; or the shared directory unit 500 stores the cache status of at least one constant in at least two processing cores. The at least one constant is a constant cached in the at least two processing cores; or in other words, the at least one constant is a constant stored in the caches of the at least two processing cores.

[0072] Exemplarily, the shared directory unit 500 is embodied as a directory structure, or a table structure. The shared directory unit 500 stores whether each constant of the at least one constant exists in the cache of each processing core of the at least two processing cores.

[0073] Optionally, the shared directory unit 500 includes at least one cache core bitmap, and the at least one cache core bitmap corresponds to at least one constant. The cache core bitmap corresponding to each constant in the at least one constant is used to indicate the processing core in the at least two processing cores that caches the constant.

[0074] The at least two processing cores are all or part of the processing cores in the processor.

[0075] Optionally, each cache core bitmap in the at least one cache core bitmap includes p bits, where p is a positive integer representing the number of processing cores in the at least two processing cores. For example, if the at least two processing cores are five processing cores, then p is 5, meaning that each cache core bitmap in the at least one cache core bitmap includes five bits. Each bit in each cache core bitmap corresponds to each processing core in the at least two processing cores, meaning that one bit corresponds to one processing core.

[0076] Among them, the cache core bitmap is used to indicate the cache status of at least two processing cores for constants, which not only ensures the reliability of the second processing core determined by the shared directory unit 500, but also reduces the amount of data compared to the identification of the stored processing core, helps to improve the query efficiency of the shared directory unit 500, and indirectly improves the reading efficiency of the first processing core for the first constant.

[0077] Exemplarily, the at least two processing cores are five processing cores, and the cache core bitmap of the first constant includes five bits. Assuming that, from right to left, the first bit corresponds to processing core 1, the second bit corresponds to processing core 2, the third bit corresponds to processing core 3, the fourth bit corresponds to processing core 4, and the fifth bit corresponds to processing core 5, if the cache core bitmap of the first constant is 00110, it means that the first constant is cached in processing core 2 and processing core 3. In this case, shared directory unit 500 can identify at least one of processing core 2 and processing core 3 as the second processing core.

[0078] Optionally, when confirming the second processing core, the shared directory unit 500 may also refer to the method shown in the above-mentioned "Determination method one: determining the adjacent processing core as the second processing core" to filter all or part of the processing cores that have the first constant cached as the second processing core through the first condition. That is, based on the query request, the shared directory unit 500 determines that the processing core that has the first constant cached and meets the first condition is the second processing core. The first condition includes at least one of the following: the load is lower than the first threshold; it is ranked in the top n places in order of load, where n is a positive integer; the distance from the first processing core is lower than the second threshold; it is ranked in the top k places in order of distance from the first processing core, where k is a positive integer.

[0079] Optionally, the shared directory unit 500 also includes at least one status flag, each of which corresponds to at least one constant. Each of the at least one status flag is used to indicate whether the corresponding constant is in an exclusive state or a shared state. The exclusive state indicates that the constant is cached by a single processing core; the shared state indicates that the constant is cached by multiple processing cores. Alternatively, the exclusive state indicates that the constant exists in the cache of a single processing core; the shared state indicates that the constant exists in the caches of multiple processing cores. That is, the configuration of the shared directory unit 500 follows the cache coherence protocol, setting a status flag for each constant to indicate whether the constant is exclusively used by a single processing core or shared by multiple processing cores. It should be understood that the cache coherence protocol is configured to optimize data access efficiency and data consistency across multiple processing cores. However, constants are inherently read-only (i.e., cannot be modified or globally shared). Therefore, setting a status flag for a constant is optional. After the status flag is set, the shared directory unit 500 can quickly know the status of the constant in at least two processing cores, so that the shared directory unit 500 can implement global adjustments to the constant based on the status flag. For example, the shared directory unit 500 determines whether the constant needs to be cached in other processing cores based on the status change of the constant and the change in the access frequency of the processing core to the constant.

[0080] For example, processing core 1 prefetches hot data into its cache based on the hot data prefetching principle. Hot data refers to frequently accessed data or data expected to be frequently accessed. Specifically, hot data is first data or data associated with first data. First data refers to data with an access frequency greater than an access threshold. The association with first data includes temporal association or spatial association. Specifically, prefetching hot data can be performed based on temporal locality or spatial locality. Temporal locality means that recently accessed data is likely to be used again, while spatial locality means that adjacent data is likely to be accessed consecutively. In other words, data temporally associated with first data refers to data that was the first data a fixed amount of time ago; data spatially associated with first data refers to data that is adjacent to or within a fixed address interval of the first data. For example, if processing core 1 frequently accesses constant 1, and constant 1 is adjacent to constant 2, processing core 1 may consider constant 2 to be hot data and therefore prefetch constant 2 into its cache in advance. At this point, constant 2 is marked as exclusive in shared directory unit 500. If constant 2 is prefetched into the cache of processing core 1 and processing core 1 starts to frequently access constant 1, the shared directory unit 500 can determine that constant 2 is hot data, and at this time constant 2 can be prefetched into the cache of other processing cores in advance. If the shared directory unit 500 determines that there are two hot data, such as constant 3 and constant 4, and constant 3 is in an exclusive state and constant 4 is in a shared state, then when the shared directory unit 500 prefetches hot data for the processing core, since the hot data prefetch priority of the constant in the exclusive state is higher than the hot data prefetch priority of the shared state, constant 3 is prefetched into the processing core first. That is, the hot data prefetch priority of the constant in the exclusive state is higher than the hot data prefetch priority of the shared state; the hot data prefetch priority is used to indicate the priority of hot data to be prefetched into the cache of the processing core. Hot data refers to data that is frequently accessed by the processing core or data that is expected to be frequently accessed by the processing core. Frequent access means that the access frequency is greater than the access threshold. By prioritizing prefetching data in an exclusive state, the prefetching accuracy can be improved, thereby utilizing the cache more efficiently. This means that the cache can prioritize reserving space for "hot and exclusive" data, thus achieving targeted and more efficient data prefetching and improving overall system performance.

[0081] Optionally, the access frequency is preset, such as automatically restoring the value after the processor is powered on; or, the access frequency is set by the user, such as through an interface opened by the processor.

[0082] Optionally, the shared directory unit 500 is configured to send a miss message to the first processing core if no cache core having the first constant is found. The first processing core is configured to request the first constant from the last-level cache or main memory; the last-level cache is the last level cache in the processor.

[0083] The main memory can also be called internal memory, primary storage, etc. The main memory is the main storage level in the processor that is directly accessed by the processor and is used to store instructions and data executed by the processor.

[0084] Optionally, the miss information is used to indicate that there is no processing core having the first constant cached in the at least two processing cores.

[0085] Optionally, after confirming the request for the first constant from the last-level cache or main memory, or after storing the first constant requested from the last-level cache or main memory in the first cache, the first processing core sends cache update information to the shared directory unit 500. The cache update information indicates that the first constant is cached in the first cache. The shared directory unit 500 then needs to update the cache status in the shared directory unit 500 based on the cache update information. For details, please refer to "2. Maintenance of Shared Directory Unit" below.

[0086] Optionally, when the shared directory unit 500 does not find a processing core that caches the first constant and meets the first condition, it will confirm the processing core that caches the first constant but does not meet the first condition as the second processing core. That is, when there is no processing core that caches the first constant and meets the first condition, the shared directory unit 500 reconfirms the second processing core based on the cached first constant as the basic condition. The first condition is a condition related to at least one of the load of the processing core and the distance between the processing core and the first processing core. The first condition is an additional condition for the shared directory unit 500 to determine the second processing core. If there is no processing core that meets both the basic condition and the additional condition, the shared directory unit can lower the screening condition for the second processing core and only needs to meet the basic condition, thereby ensuring that the first processing core can obtain the first constant within the shortest possible delay, thereby improving the reading efficiency of the constant. Especially when the first condition is related to the distance to the first processing core, although it is impossible to guarantee that the first constant will achieve the shortest first read latency when reading, compared with the second read latency of reading from the last-level cache or main memory in the traditional way, the first read latency is still much lower than the second read latency. That is, although giving up the screening of additional conditions cannot achieve the best reading efficiency, it still improves the reading efficiency compared with the traditional method.

[0087] Optionally, if no cache core is found that caches the first constant and meets the first condition, the shared directory unit 500 sends a miss message to the first processing core. That is, if no processing core both caches the first constant and meets the first condition, the shared directory unit 500 directly sends the miss message to the first processing core, informing the first processing core that it should request the first constant from the final-level cache or main memory. In scenarios where the first condition is load-related, if the loads of the processing cores that cache the first constant are all high, to avoid further load increases caused by the first processing core requesting the first constant from these other processing cores, the first processing core is directly notified of the miss message, causing it to request the first constant from the final-level cache or main memory. This not only reduces the load pressure caused by the first processing core's read request, ensuring normal operation of the processing cores that cache the first constant but do not meet the first condition, but also prevents the first processing core from having to wait for these other processing cores to meet the first condition, thereby pausing its operation. This ensures the efficiency of at least two processing cores.

[0088] The first condition is related to at least one of the load of the processing core and the distance between the processing core and the first processing core. The first condition is an additional condition for the shared directory unit 500 to determine the second processing core. If no processing core meets both the basic condition and the additional condition, the shared directory unit can reduce the screening condition for the second processing core to only meet the basic condition, thereby ensuring that the first processing core can obtain the first constant with the shortest possible latency, thereby improving the efficiency of reading the constant.

[0089] Optionally, the miss information is used to indicate that there is no processing core among the at least two processing cores that has the first constant cached and satisfies the first condition.

[0090] In summary, the processor provided in the embodiment of the present application shows two ways to set up the shared directory unit, namely setting it inside the processor or setting it outside the processor. When the shared directory unit is set outside the processor, it also supports other processors or components to use the shared directory unit.

[0091] On the other hand, a shared directory unit for determining the second processing core is shown. The shared directory unit is used to indicate the cache status of at least one constant in at least two processing cores. That is, the second processing core cache determined by the shared directory unit has the first constant, which improves the hit rate of the read request sent by the first processing core to the second processing core. In particular, compared with the above-mentioned "determination method one: determining the adjacent processing core as the second processing core", although the delay of the first processing core querying the shared directory unit for the second processing core and the delay of the shared directory unit determining the second processing core are increased, due to the connection relationship between the shared directory unit and the at least two processing cores, these increased delays are generally still shorter than the reading delay and transmission delay of the first processing core directly from the last-level cache or main memory, which can still improve the reading efficiency of the first processing core for the first constant and also ensure the hit rate of the first processing core for the first constant.

[0092] In addition, through the shared directory unit, the cache status of constants in at least two processing cores within the processor is integrated, thereby enabling at least two processing cores to share a private cache, reducing cache redundancy in at least two processing cores in the processor, and improving cache utilization of at least two processing cores.

[0093] After the shared directory unit 500 identifies the second processing core, there are two ways to request the first constant from the second processing core on behalf of the first processing core. One is for the first processing core to send a read request to the second processing core on its own, and the other is for the shared directory unit 500 to help the first processing core send a read request to the second processing core. The details are as follows.

[0094] (1) The first processing core sends a read request to the second processing core.

[0095] Optionally, the shared directory unit 500 is configured to send a query response to the first processing core, where the query response includes an identifier of the second processing core.

[0096] Optionally, the first processing core requests the first constant from the second processing core based on the identifier of the second processing core in the query response. Alternatively, the first processing core sends a read request to the second processing core based on the identifier of the second processing core in the query response, the read request carrying the address of the first constant.

[0097] Optionally, the query response also includes the address of the first constant, or the query response includes identification information, and the identification information is used to indicate that the current query response is a response of the first processing core to the query request of the first constant. In one implementation, the address of the first constant is used to confirm the correspondence between the query request and the query response, and the query response includes the address of the first constant. In another implementation, the query request and the query response respectively carry the same or corresponding identification information, and the identification information can be pre-set by the developer or generated according to preset rules, and the embodiments of the present application are not limited to this.

[0098] Optionally, the second processing core searches for the first constant in the second cache based on the address of the first constant, and returns the first constant to the first processing core in a read response.

[0099] To sum up, in the processor provided by the embodiment of the present application, the shared directory unit sends a query response to the first processing core, so that the first processing core automatically requests the first constant from the second processing core, thereby ensuring the functional unity of the shared directory unit and improving the maintainability of the shared directory unit.

[0100] (2) The shared directory unit sends a read request to the second processing core.

[0101] Optionally, the shared directory unit 500 is configured to send a read request to the second processing core, where the read request includes an address of the first constant, and the target address of the read request is the first cache.

[0102] Optionally, the target address of the read request is the processing core that requested the first constant. That is, the target address of the read request can be the first cache, or the identifier of the first processing core, etc. After receiving the read request, the second processing core reads the first constant from the second cache according to the address of the first constant and returns it to the first cache or the first processing core according to the target address of the read request.

[0103] Optionally, in addition to sending a read request to the second processing core, the shared directory unit 500 also sends a query response to the first processing core. The query response is used to indicate that there is a processing core that has cached the first constant among the at least two processing cores, or the query response is used to indicate the second processing core that has cached the first constant. That is, the query response may carry an identifier of the second processing core, as shown in "(1) The first processing core sends a read request to the second processing core" above, or may not carry an identifier related to the second processing core.

[0104] After receiving the query response, the first processing core may continue to send a read request for the first constant to the second processing core, as described in "(1) First processing core sends a read request to second processing core." The second processing core may execute a read response to at least one of the first processing core and the shared directory unit 500.

[0105] To sum up, the processor provided in the embodiment of the present application assists the first processing core in requesting the first constant from the second processing core after the shared directory unit determines the second processing core. There is no need for the shared directory unit to return a query response to the first processing core, and then the first processing core sends a read request to the second processing core. This shortens the transmission path of the first processing core requesting the first constant, reduces the transmission delay, and further improves the reading efficiency of the first constant.

[0106] In some embodiments, the second processing cores determined based on "Determination Method 2: Setting a shared directory unit, and having the shared directory unit determine the processing core that caches the first constant as the second processing core" include multiple processing cores. That is, there are multiple second processing cores. In this case, the first processing core or the shared directory unit is configured to request the first constant from a processing core among the multiple second processing cores that meets a second condition; the second condition includes at least one of the following: the lowest load; or the closest distance to the first processing core.

[0107] Because the second processing core determined by "Determination Method 2: Setting a shared directory unit, and having the shared directory unit determine the processing core that caches the first constant as the second processing core" is the processing core that caches the first constant, that is, if the second processing core is determined based on "Determination Method 2: Setting a shared directory unit, and having the shared directory unit determine the processing core that caches the first constant as the second processing core," there is no need to request the first constant from all second processing cores, as in "Determination Method 1: Determining an adjacent processing core as the second processing core." Instead, only one second processing core needs to be selected. This avoids increasing the load on the entire processor while ensuring that the first processing core can quickly read the first constant.

[0108] It should be noted that the above-mentioned "Determination Method 1: Determining the Adjacent Processing Core as the Second Processing Core" and "Determination Method 2: Setting a Shared Directory Unit, Having the Shared Directory Unit Determine the Processing Core That Has the First Constant Cached as the Second Processing Core" can be implemented as independent embodiments or as a combined embodiment. For example, the first processing core first executes "Determination Method 1: Determining the Adjacent Core as the Second Processing Core," determines the adjacent core as the second processing core, and requests the first constant from the second processing core. If the first constant is not cached in any of the adjacent cores of the first processing core, that is, the first constant is not cached in the second processing core, the first processing core then executes "Determination Method 2: Setting a Shared Directory Unit, Having the Shared Directory Unit Determine the Processing Core That Has the First Constant Cached as the Second Processing Core," queries the shared directory unit for a processing core that has the first constant cached in at least two processing cores, confirms it as the second processing core, and then requests the first constant from the second processing core. If the shared directory unit returns a miss message, this indicates that there is no processing core that has the first constant cached in the at least two processing cores, or there is no processing core that has the first constant cached and meets the first condition. At this time, the first processing core should request the first constant from the last-level cache or main memory. That is, the first processing core is configured to, if the first constant does not hit in the first cache, determine all or part of the adjacent cores as the second processing core, where the adjacent core refers to a processing core that is physically or logically adjacent to the first processing core; request the first constant from the second processing core; if the first constant does not hit in the second processing core, send a query request to the shared directory unit, the query request including the address of the first constant, the shared directory unit being configured to indicate the cache status of at least one constant in at least two processing cores; the shared directory unit is configured to, based on the query request, determine that the processing core that has the first constant cached is the second processing core; the shared directory unit is configured to request the first constant from the second processing core, or send the identifier of the second processing core to the first processing core so that the first processing core requests the first constant from the second processing core.

[0109] In other embodiments, in addition to executing "Determination Method 2: Setting a shared directory unit, whereby the shared directory unit determines the processing core that caches the first constant as the second processing core," the first processing core may also simultaneously request the first constant from the last-level cache or main memory. Specifically, if the first constant does not find a hit in the first cache, the first processing core may send a query request to the shared directory unit and a read request to the last-level cache or main memory. The first constant may be determined based on the earlier (or earliest) arriving response of the query response or the read response. The query request includes the address of the first constant, the read request includes the address of the first constant, the shared directory unit indicates the cache status of at least one constant in at least two processing cores, the query response is the response of the shared directory unit to the query request, and the read response is the response of the last-level cache or main memory to the read request. The last-level cache is the last-level cache in the processor. The later-arriving response of the query response or the read response may be ignored, treated as an invalid response, or discarded. If the query response is earlier than the read response of the last-level cache or main memory, then wait for the read response of the second processing core according to the query response, or send a read request to the second processing core; and set the read response of the last-level cache or main memory to invalid (or discard the read response of the last-level cache or main memory). If the read response of the last-level cache or main memory is earlier than the query response, then set the query response to invalid (or discard the query response). This can avoid the excessive delay caused by the first processing core still having to request the first constant from the last-level cache or main memory due to the absence of the first constant in the shared directory unit, that is, shorten the maximum transmission delay of the method provided in the embodiment of the present application to the same as the traditional method as much as possible.

[0110] For the shared directory unit 500 , in order to ensure the timeliness of the cache status of at least one constant in the shared directory unit 500 , the shared directory unit 500 should be maintained in a timely manner.

[0111] 2. Maintenance of shared directory units.

[0112] Optionally, the shared directory unit 500 is configured to update a cache status of at least one constant in the shared directory unit when a cache in any processing core of the at least two processing cores is updated.

[0113] Exemplarily, the cache status of at least one constant in the shared directory unit 500 is represented by a directory entry. The format of the directory entry is: [constant block address, cache core bitmap, status flag]. The constant block address is the address of the constant. The address of the constant can be the address of the constant in the last-level cache or main memory, or the address in the processing core.

[0114] Next, several scenarios of shared directory units maintaining directory entries are shown.

[0115] Scenario 1: There is a directory entry corresponding to the first constant, and the first constant is newly added to the first cache. For example, the first processing core reads the first constant from the cache of the second processing core in the manner shown above. When the first constant is newly added to the first cache, the shared directory unit 500 is triggered to update the directory entry corresponding to the first constant. Assume that at least two processing cores are 5 processing cores, namely processing core 1 to processing core 5. If before the first cache is updated, the directory entry corresponding to the first constant is: [address of the first constant, 00110, shared status], wherein a bit in the cache core bitmap is 0, indicating that the processing core corresponding to the bit does not cache the first constant, and a bit in the cache core bitmap is 1, indicating that the processing core corresponding to the bit caches the first constant, that is, processing core 2 and processing core 3 cache the first constant, and processing core 1, processing core 4, and processing core 5 do not cache the first constant. If the first processing core is processing core 4 and the second processing core is processing core 3, after processing core 4 reads the first constant from processing core 3, it sends a cache update message to shared directory unit 500. The cache update message is used to instruct the first cache to add the first constant. The cache update message carries the address of the first constant. Based on the cache update message, shared directory unit 500 determines that the directory entry corresponding to the first constant should be updated. The updated directory entry corresponding to the first constant is: [address of the first constant, 01110, shared status].

[0116] Scenario 2: There is no directory entry corresponding to the first constant, and the first constant is added to the first cache, that is, the first processing core reads the first constant from the final cache or main memory. Assume that at least two processing cores are 5 processing cores, namely processing core 1 to processing core 5. If the first processing core is processing core 4, then after processing core 4 reads the first constant from the final cache or main memory, cache update information is sent to the shared directory unit 500. The cache update information is used to indicate that the first constant is added to the first cache, and the cache update information carries the address of the first constant. The shared directory unit 500 determines that there is no directory entry corresponding to the first constant based on the cache update information, then the shared directory unit 500 adds a directory entry corresponding to the first constant, and the newly added directory entry corresponding to the first constant is: [address of the first constant, 01000, exclusive state]. At this time, since the first constant is cached in only one processing core, the state of the first constant is exclusive.

[0117] Scenario three: The first cache clears or replaces the first constant, and does not delete the directory entry corresponding to the first constant. Assume that at least two processing cores are 5 processing cores, namely processing core 1 to processing core 5. If the first processing core is processing core 4, processing core 4 replaces the cache line corresponding to the first constant due to insufficient space in the first cache. At this time, the shared directory unit 500 should be notified to update the directory entry corresponding to the first constant, such as sending cache update information to the shared directory unit 500. The cache update information is used to instruct the first cache to clear or replace the first constant. The cache update information carries the address of the first constant. The shared directory unit 500 determines to update the directory entry corresponding to the first constant based on the cache update information. For example, the directory entry corresponding to the first constant before the update is: [address of the first constant, 01110, shared status], then the directory entry corresponding to the first constant after the update is: [address of the first constant, 00110, shared status].

[0118] Scenario 4: The first cache clears or replaces the first constant, deleting the directory entry corresponding to the first constant. Assume that at least two processing cores are five processing cores, namely processing core 1 through processing core 5. If the first processing core is processing core 4, and processing core 4 replaces the cache line corresponding to the first constant due to insufficient space in the first cache, the shared directory unit 500 should be notified to update the directory entry corresponding to the first constant. For example, cache update information should be sent to the shared directory unit 500. The cache update information is used to instruct the first cache to clear or replace the first constant, and the cache update information carries the address of the first constant. The shared directory unit 500 determines to update the directory entry corresponding to the first constant based on the cache update information. For example, the directory entry corresponding to the first constant before the update is: [address of the first constant, 01000, exclusive state]. Since the current state of the first constant is exclusive, and the processing core 4 indicates that the first constant has been cleared or replaced, that is, there is currently no processing core cache with the first constant, the directory entry corresponding to the first constant can be deleted to reduce the memory usage of invalid information in the shared directory unit 500, and avoid excessive invalid information causing a reduction in the query speed of the shared directory unit 500.

[0119] In summary, the processor provided by the embodiment of the present application illustrates the maintenance process of the shared directory unit. The cache status of at least one constant stored in the shared directory unit must ensure real-time performance, thereby ensuring the reliability of the second processing core determined by the shared directory unit. Therefore, when the cache of each processing core in at least two processing cores is updated, the shared directory unit should also update the cache status of at least one constant, thereby ensuring that the first processing core can determine a reliable second processing core based on the shared directory unit, thereby achieving rapid reading of the first constant.

[0120] It should be noted that after the first processing core reads the first constant from the second processing core or the last-level cache or the main memory, the first constant may be cached in the first cache, or the first constant may not be cached in the first cache. For example, the first processing core saves the first constant in a register, performs an operation based on the first constant stored in the register, and releases the first constant stored in the register after the operation is completed; or, the first processing core caches the first constant in the L1 cache, releases the first constant in the L1 cache after the operation is completed, or determines whether to release the first constant in the L1 cache based on the actual execution situation (such as releasing it according to the release policy after the L1 cache is full). In some embodiments, the first constant is not cached in the first cache, but the first constant is cached in the second processing core. Compared with the traditional method, the method provided in the embodiment of the present application supports the first processing core to read the first constant from the second processing core, so that the reading efficiency of the first constant is higher than that of the traditional method. In another embodiment, after the first constant is cached in the first cache, the first processing core can subsequently read the first constant directly from its corresponding first cache, without having to read the first constant from the second processing core, the last-level cache, or the main memory. This can greatly improve the efficiency of the first processing core in reading the first constant and reduce the latency of reading the first constant. However, to ensure high read and write speeds, the capacity of the first cache of the first processing core is often small, so the first processing core needs to determine whether to cache the first constant in the first cache.

[0121] Exemplarily, the first processing core is configured to cache the first constant in the first cache when a third condition is met; wherein the third condition includes at least one of the following: an access frequency of the first constant is greater than a third threshold, the access frequency being used to indicate a frequency at which the first processing core requests the first constant from the third processing core; and an occupancy rate of the first cache is less than a fourth threshold.

[0122] Optionally, the third threshold and the fourth threshold may be set by a software developer through an interface, or may be set by the processor before leaving the factory.

[0123] Optionally, the first processing core caches the first constant in the first cache if the access frequency of the first constant is greater than a third threshold. If the first processing core frequently accesses the first constant in the third processing core, for example, the number of accesses within the first time period is greater than the third threshold, the first constant currently read from the third processing core is cached in the first cache, i.e., stored in a cache line of the first cache.

[0124] Optionally, the first processing core caches the first constant in the first cache when the occupancy of the first cache is less than a fourth threshold. That is, when the first cache is relatively idle, if the first constant is hit in the third processing core, the first constant is also cached in the first processing core when the first constant is read into the first processing core.

[0125] Optionally, the first processing core caches the first constant in the first cache if the occupancy of the first cache is greater than a fifth threshold and the access frequency of the first constant is greater than a third threshold. That is, if the capacity of the first cache is relatively limited, only the first constant with a higher access frequency is cached in the first cache.

[0126] Optionally, when the first processing core reads the first constant from the final cache or main memory, or in other words, when the shared directory unit 500 does not include a cache of the first constant, that is, when there is no processing core with the first constant cached among the at least two processing cores, the first processing core caches the first constant in the first cache. Furthermore, a cache of the first constant is newly added in the shared directory unit 500. That is, if there is no processing core with the first constant cached among the at least two processing cores, the above-mentioned access frequency and occupancy rate restrictions can be ignored, and the first constant can be directly cached in the first cache. First, ensure that there is a processing core with the first constant cached among the at least two processing cores, so that other processing cores do not need to read from the final cache or main memory when they need the first constant.

[0127] In some embodiments, the shared directory unit 500 also supports pre-fetching constants into some cache cores, as shown below.

[0128] In some embodiments, the shared directory unit 500 is used to send a prefetch instruction to at least one fourth processing core when a second constant among the at least one constant satisfies a fourth condition, and the prefetch instruction is used to cache the second constant in the cache of the at least one fourth processing core, and the distance between each fourth processing core among the at least one fourth processing core and each fifth processing core among the at least two fifth processing cores is less than a fifth threshold, and the at least two fifth processing cores are processing cores that access the second constant.

[0129] The fourth condition includes at least one of the following: the number of at least two fifth processing cores is greater than the first number; and the access frequency of the second constant is greater than a sixth threshold.

[0130] Optionally, the access frequency of the second constant is an average, a sum, a median, etc. of the access frequencies of the at least two fifth processing cores to the second constant.

[0131] Exemplarily, when there are a large number of fifth processing cores frequently accessing the second constant, the shared directory unit 500 can pre-fetch the second constant to the fourth processing core in advance, and the distance between the fourth processing core and the fifth processing core is less than the fifth threshold. For example, assuming that at least two processing cores are five processing cores, namely processing core 1 to processing core 5, if the shared directory unit 500 detects that processing core 2 and processing core 4 will frequently access the second constant, the shared directory unit 500 selects some processing cores from the at least two processing cores as the first processing core, such as processing core 3 as the fourth processing core, and caches the second constant in the cache of processing core 3. At this time, the distance between processing core 3 and processing core 2 is 1, and the distance between processing core 3 and processing core 4 is also 1, that is, the fifth threshold is 2. Of course, the fifth threshold can also be set to 1, that is, processing core 2 and processing core 4 are regarded as the fourth processing core.

[0132] Optionally, at least one fourth processing core satisfies a fifth condition, which includes at least one of the following: a distance from each of the at least one fifth processing core being less than a fifth threshold; and a load being less than a sixth threshold. That is, the fourth processing core may also prefetch the second constant to a core with a lower load based on load, thereby ensuring that at least two processing cores can read the second constant from the fourth processing core while also ensuring that the fourth processing core is not overloaded.

[0133] It should be understood that the fifth processing core here can be understood as the first processing core mentioned above, and the fourth processing core can be understood as the second processing core mentioned above.

[0134] In summary, the processor provided by the embodiment of the present application supports a shared directory unit that pre-fetches the second constant to at least one fourth processing core in advance. The fourth processing core can be a processing core with a high-frequency access demand for the second constant, or a neighboring core of a processing core with a high-frequency access demand, or a processing core within a fifth threshold distance. On the one hand, the second constant is a known shared constant. Generally speaking, there is a higher probability of sharing constants between neighboring cores. Therefore, when it is confirmed that the fifth processing core has a high access demand, the second constant can be pre-loaded to its neighboring core (i.e., at least one fourth processing core). If the fifth processing core does not cache the second constant, it can ensure that the fifth processing core can quickly obtain the second constant from the fourth processing core; if the fifth processing core has the second constant cached, it can ensure that the fourth processing core can quickly read the second constant pre-fetched into the cache of the fourth processing core when it needs the second constant.

[0135] Specifically, prefetching the second constant to the fourth processing core is intended to ensure that multiple fifth processing cores can quickly obtain the second constant from adjacent fourth processing cores. If the multiple fifth processing cores are closely spaced, forming a core cluster, they are equivalent to sharing a single fourth processing core. In this case, pre-storing a copy of the second constant in the fourth processing core allows all fifth processing cores to promptly obtain the second constant from the fourth processing core, reducing cache usage while ensuring read speed. If the multiple fifth processing cores are farther apart, forming multiple core clusters, the multiple fifth processing cores within each core cluster share a single fourth processing core, with the fifth processing cores obtaining the second constant from the fourth processing core within their core cluster. The fourth processing core can be selected so that the sum of routing paths from the fourth processing core to each fifth processing core within each core cluster is minimized. In other words, under normal circumstances, no cache line is occupied within the fifth processing core to store the second constant. However, in some special cases, to improve the processing efficiency of the fifth processing core, the second constant can be cached in some processing cores within the core cluster (e.g., the fifth core). This is especially true when the access frequency of the second constant from a particular fifth processing core exceeds a predetermined threshold.

[0136] In related technologies, each core independently caches the same constant data, resulting in storage redundancy and reduced overall cache utilization. The last-level cache shared by all cores is often located behind the on-chip network, resulting in high latency.

[0137] Therefore, the embodiments of the present application illustrate a processor related to the field of computer architecture, particularly a shared constant cache coherence protocol in multi-core / many-core processors (such as CPUs and GPUs). The protocol aims to reduce access latency during cache misses and improve cache utilization by optimizing the sharing mechanism of constant data between multiple cores. The embodiments of the present application improve the overall utilization of the constant cache through a core constant cache coherence interconnect network.

[0138] Specifically, when a processing core experiences a constant cache miss, it prioritizes requesting data from physically / logically adjacent processing cores. If none of the adjacent cores miss, it queries the cache status of other non-adjacent cores through the globally shared shared directory unit. If there is no cache record in the shared directory unit, it finally accesses the downstream last-level cache or main memory. The details are shown below.

[0139] 1. Shared directory.

[0140] All processing cores share a shared directory unit, which includes a centralized directory that records the cache distribution of each constant block.

[0141] The format of directory entries in a shared directory unit is: [constant block address, cache core bitmap, status flags].

[0142] The cache core bitmap uses a bit mask to indicate which cores have the constant block stored in their private constant cache (for example, an 8-core system uses 8 bits, with each bit representing one core).

[0143] Status flag: marks the sharing status of the constant block (such as "exclusive" for hot data prefetching optimization).

[0144] 2. Cache miss handling process.

[0145] When processing core A accesses constant data, the local core (i.e., processing core A) misses the private constant cache. It requests the data from the neighboring core. If the neighboring core's private constant cache hits, the following steps are not continued.

[0146] If the private constant cache of the neighboring core misses, a query request containing the target address (i.e., the constant block address of the constant data) is sent to the centralized directory.

[0147] After receiving the request, the shared directory unit checks the directory entry.

[0148] If the shared directory unit hits, the shared directory unit can return the target cache (such as the private constant cache of core B) containing the target address to the private constant cache of processing core A. The private constant cache of processing core A initiates a read request to the private constant cache of processing core B.

[0149] In other embodiments, the shared directory unit may simultaneously send a read request corresponding to the constant block address to the private constant cache of processing core B, and mark the destination of the read request as the private constant cache of processing core A.

[0150] If the shared directory unit misses, the miss information is returned to the processing core A, and the processing core A normally initiates a read request to the last-level cache.

[0151] In other embodiments, processing core A may simultaneously initiate read requests to the shared directory unit and the last-level cache, and based on the cache hit status of adjacent cores, take the first returned result as the result and discard the later returned result.

[0152] 3. Data cache and directory update.

[0153] Whenever the content of the private constant cache of the processing core changes, the directory entry is updated, which includes the following two situations.

[0154] New cache: Mark core A in the cache core bitmap of the directory entry.

[0155] Cache replacement: If core A replaces the constant block due to insufficient space, it notifies the directory to clear its own ID.

[0156] In addition, further optimization can be achieved through the following methods.

[0157] 4. Dynamic grouping of adjacent cores.

[0158] Dynamically adjust the range of neighboring cores based on physical topology (such as intra-chip NoC adjacency) or runtime access patterns (such as the same thread group in a GPU). This prioritizes data transfers to less-loaded cores or processing cores that are physically closer (such as within a GPU cluster), reducing latency.

[0159] 5. Dynamic application of access frequency.

[0160] The local cache decides whether to store data read back from the neighboring core's cache into the cache line based on the address access hit rate. If there are no repeated hits on the same address, the data in the neighboring core's cache will not be written to the local cache line if the cache is full.

[0161] 6. Combination with prefetching.

[0162] When a constant block is frequently accessed by multiple cores, the directory proactively triggers prefetch instructions to load it into the cache of neighboring cores in advance.

[0163] Figure 6 A flowchart of a method for reading a constant provided by an exemplary embodiment of the present application is shown. The method is executed by a processor including at least two processing cores. The method includes step 210.

[0164] Step 210 : When the first constant does not hit in the first cache, the first processing core of the at least two processing cores obtains the first constant cached in the second processing core.

[0165] The first cache is a cache of the first processing core.

[0166] Optionally, each of the at least two processing cores may communicate with each other via a bus, or via a bus such as Figure 2 The on-chip network shown here can achieve mutual communication, and mutual communication can also be achieved through traces on the processor circuit board, etc. It should be noted that the above only shows some examples of the connection between at least two processing cores, and does not constitute a limitation on the connection between at least two processing cores in the embodiments of the present application.

[0167] Optionally, when the first constant does not hit in the first cache, the first processing core of the at least two processing cores requests the first constant from the second cache, where the second cache is the cache of the second processing core.

[0168] Optionally, the first cache and the second cache may be dedicated constant caches, i.e., the data cached in the first cache and the second cache is read-only data. In this case, the first cache may be referred to as a first constant cache, and the second cache may similarly be referred to as a second constant cache. Alternatively, the first cache and the second cache may be general caches, i.e., they are used to cache both constants and variables. In other words, the data cached in the first cache and the second cache may be both read-only data and data that supports write operations (i.e., writable data). For example, the first cache and the second cache may be the first-level caches in the first processing core and the second processing core, respectively.

[0169] Optionally, the first cache is a private cache of the first processing core, that is, the first cache is a cache exclusively used by the first processing core. The first cache is usually only accessible to the first processing core, or in other words, data in the first cache cannot be directly queried by other processing cores.

[0170] Optionally, for the first processing core among the at least two processing cores, when the first constant needs to be used, the first cache is first accessed. If the first constant hits in the first cache, no subsequent steps are required. If the first constant does not hit in the first cache, a second processing core is selected from the at least two processing cores, and a read request carrying the address of the first constant and the identifier of the first processing core is sent to the second processing core.

[0171] Optionally, the second processing core is one processing core or multiple processing cores, that is, the first processing core can identify one or more processing cores of the at least two processing cores as the second processing core.

[0172] Optionally, the second processing core accesses the second cache based on the read request sent by the first processing core; if the first constant hits the second cache, the first constant is sent to the first processing core according to the identifier of the first processing core indicated in the read request; if the first constant does not hit the second cache, miss information is returned to the first processing core according to the identifier of the first processing core indicated in the read request.

[0173] Optionally, the first constant is data frequently read by the at least two processing cores, or in other words, the at least two processing cores read the first constant at a frequency greater than a frequency threshold. That is, when the at least two processing cores read the constant at a frequency greater than the frequency threshold, the method provided in the embodiments of the present application can be used to improve the efficiency of the at least two processing cores in reading the constant.

[0174] It should be noted that the "requests" and "responses" shown in the embodiments of the present application, whether they are read requests, query requests, read responses, query responses, etc., usually exist in the form of instructions in the processor, but the embodiments of the present application are not limited to this.

[0175] In summary, the processor provided by the embodiment of the present application supports the first processing core to request the first constant from other processing cores when the first constant is not hit. Compared with the traditional method of requesting the first constant from the last-level cache or main memory, on the one hand, the read and write speed of the cache in the processing core is usually higher than the last-level cache and main memory, that is, the speed at which the second processing core searches for the first constant from the second cache is faster than the speed at which the last-level cache or main memory searches for the first constant, or in other words, the read latency of the second processing core is lower than the read latency of the last-level cache or main memory, which indirectly makes it faster to request the first constant through the second processing core, thereby improving the reading efficiency of the first processing core for the constant. On the other hand, due to the architectural problem of the processor, the transmission path (or transmission delay) between the first processing core and the last-level cache or main memory is often longer than the transmission path (or transmission delay) between the first processing core and the second processing core. For example, the transmission path between the first processing core and the last-level cache is longer than the transmission path (or transmission delay) between the first processing core and the second processing core. Figure 1 As shown in the path 10 in FIG, the transmission path between the first processing core and the second processing core is as follows Figure 2 As shown in path 20, the transmission delay of the first processing core obtaining the read request of the first constant from the second processing core and the second processing core returning the read response of the first constant to the first processing core is also lower than that of the traditional method, so that the speed of obtaining the first constant through the second processing core request is faster, and the reading efficiency of the constant by the first processing core is improved.

[0176] Furthermore, the processors shown in the embodiments of the present application implement cache sharing between at least two processing cores. By sharing the cache, each processor does not need to ensure that constants are cached in its own cache. Instead, it only needs to ensure that a processing core with the constant cache exists among at least two processing cores. For example, when a first processing core needs a first constant, it reads the first constant from a second processing core, thereby implementing private cache sharing between at least two processing cores. Because the private cache sharing mechanism is primarily for constants, which are read-only, cache inconsistencies are less likely to occur. This ensures the reliability of constant data while avoiding redundancy of the first constant in at least two processing cores, thereby improving cache utilization in at least two processing cores.

[0177] 1. Determination of the second processing core.

[0178] For the first processing core, it is preferred to request the first constant from the second processing core to improve reading efficiency. However, if the first constant is still not hit in the second processing core, the first processing core may need to select the second processing core again or use the traditional method to read the first constant from the last-level cache or main memory. Therefore, additional design is required for the selection of the second processing core. The embodiments of the present application provide two methods for determining the second processing core.

[0179] Determination method one: determining the adjacent processing core as the second processing core.

[0180] Determination method two: setting a shared directory unit, and having the shared directory unit determine the processing core having the first constant cached therein as the second processing core.

[0181] Next, the two methods for determining the second processing core will be introduced one by one. It should be understood that the order of introduction does not represent the superiority or inferiority of the determination methods.

[0182] Determination method one: determining the adjacent processing core as the second processing core.

[0183] Optionally, step 210 may be implemented as follows: if the first constant does not hit the first cache, the first processing core, the processor, or the shared directory unit may determine all or part of the adjacent cores as second processing cores, where the adjacent cores are processing cores that are physically or logically adjacent to the first processing core; and request the first constant from the second processing core, or send an identifier of the second processing core to the first processing core so that the first processing core requests the first constant from the second processing core. In other words, the second processing cores are all or part of the adjacent cores.

[0184] It should be noted that the unit that determines the adjacent processing core as the second processing core may be the first processing core, the processor, the shared directory unit, etc. shown above, or may be other control modules in the processor, other processing cores (e.g., processing cores other than the first processing core), etc. The embodiments of the present application do not limit the unit that determines the adjacent processing core as the second processing core.

[0185] Optionally, for each of the at least two processing cores, there is at least one adjacent core, wherein the adjacent core is a processing core that is physically or logically adjacent to the first processing core.

[0186] Optionally, physical proximity refers to the presence of a direct-connected trace on a circuit board of the processor. That is, when a direct-connected trace exists between the first processing core and the second processing core on the circuit board, the first processing core and the second processing core can be considered to be physically adjacent.

[0187] Alternatively, logical proximity refers to a proximity relationship indicated by a driver, firmware, or other means. For example, for a register corresponding to a neighboring core in a first processing core, a driver may indicate that one or more processing cores are neighboring cores of the first processing core. The one or more processing cores may be indicated by physical identification or logical identification.

[0188] Optionally, the adjacent core is a processing core that is logically adjacent to the first processing core. The adjacent core of the first processing core supports dynamic adjustment of the processor based on the status of at least two processing cores. Specifically, the adjacent processing core is a processing core that meets a first condition. The first condition includes at least one of the following: a load below a first threshold; a core that is ranked in the top n positions in ascending order of load, where n is a positive integer; a distance from the first processing core below a second threshold; or a core that is ranked in the top k positions in ascending order of distance from the first processing core, where k is a positive integer.

[0189] In some embodiments, the method further includes: the processor or shared directory unit determines the processing core that meets the first condition as an adjacent core of the first processing core; wherein the first condition includes at least one of the following: the load is lower than the first threshold; the core is ranked in the top n positions in ascending order of load, where n is a positive integer; the distance to the first processing core is lower than the second threshold; the core is ranked in the top k positions in ascending order of distance to the first processing core, where k is a positive integer.

[0190] In some embodiments, the above-mentioned "when the first constant does not hit the first cache, the first processing core determines all or part of the adjacent cores as the second processing core" can be implemented as follows: when the first constant does not hit the first cache, the first processing core determines the adjacent core that meets the first condition as the second processing core.

[0191] For specific details, please refer to the relevant content regarding “Determination method one: determining the adjacent processing core as the second processing core” in the processor shown above, and will not be repeated here.

[0192] Determination method two: setting a shared directory unit, and having the shared directory unit determine the processing core having the first constant cached therein as the second processing core.

[0193] Optionally, the processor includes a shared directory unit, and the shared directory unit is connected to at least two processing cores; or, the shared directory unit is independent of the processor, and the processor is connected to the shared directory unit.

[0194] In some embodiments, step 210 may be implemented as follows: if the first constant does not hit the first cache, the first processing core sends a query request to the shared directory unit, the query request including the address of the first constant, and the shared directory unit is used to indicate the cache status of at least one constant in at least two processing cores. Based on the query request, the shared directory unit determines that the processing core that caches the first constant is the second processing core. The shared directory unit is used to request the first constant from the second processing core, or to send the identifier of the second processing core to the first processing core so that the first processing core requests the first constant from the second processing core.

[0195] Optionally, the shared directory unit is used to indicate a cache status of at least one constant in at least two processing cores; or the shared directory unit stores a cache status of at least one constant in at least two processing cores.

[0196] Exemplarily, the shared directory unit is embodied as a directory structure, or a table structure, and stores information on whether each constant in at least one constant exists in a cache of each processing core in the at least two processing cores.

[0197] Optionally, the shared directory unit includes at least one cache core bitmap, the at least one cache core bitmap corresponding to the at least one constant, and the cache core bitmap corresponding to each of the at least one constant is used to indicate the processing core in the at least two processing cores that has the constant cached.

[0198] The at least two processing cores are all or part of the processing cores in the processor.

[0199] Optionally, when confirming the second processing core, the shared directory unit may also refer to the method shown in the above "Determination method one: determining the adjacent processing core as the second processing core" and filter all or part of the processing cores that have the first constant cached as the second processing core through the first condition. That is, based on the query request, the shared directory unit determines that the processing core that has the first constant cached and meets the first condition is the second processing core. The first condition includes at least one of the following: the load is lower than the first threshold; the load is ranked in the top n places from small to large, where n is a positive integer; the distance from the first processing core is lower than the second threshold; the distance from the first processing core is ranked in the top k places from small to large, where k is a positive integer.

[0200] Optionally, the shared directory unit also includes at least one status flag, each of which corresponds one-to-one to at least one constant. Each of the at least one status flag is used to indicate whether the corresponding constant is in an exclusive or shared state. The exclusive state indicates that the constant is cached by a single processing core; the shared state indicates that the constant is cached by multiple processing cores. Alternatively, the exclusive state indicates that the constant exists in the cache of a single processing core; the shared state indicates that the constant exists in the caches of multiple processing cores. That is, the configuration of the shared directory unit follows the cache coherence protocol, setting a status flag for each constant to indicate whether the constant is exclusively used by a single processing core or shared by multiple processing cores. It should be understood that the cache coherence protocol is designed to optimize data access efficiency and data consistency across multiple processing cores. However, constants are inherently read-only (i.e., cannot be modified or globally shared). Therefore, setting a status flag for a constant is optional. After the status flag is set, the shared directory unit can quickly know the status of the constant in at least two processing cores, so that the shared directory unit can implement global adjustments for the constant based on the status flag. For example, the shared directory unit determines whether the constant needs to be cached in other processing cores based on the status change of a constant and the change in the access frequency of the processing core to the constant. The hot data prefetch priority of the constant in the exclusive state is higher than the hot data prefetch priority of the constant in the shared state; the hot data prefetch priority is used to indicate the priority of hot data to be prefetched into the cache of the processing core. Hot data refers to data that is frequently accessed by the processing core or data that is expected to be frequently accessed by the processing core. By prioritizing the prefetching of data in the exclusive state, the prefetching accuracy can be improved, thereby making more efficient use of the cache, that is, the cache can prioritize reserving space for "hot and exclusive" data, that is, achieving targeted and more efficient data prefetching, thereby improving the overall performance of the system.

[0201] In some embodiments, the method further includes: the shared directory unit being configured to send a miss message to the first processing core if no cache core having the first constant is found in the query. The first processing core is configured to request the first constant from a last-level cache or main memory; the last-level cache being the last level cache in the processor.

[0202] The main memory can also be called internal memory, primary storage, etc. The main memory is the main storage level in the processor that is directly accessed by the processor and is used to store instructions and data executed by the processor.

[0203] Optionally, the miss information is used to indicate that there is no processing core having the first constant cached in the at least two processing cores.

[0204] Optionally, after confirming the request for the first constant from the final-level cache or main memory, or after storing the first constant requested from the final-level cache or main memory in the first cache, the first processing core sends cache update information to the shared directory unit. The cache update information is used to indicate that the first constant is cached in the first cache. The shared directory unit is then required to update the cache status in the shared directory unit based on the cache update information. For details, please refer to "2. Maintenance of Shared Directory Unit" below.

[0205] Optionally, when the shared directory unit does not find a processing core that caches the first constant and meets the first condition, the shared directory unit confirms the processing core that caches the first constant but does not meet the first condition as the second processing core. That is, when there is no processing core that caches the first constant and meets the first condition, the shared directory unit reconfirms the second processing core based on the cache of the first constant as the basic condition. The first condition is a condition related to at least one of the load of the processing core and the distance between the processing core and the first processing core. The first condition is an additional condition for the shared directory unit to determine the second processing core. If there is no processing core that meets both the basic condition and the additional condition, the shared directory unit can lower the screening condition for the second processing core and only needs to meet the basic condition, thereby ensuring that the first processing core can obtain the first constant within the shortest possible delay, thereby improving the reading efficiency of the constant. Especially when the first condition is related to the distance to the first processing core, although it is impossible to guarantee that the first constant will achieve the shortest first read latency when reading, compared with the second read latency of reading from the last-level cache or main memory in the traditional way, the first read latency is still much lower than the second read latency. That is, although giving up the screening of additional conditions cannot achieve the best reading efficiency, it still improves the reading efficiency compared with the traditional method.

[0206] Optionally, if no cache core is found that caches the first constant and meets the first condition, the shared directory unit sends a miss message to the first processing core. That is, if no processing core both caches the first constant and meets the first condition, the shared directory unit directly sends a miss message to the first processing core, informing the first processing core that it should request the first constant from the final-level cache or main memory. In scenarios where the first condition is load-related, if the loads of the processing cores that cache the first constant are all high, to avoid further load increases caused by the first processing core requesting the first constant from these other processing cores, the first processing core is directly notified of the miss message, causing it to request the first constant from the final-level cache or main memory. This not only reduces the load pressure caused by the first processing core's read request, ensuring normal operation of the processing cores that cache the first constant but do not meet the first condition, but also prevents the first processing core from having to wait for these other processing cores to meet the first condition, thereby pausing its operation and thus ensuring the efficiency of at least two processing cores.

[0207] For specific details, please refer to the relevant content of "Determination method two: setting a shared directory unit, and having the shared directory unit determine the processing core that caches the first constant as the second processing core" in the processor shown above, which will not be repeated here.

[0208] After the shared directory unit identifies the second processing core, there are two ways to request the first constant from the second processing core on behalf of the first processing core. One is for the first processing core to send a read request to the second processing core on its own, and the other is for the shared directory unit to help the first processing core send a read request to the second processing core. The details are as follows.

[0209] (1) The first processing core sends a read request to the second processing core.

[0210] In some embodiments, the method further includes: the shared directory unit sending a query response to the first processing core, where the query response includes an identifier of the second processing core.

[0211] Optionally, the second processing core searches for the first constant in the second cache based on the address of the first constant, and returns the first constant to the first processing core in a read response.

[0212] For details, please refer to the relevant content regarding “(1) the first processing core sends a read request to the second processing core” in the processor shown above, which will not be repeated here.

[0213] (2) The shared directory unit sends a read request to the second processing core.

[0214] In some embodiments, the method further includes: the shared directory unit sending a read request to the second processing core, the read request including the address of the first constant, and the target address of the read request being the first cache.

[0215] In some embodiments, the second processing cores determined based on "Determination Method 2: Setting a shared directory unit, and having the shared directory unit determine the processing core that caches the first constant as the second processing core" include multiple processing cores. That is, there are multiple second processing cores. In this case, "the first processing core or the shared directory unit requests the first constant from the second processing core" can be implemented as follows: the first processing core or the shared directory unit requests the first constant from a processing core among the multiple second processing cores that meets a second condition; the second condition includes at least one of the following: the lowest load; the closest distance to the first processing core.

[0216] It should be noted that the above-mentioned "Determination Method 1: Determining the Adjacent Processing Core as the Second Processing Core" and "Determination Method 2: Setting a Shared Directory Unit, Having the Shared Directory Unit Determine the Processing Core That Has the First Constant Cached as the Second Processing Core" can be implemented as independent embodiments or as a combined embodiment. For example, the first processing core first executes "Determination Method 1: Determining the Adjacent Core as the Second Processing Core," determines the adjacent core as the second processing core, and requests the first constant from the second processing core. If the first constant is not cached in any of the adjacent cores of the first processing core, that is, the first constant is not cached in the second processing core, the first processing core then executes "Determination Method 2: Setting a Shared Directory Unit, Having the Shared Directory Unit Determine the Processing Core That Has the First Constant Cached as the Second Processing Core," queries the shared directory unit for a processing core that has the first constant cached in at least two processing cores, confirms it as the second processing core, and then requests the first constant from the second processing core. If the shared directory unit returns a miss message, this indicates that there is no processing core that has the first constant cached in the at least two processing cores, or there is no processing core that has the first constant cached and meets the first condition. At this time, the first processing core should request the first constant from the last-level cache or main memory. That is, the first processing core is configured to, if the first constant does not hit in the first cache, determine all or part of the adjacent cores as the second processing core, where the adjacent core refers to a processing core that is physically or logically adjacent to the first processing core; request the first constant from the second processing core; if the first constant does not hit in the second processing core, send a query request to the shared directory unit, the query request including the address of the first constant, the shared directory unit being configured to indicate the cache status of at least one constant in at least two processing cores; the shared directory unit is configured to, based on the query request, determine that the processing core that has the first constant cached is the second processing core; and the first processing core or the shared directory unit is configured to request the first constant from the second processing core.

[0217] In other embodiments, in addition to executing "Determination Method 2: Setting a shared directory unit, whereby the shared directory unit determines the processing core that caches the first constant as the second processing core," the first processing core may also simultaneously request the first constant from the last-level cache or main memory. Specifically, if the first constant does not find a hit in the first cache, the first processing core may send a query request to the shared directory unit and a read request to the last-level cache or main memory. The first constant may be determined based on the earlier (or earliest) arriving response of the query response or the read response. The query request includes the address of the first constant, the read request includes the address of the first constant, the shared directory unit indicates the cache status of at least one constant in at least two processing cores, the query response is the response of the shared directory unit to the query request, and the read response is the response of the last-level cache or main memory to the read request. The last-level cache is the last-level cache in the processor. The later-arriving response of the query response or the read response may be ignored, treated as an invalid response, or discarded. If the query response is earlier than the read response of the last-level cache or main memory, then wait for the read response of the second processing core according to the query response, or send a read request to the second processing core; and set the read response of the last-level cache or main memory to invalid (or discard the read response of the last-level cache or main memory). If the read response of the last-level cache or main memory is earlier than the query response, then set the query response to invalid (or discard the query response). This can avoid the excessive delay caused by the first processing core still having to request the first constant from the last-level cache or main memory due to the absence of the first constant in the shared directory unit, that is, shorten the maximum transmission delay of the method provided in the embodiment of the present application to the same as the traditional method as much as possible.

[0218] For a shared directory unit, in order to ensure the timeliness of the cache status of at least one constant in the shared directory unit, the shared directory unit should be maintained in a timely manner.

[0219] 2. Maintenance of shared directory units.

[0220] In some embodiments, the method further includes: when the shared directory unit updates a cache status of at least one constant in any processing core of the at least two processing cores, the shared directory unit updates a cache status of at least one constant in the shared directory unit.

[0221] It should be noted that after the first processing core reads the first constant from the second processing core or the last level cache or the main memory, the first constant may be cached in the first cache or not. In some embodiments, the first constant is not cached in the first cache, but the first constant is cached in the second processing core. Compared with the traditional method, the method provided in the embodiment of the present application supports the first processing core to read the first constant from the second processing core, so that the reading efficiency of the first constant is higher than the traditional method. In another embodiment, after the first constant is cached in the first cache, the first processing core can subsequently read the first constant directly from its corresponding first cache without having to read the first constant from the second processing core or the last level cache or the main memory. This can greatly improve the efficiency of the first processing core in reading the first constant and reduce the delay in reading the first constant. However, the capacity of the first cache of the first processing core is often small in order to ensure a higher read and write speed, so the first processing core needs to determine whether to cache the first constant in the first cache.

[0222] In some embodiments, the method further includes: the first processing core caching the first constant in the first cache when a third condition is satisfied; wherein the third condition includes at least one of the following: an access frequency of the first constant is greater than a third threshold, the access frequency being used to indicate a frequency at which the first processing core requests the first constant from the third processing core; and an occupancy rate of the first cache is less than a fourth threshold.

[0223] In some embodiments, the shared directory unit also supports pre-fetching constants into some cache cores, as shown below.

[0224] In some embodiments, the method further includes: when a second constant among at least one constant satisfies a fourth condition, the shared directory unit sends a prefetch instruction to at least one fourth processing core, the prefetch instruction being used to cache the second constant in the cache of the at least one fourth processing core, the distance between each fourth processing core among the at least one fourth processing core and each fifth processing core among the at least two fifth processing cores being less than a fifth threshold, and the at least two fifth processing cores being processing cores for accessing the second constant.

[0225] The fourth condition includes at least one of the following: the number of at least two fifth processing cores is greater than the first number; and the access frequency of the second constant is greater than a sixth threshold.

[0226] On the other hand, an embodiment of the present application provides a graphics card, comprising the processor described in the above embodiments. Optionally, the processor is a GPU.

[0227] In another aspect, an embodiment of the present application provides a computer device comprising the processor described above. Optionally, the processor is a GPU. The computer device can be at least one of a portable computer, a desktop computer, a server, a server cluster, an artificial intelligence (AI) computing cluster, and a cloud computing cluster. The AI ​​computing cluster may also be referred to as an intelligent computing cluster or a smart computing cluster.

[0228] It should be understood that the "multiple" mentioned in this article refers to two or more. The character " / " generally indicates that the objects associated with each other are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution order between the steps. In some other embodiments, the above steps may also be executed in a non-numbered order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the opposite order to that shown in the figure. This embodiment of the application is not limited to this.

[0229] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A processor, characterized in that: The processor includes at least two processing cores; The first processing core of the at least two processing cores is configured to obtain the first constant cached in the second processing core if the first constant does not hit in the first cache; The first cache is a cache of the first processing core.

2. The processor according to claim 1, wherein: The second processing core is a whole or part of an adjacent core, and the adjacent core is a processing core that is physically or logically adjacent to the first processing core.

3. The processor according to claim 2, wherein: The adjacent core is a processing core that is logically adjacent to the first processing core; the adjacent processing core is a processing core that meets the first condition; The first condition includes at least one of the following: The load is lower than a first threshold; Sort the loads in descending order in the first n positions, where n is a positive integer; The distance to the first processing core is lower than a second threshold; The first k positions are sorted in ascending order according to the distance from the first processing core, where k is a positive integer.

4. The processor according to any one of claims 1 to 3, characterized in that: the first processing core being configured to, when the first constant does not hit in the first cache, send a query request to a shared directory unit, the query request including an address of the first constant, the shared directory unit being configured to indicate a cache status of at least one constant in the at least two processing cores; The shared directory unit is configured to determine, based on the query request, that the processing core having the first constant cached therein is the second processing core; The shared directory unit is configured to request the first constant from the second processing core, or to send an identifier of the second processing core to the first processing core so that the first processing core requests the first constant from the second processing core.

5. The processor according to any one of claims 1 to 3, characterized in that: the first processing core being configured to, if the first constant does not hit in the first cache, determine all or part of adjacent cores as second processing cores, the adjacent cores being processing cores that are physically or logically adjacent to the first processing core; requesting the first constant from the second processing core; if the first constant is not found in the second processing core, sending a query request to a shared directory unit, the query request including an address of the first constant, the shared directory unit being configured to indicate a cache status of at least one constant in the at least two processing cores; The shared directory unit is configured to determine, based on the query request, that the processing core having the first constant cached therein is the second processing core; The shared directory unit is configured to request the first constant from the second processing core, or to send an identifier of the second processing core to the first processing core so that the first processing core requests the first constant from the second processing core.

6. The processor according to any one of claims 1 to 3, characterized in that: The first processing core is configured to send a query request to a shared directory unit if the first constant does not hit in the first cache; and sending read requests to the last-level cache or main memory; determining the first constant based on the query response or the read response, whichever arrives earlier; The query request includes the address of the first constant, the read request includes the address of the first constant, and the shared directory unit is used to indicate a cache status of at least one constant in the at least two processing cores; The query response is the response of the shared directory unit to the query request, and the read response is the response of the last-level cache or the main memory to the read request; the last-level cache is the last-level cache in the processor.

7. The processor according to claim 4, wherein: The processor includes the shared directory unit, and the shared directory unit is connected to the at least two processing cores; Alternatively, the shared directory unit is independent of the processor, and the processor is connected to the shared directory unit.

8. The processor according to claim 4, wherein: The shared directory unit includes at least one cache core bitmap, and the at least one cache core bitmap has a one-to-one correspondence with the at least one constant; The cache core bitmap corresponding to each constant in the at least one constant is used to indicate the processing core in the at least two processing cores that caches the constant.

9. The processor according to claim 6, wherein: The shared directory unit further includes at least one status flag, the at least one status flag corresponds to the at least one constant in a one-to-one manner, and each status indicator in the at least one status flag is used to indicate whether the constant corresponding to the status flag is in an exclusive state or a shared state.

10. The processor according to claim 9, wherein: The hot data prefetch priority of the constant in the exclusive state is higher than the hot data prefetch priority of the shared state; the hot data prefetch priority is used to indicate the priority of the hot data to be prefetched into the cache of the processing core, and the hot data refers to data that is frequently accessed by the processing core or data that is expected to be frequently accessed by the processing core, and the frequent access means that the access frequency is greater than the access threshold.

11. The processor according to claim 4, wherein: The shared directory unit is used to send a read request to the second processing core, where the read request includes the address of the first constant, and the target address of the read request is the first cache.

12. The processor according to claim 4, wherein: The shared directory unit is configured to send miss information to the first processing core when no processing core having the first constant cached is found; The first processing core is configured to request the first constant from a final-level cache or a main memory, where the final-level cache is the last level cache in the processor.

13. The processor according to claim 4, wherein: The shared directory unit is configured to update a cache status of the at least one constant in the shared directory unit when a cache in any processing core of the at least two processing cores is updated.

14. The processor according to claim 4, wherein: There are multiple second processing cores; The first processing core or the shared directory unit is configured to request the first constant from a processing core among the plurality of second processing cores that meets a second condition; The second condition includes at least one of the following: Minimum load; The distance to the first processing core is the shortest.

15. The processor according to any one of claims 1 to 3, characterized in that: the first processing core is configured to cache the first constant in the first cache when a third condition is satisfied; The third condition includes at least one of the following: An access frequency of the first constant is greater than a third threshold, where the access frequency indicates a frequency at which the first processing core requests the first constant from the third processing core; The occupancy rate of the first cache is less than a fourth threshold.

16. The processor according to claim 4, wherein: the shared directory unit being configured to send a prefetch instruction to at least one fourth processing core if a second constant among the at least one constant satisfies a fourth condition, the prefetch instruction being configured to cache the second constant in a cache of the at least one fourth processing core, a distance between each of the at least one fourth processing core and each of the at least two fifth processing cores being processing cores that access the second constant being less than a fifth threshold; The fourth condition includes at least one of the following: The number of the at least two fifth processing cores is greater than the first number; The access frequency of the second constant is greater than a sixth threshold.

17. A graphics card, characterized in that: The graphics card includes the processor according to any one of claims 1 to 16.

18. A computer device, characterized in that: The computer device comprises the processor according to any one of claims 1 to 16.

19. A method for reading a constant, characterized in that: The method is executed by a processor, wherein the processor includes at least two processing cores; the method includes: A first processing core among the at least two processing cores obtains the first constant cached in a second processing core when the first constant does not hit in the first cache; The first cache is a cache of the first processing core.

Citation Information

Patent Citations

  • Prefetch configuration optimization method for sharing prefetcher in multiple core groups

    CN119917443A

  • Cache management method and device for memory access, equipment and storage medium

    CN119988255A

  • Data caching method and device, storage system, electronic equipment and storage medium

    CN120256332A

  • Request processing method of multi-core processor, multi-core processor, product and server

    CN120336213A