Method, apparatus, electronic device, and storage medium for selecting a processor core
By sorting based on the distance between the processor core and the listening filter in an on-chip network, and selecting the processor core with the smallest distance as the target processor core, the problems of slow data response speed and low query efficiency in the prior art are solved, and more efficient data query is achieved.
Patent Information
- Application Number
- CN202510152363.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The prior art method of selecting processor cores in on-chip networks results in slow data response speed and low query efficiency.
By receiving the query request sent by the query node, the query address corresponding to the query request is obtained, based on the correspondence between the address recorded by the listening filter and the processor core, and classifying the processor core based on the distance between the processor core and the listening filter, and determining the target processor core stored in the data corresponding to the query address.
By shortening the data transmission path, improving the data response speed and query efficiency, and improving the performance of the NOC system.
Smart Images

Figure CN119621652B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device, electronic device and computer-readable storage medium for selecting a processor core. Background Art
[0002] In a network-on-chip (NOC), in order to improve system performance, a snoop filter (SF) is usually used to send a snoop request to a request node (RN) recorded in the SF to obtain query data of the snoop request, thereby reducing system resource consumption and improving overall performance.
[0003] Taking the requesting node as a processor core as an example, the relevant technology usually randomly selects a processor core from the processor cores recorded in the SF as the target processor core for the listening request, or uses the processor core that has most recently accessed the data as the target processor core for the listening request and sends a listening request to the target processor core.
[0004] The processor core selection method proposed in the related art has a slow data response speed and low query efficiency when querying data. Summary of the invention
[0005] Embodiments of the present application provide a method, device, electronic device, and computer-readable storage medium for selecting a processor core to solve problems in related technologies.
[0006] In a first aspect, an embodiment of the present application provides a method for selecting a processor core, the method comprising:
[0007] Receive a query request sent by a query node, and obtain a query address corresponding to the query request; the network on chip includes multiple nodes; the node includes at least one listening filter and multiple processor cores; the query node is any node in the network on chip except the listening filter;
[0008] Based on the correspondence between the address recorded by the listening filter and the processor core, and multiple classifications, the target processor core where the data corresponding to the query address is stored is determined; the multiple classifications are obtained by classifying the processor core according to the distance between the processor core and the listening filter; each of the classifications corresponds to a distance range; different classifications correspond to different distance ranges.
[0009] In a second aspect, an embodiment of the present application provides a device for selecting a processor core, the device comprising:
[0010] A first acquisition module is used to receive a query request sent by a query node and acquire a query address corresponding to the query request; the network on chip includes a plurality of nodes; the node includes at least one listening filter and a plurality of processor cores; the query node is any node in the network on chip except the listening filter;
[0011] A determination module is used to determine the target processor core where the data corresponding to the query address is stored based on the correspondence between the address recorded by the listening filter and the processor core, and multiple classifications; the multiple classifications are obtained by classifying the processor cores according to the distance between the processor cores and the listening filter; each of the classifications corresponds to a distance range; different classifications correspond to different distance ranges.
[0012] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the method of the first aspect.
[0013] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to execute the method of the first aspect.
[0014] In an embodiment of the present application, by receiving a query request sent by a query node, obtaining a query address corresponding to the query request, and determining a target processor core where the data corresponding to the query address is stored based on the correspondence between the address recorded by the listening filter and the processor core, and multiple classifications obtained by classifying the processor core according to the distance between the processor core and the listening filter, multiple distance ranges and classifications corresponding to the multiple distance ranges can be determined based on the distance between the processor core and the listening filter, so that the target processor core can be determined from the classifications corresponding to the distance ranges according to the distance ranges. For example, the target processor core is determined from the processor core in the classification with the smallest distance range. Since the distance from the processor core in the classification with the smallest distance range to the listening filter is the smallest, the data transmission path can be shortened, and when querying data, the data response speed and query efficiency can be improved.
[0015] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0017] Figure 1 is a flowchart of the steps of a method for selecting a processor core provided in an embodiment of the present application;
[0018] Figure 2 It is a topological structure diagram of a network on chip provided in an embodiment of the present application;
[0019] Figure 3 It is a schematic diagram of a listening filter provided in an embodiment of the present application;
[0020] Figure 4 is a schematic diagram of a processor core query data provided by an embodiment of the present application;
[0021] Figure 5 is a flowchart of specific steps of a method for selecting a processor core provided in an embodiment of the present application;
[0022] Figure 6 This is a flowchart of the steps for processing a query request provided by an embodiment of the present application;
[0023] Figure 7 This is a block diagram of a processor core selection device provided in an embodiment of the present application;
[0024] Figure 8 is a block diagram of an electronic device provided by an embodiment of the present invention;
[0025] Fig. 9 is a block diagram of another electronic device according to another embodiment of the present invention. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0027] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally a class, and the number of objects is not limited. For example, the first object can be one or more. In addition, the term "and / or" in the specification and claims is used to describe the association relationship of associated objects, indicating that three kinds of relationships can exist, for example, A and / or B can be represented: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the front and back associated objects are a kind of "or" relationship. In the embodiment of the present application, the term "multiple" refers to two or more, and other quantifiers are similar.
[0028] NOC is a communication network structure implemented on an integrated circuit chip, which is used to connect various functional modules, processor cores, storage units and other important components on the chip. With the increase in chip integration and the rise of multi-core processors, NOC has become increasingly important because it provides an efficient and low-latency communication method within the chip.
[0029] In NOC, in order to improve system performance, the SF functional module is usually used to record which RNs store data related to specific cache lines. When a snoop request needs to be initiated, it is only necessary to send a snoop request to the RN recorded in the SF without broadcasting the snoop request to all processor cores in the system. By reducing unnecessary snoop requests, system resource consumption can be significantly reduced and overall performance can be improved. Among them, RN refers to the node responsible for generating transaction requests and sending these requests to other nodes in the system. In the coherent hub interface (CHI) protocol, RN is the key component responsible for initiating transaction requests and interacting with other system nodes. Its main functions include generating transaction requests, address translation and routing, transaction management, and maintaining cache consistency. Through these functions, RN ensures the efficiency and consistency of data transmission in high-performance computing systems.
[0030] Taking the requesting node as a processor core as an example, the common schemes for selecting the target processor core for the listening request mainly include the following two: randomly selecting a processor core and the last accessed processor core. In the method of randomly selecting a processor core, one processor core is selected as the target processor core for the listening request from the processor cores recorded by the SF functional module, and the listening request is only sent to the target processor core. However, this method does not consider the physical distance between the target processor core and the SF. Random selection may cause the selected processor core to be far away from the SF, thus requiring a longer transmission path, resulting in a longer data response time, which in turn affects the NOC performance.
[0031] In the last accessed processor core method, each time a processor core accesses a specific cache line, the target processor core recorded in the SF is updated to the most recently accessed processor core. Subsequently, the SF sends a snoop request to this target processor core. However, this method also does not consider the distance between the target processor core and the SF. If the most recently accessed processor core is far away from the SF, the data response time may also be long, resulting in reduced system performance.
[0032] In an embodiment of the present application, by receiving a query request sent by a query node, obtaining a query address corresponding to the query request, and determining a target processor core where the data corresponding to the query address is stored based on the correspondence between the address recorded by the listening filter and the processor core, and multiple classifications obtained by classifying the processor core according to the distance between the processor core and the listening filter, multiple distance ranges and classifications corresponding to the multiple distance ranges can be determined based on the distance between the processor core and the listening filter, so that the target processor core can be determined from the classifications corresponding to the distance ranges according to the distance ranges. For example, the target processor core is determined from the processor core in the classification with the smallest distance range. Since the distance from the processor core in the classification with the smallest distance range to the listening filter is the smallest, the data transmission path can be shortened, and when querying data, the data response speed and query efficiency can be improved.
[0033] Figure 1 , is a flowchart of a method for selecting a processor core provided by an embodiment of the present application, such as Figure 1 As shown, the method may include:
[0034] Step 101, receiving a query request sent by a query node, and obtaining a query address corresponding to the query request; the on-chip network includes multiple nodes; the node includes at least one listening filter and multiple processor cores; the query node is any node in the on-chip network except the listening filter.
[0035] For example, the on-chip network includes a snoop filter and multiple processor cores, wherein each processor core has a unique number. Figure 2 , Figure 2 Four processor cores are shown, numbered RN-0, RN-1, RN-2 and RN-3 respectively.
[0036] For example, the query node is any node in the on-chip network except the listening filter, which is used to send a query request to the on-chip network to read data from the on-chip network. Figure 2,Each processor core has a unique number, and the query node can be the processor core RN-0, which sends a query request to the snoop filter in the ,on-chip network to read data from the on-chip network.
[0037] For example, refer to Figure 2 After the on-chip network receives the query request sent by the processor core RN-0, the query address 6 corresponding to the query request can be obtained, thereby obtaining the target processor core where the data corresponding to the query address 6 is stored based on the correspondence between the address recorded in the snoop filter and the processor core. For example, if the snoop filter records that the data at address 6 is in the processor core RN-1, then the target processor core where the data corresponding to the query address 6 is stored is the processor core RN-1.
[0038] Step 102: Based on the correspondence between the address recorded by the listening filter and the processor core, and multiple classifications, determine the target processor core where the data corresponding to the query address is stored; the multiple classifications are obtained by classifying the processor core according to the distance between the processor core and the listening filter; each of the classifications corresponds to a distance range; different classifications correspond to different distance ranges.
[0039] For example, the snoop filter can record the correspondence between addresses and processor cores through a log book, that is, record which addresses of data are in which processor cores. Figure 3 The information recorded by the snoop filter includes: snoop filter-tag information. The snoop filter-tag information is part of the address and is used to specify which processor cores the data of the corresponding address is stored in. For example, the data at address 6 is in processor core RN-1, the data at address 7 is in processor core RN-2, the data at address 8 is in processor core RN-1 and processor core RN-2, and the data at address 8 is in processor core RN-2 and processor core RN-3.
[0040] For example, the present application calculates the distance between the processor core and the snoop filter, and classifies the processor cores according to the distance, thereby obtaining multiple classifications of the processor cores. The distance between the processor core and the snoop filter can be the distance between the routing module connected to the processor core and the routing module connected to the snoop filter, referring to Figure 4 The distance between the routing module connected to the processor core RN-1 and the routing module connected to the snoop filter is one routing module, so the distance between the processor core RN-1 and the snoop filter is 1.
[0041] For example, the present application classifies the ones with the same distance range into one category. The distance range can be pre-set, the distance range can be a fixed value, for example, the distance range can be 1, or the distance range can be a range, for example, the distance range can be 1-2, that is, in the category of the minimum distance range, the distance from each processor core to the listening filter is 1 or 2. Taking the distance range of 1 as an example, the processor core of one routing module of the distance listening filter is set as a short-distance processor core, the processor cores of two routing modules of the distance listening filter are set as medium-distance processor cores, and the rest are set as long-distance processor cores. Refer to Figure 4 , Figure 4 The processor cores in include processor cores filled with horizontal lines, processor cores filled with vertical lines, and processor cores without filling. Taking the processor core as a processor core filled with horizontal lines as an example, the routing module connected to the processor core filled with horizontal lines and the routing module connected to the listening filter differ by one routing module number, therefore, the processor core filled with horizontal lines is set as a processor core at a short distance. Taking the processor core as a processor core filled with vertical lines as an example, the routing module connected to the processor core filled with vertical lines and the routing module connected to the listening filter differ by two routing modules number, therefore, the processor core filled with vertical lines is set as a processor core at a medium distance. Taking the processor core as an unfilled processor core as an example, the routing module connected to the processor core without filling and the routing module connected to the listening filter differ by three or more routing modules number, therefore, the processor core without filling is set as a processor core at a long distance.
[0042] Taking the distance range of 1-2 as an example, the processor core of one routing module or two routing modules of the distance listening filter is set as a short-distance processor core, the processor core of three routing modules or four routing modules of the distance listening filter is set as a medium-distance processor core, and the rest is set as a long-distance processor core. In addition, it should be noted that the processor cores can be divided into three categories, or the processor cores can be divided into categories greater than three, for example, five categories. Taking the division of processor cores into five categories as an example, the processor core of one routing module of the distance listening filter is set as a first-category processor core, the processor core of two routing modules of the distance listening filter is set as a second-category processor core, the processor core of three routing modules of the distance listening filter is set as a third-category processor core, the processor core of four routing modules of the distance listening filter is set as a fourth-category processor core, and the rest is set as a fifth-category processor core.
[0043] For example, refer to Figure 3After the processor cores are divided into short-distance processor cores, medium-distance processor cores, and long-distance processor cores, the identifiers of the short-distance processor cores, medium-distance processor cores, and long-distance processor cores can be recorded in the snoop filter. The content recorded by the snoop filter also includes: snoop filter logic control. The snoop filter logic control is a functional component that reads the random access memory (RAM) of the snoop filter label, writes the RAM of the snoop filter label, and compares the hit or miss of the snoop filter label. It mainly controls the read and write operations of the RAM of the snoop filter label.
[0044] For example, after classifying the processor cores and obtaining multiple classifications of the processor cores, the processor core with the smallest distance from the listening filter is selected from the processor cores of the multiple classifications according to the correspondence between the address and the processor core, and the multiple classifications, as the target processor core for storing the data corresponding to the query address. This can shorten the data transmission path and improve the data response speed and query efficiency when querying data.
[0045] In summary, in an embodiment of the present application, by receiving a query request sent by a query node, obtaining a query address corresponding to the query request, and determining a target processor core where the data corresponding to the query address is stored based on the correspondence between the address recorded by the listening filter and the processor core, and based on the distance between the processor core and the listening filter, classifying the processor core to obtain multiple categories, multiple distance ranges and categories corresponding to the multiple distance ranges can be determined based on the distance between the processor core and the listening filter, so that the target processor core can be determined from the categories corresponding to the distance ranges according to the distance ranges. For example, the target processor core is determined from the processor core in the category with the smallest distance range. Since the distance from the processor core in the category with the smallest distance range to the listening filter is the smallest, the data transmission path can be shortened, and when querying data, the data response speed and query efficiency can be improved.
[0046] Figure 5 , is a flowchart of specific steps of a method for selecting a processor core provided in an embodiment of the present application, such as Figure 5 As shown, the method may include:
[0047] Step 201, receiving a query request sent by a query node, and obtaining a query address corresponding to the query request; the on-chip network includes multiple nodes; the node includes at least one listening filter and multiple processor cores; the query node is any node in the on-chip network except the listening filter.
[0048] This step may specifically refer to the above step 101, and will not be described in detail here.
[0049] Step 202: Based on the correspondence between the address recorded by the listening filter and the processor core, determine the candidate processor core that matches the query address; the multiple categories are obtained by classifying the processor core according to the distance between the processor core and the listening filter; each category corresponds to a distance range; different categories correspond to different distance ranges.
[0050] For example, after obtaining the query address corresponding to the query request, the query address is matched with the address recorded in the listening filter. If the address recorded in the listening filter is matched, it is a hit. The processor core corresponding to the address recorded in the listening filter is determined to be a candidate processor core matching the query address. When the query address is queried, a hit occurs. If the address recorded in the listening filter is not matched, a miss occurs when the query address is queried.
[0051] For example, the snoop filter records that the data at address 6 is in processor core RN-1, the data at address 7 is in processor core RN-2, the data at address 8 is in processor core RN-1 and processor core RN-2, and the data at address 8 is in processor core RN-2 and processor core RN-3. For example, if the query address corresponding to the query request is address 6, address 6 is recorded in the snoop filter, and the processor core RN-1 corresponding to address 6 is determined as a candidate processor core that matches query address 6. When querying query address 6, a hit occurs. For example, if the query address corresponding to the query request is address 5, address 5 is not recorded in the snoop filter, and when querying query address 5, a miss occurs.
[0052] Step 203: The candidate classification to which the candidate processor core belongs and the classification with the smallest distance range is taken as the target classification.
[0053] For example, the present application calculates the distance between all processor cores and the snoop filter, and classifies the processor cores according to the distance, to obtain multiple classifications of the processor cores. After determining the candidate processor cores matching the query address based on the correspondence between the addresses recorded by the snoop filter and the processor cores, the classification to which the candidate processor cores belong is determined, and the classification with the smallest distance is used as the target classification.
[0054] For example, taking the query address of 6 as an example, the candidate processor core that matches the query address 6 is processor core RN-1. Since there is only one candidate processor core at the query address 6, the candidate category to which processor core RN-1 belongs is determined as the target category. Taking the query address of 8 as an example, the candidate processor cores that match the query address 8 are processor core RN-1 and processor core RN-2. Since there are two candidate processor cores at the query address 6, the candidate categories to which processor core RN-1 and processor core RN-2 belong are first determined. If the candidate categories to which processor core RN-1 and processor core RN-2 belong are the same, the candidate category is used as the target category. If the candidate categories to which processor core RN-1 and processor core RN-2 belong are different, the category with the smallest distance range among the candidate categories is used as the target category.
[0055] Optionally, the candidate categories include: a first category, a second category, and a third category; in the first category, the distance between the processor core and the listening filter is a first distance range; in the second category, the distance between the processor core and the listening filter is a second distance range; in the third category, the distance between the processor core and the listening filter is a third distance range; the distance in the first distance range is smaller than the distance in the second distance range, which is smaller than the distance in the third distance range; step 203 may specifically include:
[0056] Sub-step 2031: if there is a candidate classification that is the first classification, then use the first classification as the target classification;
[0057] Sub-step 2032: if there is no candidate classification for the first classification, and there is a candidate classification for the second classification, then the second classification is used as the target classification;
[0058] Sub-step 2033: If there are no candidate categories for the first category and the second category, and there is a candidate category for the third category, then the third category is used as the target category.
[0059] For sub-steps 2031-2033, taking the first distance range of 1, the second distance range of 2, and the third distance range of 3 as an example, the distance between the processor core and the listening filter in the first category is 1, the distance between the processor core and the listening filter in the second category is 2, and the distance between the processor core and the listening filter in the third category is 3.
[0060] For example, the first category includes processor core RN-1, processor core RN-3, and processor core RN-4. The candidate processor core of query address 6 is processor core RN-1, so the candidate category of processor core RN-1 is the first category, and the first category is used as the target category. The second category includes processor core RN-2 and processor core RN-6. The candidate processor core of query address 7 is processor core RN-2, so the candidate category of processor core RN-2 is the second category. Since there is no candidate category for the first category, and there is a candidate category for the second category, the second category is used as the target category. When the listening filter needs to send a listening request, it first queries whether the processor core in the vicinity is hit, and then decides whether to send a listening request to the processor core in the vicinity. The listening request is sent to the processor core in the vicinity first. Since the listening filter selects the processor core in the vicinity as the target processor core, the listening request only needs to be forwarded and transmitted through a small number of modules, so the listening request can reach the target processor core faster. On the other hand, since the target processor core is closer to the snoop filter in distance, the response of the processor core can reach the snoop filter faster. In summary, the snoop filter based on distance priority can speed up the processing and response of snoop requests, thereby improving the throughput of the NOC system and enhancing the performance of the system.
[0061] Step 204: Select one processor core from the processor cores of the target classification as the target processor core for storing the data corresponding to the query address.
[0062] For example, after obtaining the target classification, since the processor cores in the target classification are all processor cores that match the query address, a processor core can be selected from the processor cores in the target classification as the target processor core.
[0063] Optionally, step 204 may specifically include:
[0064] Sub-step 2041: If the target category includes multiple processor cores, randomly select one processor core from the multiple processor cores as the target processor core.
[0065] For example, taking the target category as the first category, if the first category includes multiple processor cores, a processor core is randomly selected from the multiple processor cores as the target processor core. For example, the first category includes processor cores RN-1, processor core RN-2, and processor core RN-3, then a processor core is randomly selected from processor core RN-1, processor core RN-2, and processor core RN-3 as the target processor core.
[0066] Optionally, after step 204, the method further includes:
[0067] Step 205: extract data corresponding to the query address from the target processor core through the intercept filter, and send the data to the query node.
[0068] For example, after determining the target processor core where the data corresponding to the query address is stored, the listening filter sends a listening request to the target processor core. After the target processor core receives the listening request, it queries the data corresponding to the query address and sends the data to the listening filter, which then sends the data to the processor core RN-0.
[0069] For example, refer to Figure 2 Taking the target processor core as processor core RN-2 as an example, the listening filter sends a listening request to processor core RN-2. After processor core RN-2 receives the listening request, it queries the data corresponding to the query address and sends the data to the listening filter. The listening filter then sends the data to processor core RN-0.
[0070] Optionally, the network on chip further includes routing modules arranged in a grid array, and the routing modules correspond to the nodes one by one; the method further includes:
[0071] Step 206: Acquire a first routing module connected to the processor core and a second routing module connected to the listening filter;
[0072] Step 207: Calculate the number of routing modules that differs between the first routing module and the second routing module, and determine the number of routing modules as the distance from the processor core to the snoop filter.
[0073] For steps 206 and 207, refer to Figure 3 , taking the processor core filled with horizontal lines as an example, the first routing module connected to the processor core filled with horizontal lines and the second routing module connected to the listening filter differ by one routing module number, so the distance from the processor core filled with horizontal lines to the listening filter is 1. Taking the processor core filled with vertical lines as an example, the first routing module connected to the processor core filled with vertical lines and the second routing module connected to the listening filter differ by two routing modules number, so the distance from the processor core filled with vertical lines to the listening filter is 2. Taking the processor core as an unfilled processor core as an example, the first routing module connected to the unfilled processor core and the second routing module connected to the listening filter differ by three or more routing modules number, so the distance from the unfilled processor core to the listening filter is 3 or a number greater than 3.
[0074] Optionally, the method further includes:
[0075] Step 208: If the identifier of the processor core in the category is the same as the identifier of the candidate processor core, then add the processor core in the category to the set of candidate processor cores corresponding to the category;
[0076] Step 209: randomly select a processor core from the set of candidate processor cores corresponding to each of the categories, and obtain the distance from each of the processor cores to the listening filter;
[0077] Step 210: Determine the processor core with the smallest distance as the target processor core.
[0078] For steps 208-210, after the processor cores are classified according to the distance between the processor cores and the listening filter to obtain multiple categories of processor cores, the identifiers of the processor cores in each category are compared with the identifiers of the candidate processor cores. If the identifiers of the processor cores in the category are the same as the identifiers of the candidate processor cores, the processor cores in the category are added to the set of candidate processor cores corresponding to the category. For example, the identifier of the processor core is the number of the processor core, and the categories of the processor cores are the first category, the second category, and the third category. The first category includes processor core RN-1, processor core RN-3, and processor core RN-4. The candidate processor core at query address 6 is processor core RN-1, so processor core RN-1 in the first category is added to the set of candidate processor cores corresponding to the first category. The second category includes processor core RN-2 and processor core RN-6. The candidate processor cores for query address 8 are processor core RN-1 and processor core RN-2. Therefore, processor core RN-1 in the first category is added to the candidate processor core set corresponding to the first category, and processor core RN-2 in the second category is added to the candidate processor core set corresponding to the second category.
[0079] For example, a processor core is randomly selected from the candidate processor core set corresponding to each classification, and the distance from each processor core to the listening filter is obtained, and the processor core with the smallest distance is determined as the target processor core. For example, the candidate processor core set of the first classification includes processor core RN-1, and the candidate processor core set of the second classification includes processor core RN-2. From the candidate processor core set corresponding to each classification, a processor core is randomly selected to obtain processor core RN-1 and processor core RN-2. Since the distance from the processor core in the first classification to the listening filter is less than the distance from the processor core in the second classification to the listening filter, processor core RN-1 is determined as the target processor core.
[0080] In summary, in an embodiment of the present application, by receiving a query request sent by a query node, obtaining a query address corresponding to the query request, and determining a target processor core where the data corresponding to the query address is stored based on the correspondence between the address recorded by the listening filter and the processor core, and based on the distance between the processor core and the listening filter, classifying the processor core to obtain multiple categories, multiple distance ranges and categories corresponding to the multiple distance ranges can be determined based on the distance between the processor core and the listening filter, so that the target processor core can be determined from the categories corresponding to the distance ranges according to the distance ranges. For example, the target processor core is determined from the processor core in the category with the smallest distance range. Since the distance from the processor core in the category with the smallest distance range to the listening filter is the smallest, the data transmission path can be shortened, and when querying data, the data response speed and query efficiency can be improved.
[0081] Figure 6 This is a flowchart of a process for processing a query request provided by an embodiment of the present application, referring to Figure 6 , this step process includes:
[0082] Step S1, determine whether the processor core at a close distance is hit, if it is hit, jump to step S2, if not, jump to step S3;
[0083] Step S2: selecting a processor core at a close distance as a target processor core of the intercept request, thereby sending the intercept request to the processor core at a close distance;
[0084] Step S3, judging whether the processor core in the middle distance is hit, if it is hit, jumping to step S4, if not, jumping to step S5;
[0085] Step S4: Selecting a processor core at a medium distance as a target processor core of the snoop request, thereby sending a snoop request to the processor core at the medium distance;
[0086] Step S5, determine whether the remote processor core is hit, if it is hit, jump to step S6, if not, jump to step S7;
[0087] Step S6, selecting a remote processor core as a target processor core for intercepting the request;
[0088] Step S7: Do not send a listening request.
[0089] The present application uses a distance-priority based snooping filter to give priority to sending snooping requests to processor cores that are close when sending snooping requests. Since the snooping filter selects processor cores that are close as target processor cores, the snooping request only needs to be forwarded and transmitted through a small number of modules, so the snooping request can reach the target processor core faster. On the other hand, since the target processor core is closer to the snooping filter, the response of the processor core can reach the snooping filter faster. In summary, the distance-priority based snooping filter can speed up the processing and response of snooping requests, thereby improving the throughput of the NOC system and enhancing the performance of the system.
[0090] Figure 7 is a block diagram of a processor core selection device 30 provided in an embodiment of the present application, the device comprising:
[0091] The first acquisition module 301 is used to receive a query request sent by a query node and obtain a query address corresponding to the query request; the network on chip includes multiple nodes; the node includes at least one listening filter and multiple processor cores; the query node is any node in the network on chip except the listening filter;
[0092] The determination module 302 is used to determine the target processor core where the data corresponding to the query address is stored based on the correspondence between the address recorded by the listening filter and the processor core, and multiple classifications; the multiple classifications are obtained by classifying the processor cores according to the distance between the processor cores and the listening filter; each of the classifications corresponds to a distance range; different classifications correspond to different distance ranges.
[0093] Optionally, the determining module includes:
[0094] A first determination submodule, configured to determine a candidate processor core matching the query address based on the correspondence between the address recorded by the intercept filter and the processor core;
[0095] A first selection submodule, configured to select a category from the candidate categories to which the candidate processor cores belong as a target category;
[0096] The second selection submodule is used to select a processor core from the processor cores of the target classification as the target processor core.
[0097] Optionally, the first selection submodule includes:
[0098] The determination unit is used to select the category with the smallest distance range among the candidate categories as the target category.
[0099] Optionally, the candidate categories include: a first category, a second category, and a third category; in the first category, the distance between the processor core and the snoop filter is a first distance range; in the second category, the distance between the processor core and the snoop filter is a second distance range; in the third category, the distance between the processor core and the snoop filter is a third distance range; the distance in the first distance range is smaller than the distance in the second distance range, which is smaller than the distance in the third distance range;
[0100] Optionally, the determining unit includes:
[0101] A first determining subunit, configured to use the first category as the target category if any candidate category is the first category;
[0102] A second determining subunit is configured to use the second classification as the target classification if there is no candidate classification for the first classification and there is a candidate classification for the second classification;
[0103] The third determining subunit is configured to use the third category as the target category if there are no candidate categories for the first category and the second category, and there is a candidate category for the third category.
[0104] Optionally, the second selection submodule includes:
[0105] The selection unit is configured to randomly select a processor core from the multiple processor cores as the target processor core if the target category includes multiple processor cores.
[0106] Optionally, the determining module includes:
[0107] A second determination submodule, configured to determine a candidate processor core matching the query address based on a correspondence between the address recorded by the intercept filter and the processor core;
[0108] an adding submodule, configured to add the processor core in the category to a set of candidate processor cores corresponding to the category if the identifier of the processor core in the category is the same as the identifier of the candidate processor core;
[0109] A third selection submodule, configured to randomly select a processor core from the set of candidate processor cores corresponding to each of the categories, and obtain a distance from each of the processor cores to the listening filter;
[0110] The third determining submodule is used to determine the processor core with the smallest distance as the target processor core.
[0111] Optionally, the network on chip further includes routing modules arranged in a grid array, and the routing modules correspond to the nodes one by one; the device further includes:
[0112] A second acquisition module, used to acquire a first routing module connected to the processor core and a second routing module connected to the listening filter;
[0113] The calculation module is used to calculate the number of routing modules that differs between the first routing module and the second routing module, and determine the number of routing modules as the distance from the processor core to the snooping filter.
[0114] Optionally, the device further comprises:
[0115] The sending module is used to extract the data corresponding to the query address from the target processor core through the listening filter and send it to the query node.
[0116] In summary, in an embodiment of the present application, by receiving a query request sent by a query node, obtaining a query address corresponding to the query request, and determining a target processor core where the data corresponding to the query address is stored based on the correspondence between the address recorded by the listening filter and the processor core, and based on the distance between the processor core and the listening filter, classifying the processor core to obtain multiple categories, multiple distance ranges and categories corresponding to the multiple distance ranges can be determined based on the distance between the processor core and the listening filter, so that the target processor core can be determined from the categories corresponding to the distance ranges according to the distance ranges. For example, the target processor core is determined from the processor core in the category with the smallest distance range. Since the distance from the processor core in the category with the smallest distance range to the listening filter is the smallest, the data transmission path can be shortened, and when querying data, the data response speed and query efficiency can be improved.
[0117] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts may refer to the partial description of the method embodiment.
[0118] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0119] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0120] An embodiment of the present application provides a processor core selection device, including a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors to include methods for performing one or more of the above embodiments.
[0121] Figure 8 4 is a block diagram of an electronic device 400 according to an exemplary embodiment. For example, the electronic device 400 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0122] Reference Figure 8 , the electronic device 400 may include one or more of the following components: a processing component 402 , a memory 404 , a power component 406 , a multimedia component 408 , an audio component 410 , an input / output (I / O) interface 412 , a sensor component 414 , and a communication component 416 .
[0123] The processing component 402 generally controls the overall operation of the electronic device 400, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 402 may include one or more processors 420 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 402 may include one or more modules to facilitate the interaction between the processing component 402 and other components. For example, the processing component 402 may include a multimedia module to facilitate the interaction between the multimedia component 408 and the processing component 402.
[0124] The memory 404 is used to store various types of data to support the operation of the electronic device 400. Examples of such data include instructions for any application or method operating on the electronic device 400, contact data, phone book data, messages, pictures, multimedia, etc. The memory 404 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0125] The power supply component 406 provides power to the various components of the electronic device 400. The power supply component 406 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 400.
[0126] The multimedia component 408 includes a screen that provides an output interface between the electronic device 400 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 408 includes a front camera and / or a rear camera. When the electronic device 400 is in an operating mode, such as a shooting mode or a multimedia mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0127] The audio component 410 is used to output and / or input audio signals. For example, the audio component 410 includes a microphone (MIC), and when the electronic device 400 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is used to receive an external audio signal. The received audio signal can be further stored in the memory 404 or sent via the communication component 416. In some embodiments, the audio component 410 also includes a speaker for outputting audio signals.
[0128] I / O interface 412 provides an interface between processing component 402 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0129] The sensor assembly 414 includes one or more sensors for providing various aspects of status assessment for the electronic device 400. For example, the sensor assembly 414 can detect the open / closed state of the electronic device 400, the relative positioning of the components, such as the display and keypad of the electronic device 400, and the sensor assembly 414 can also detect the position change of the electronic device 400 or a component of the electronic device 400, the presence or absence of contact between the user and the electronic device 400, the orientation or acceleration / deceleration of the electronic device 400, and the temperature change of the electronic device 400. The sensor assembly 414 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 414 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 414 may also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0130] The communication component 416 is used to facilitate wired or wireless communication between the electronic device 400 and other devices. The electronic device 400 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 416 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 416 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0131] In an exemplary embodiment, the electronic device 400 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to implement the methods provided in the embodiments of the present application.
[0132] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 404 including instructions, and the instructions can be executed by a processor 420 of an electronic device 400 to perform the above method. For example, the non-transitory storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0133] Fig. 9 is a block diagram of an electronic device 500 according to an exemplary embodiment. For example, the electronic device 500 may be provided as a server. Fig. 9 , the electronic device 500 includes a processing component 522, which further includes one or more processors, and a memory resource represented by a memory 532, for storing instructions that can be executed by the processing component 522, such as an application. The application stored in the memory 532 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 522 is configured to execute instructions to perform the method provided in the embodiment of the present application.
[0134] The electronic device 500 may also include a power supply component 526 configured to perform power management of the electronic device 500, a wired or wireless network interface 550 configured to connect the electronic device 500 to a network, and an input / output (I / O) interface 558. The electronic device 500 may operate based on an operating system stored in the memory 532, such as Windows Server TM, Mac OS X TM, Unix TM, Linux TM, FreeBSD TM or the like.
[0135] An embodiment of the present application further provides a computer program product, including a computer program, which implements the method described in the above embodiment when executed by a processor.
[0136] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0137] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A method for selecting a processor core, characterized in that: The method comprises: Receive a query request sent by a query node, and obtain a query address corresponding to the query request; the network on chip includes multiple nodes; the node includes at least one listening filter and multiple processor cores; the query node is any node in the network on chip except the listening filter; Based on the correspondence between the address recorded by the intercept filter and the processor core, and multiple classifications, determining the target processor core where the data corresponding to the query address is stored includes: Determine a candidate processor core that matches the query address based on the correspondence between the address recorded by the snoop filter and the processor core; If the identifier of the processor core in the category is the same as the identifier of the candidate processor core, then adding the processor core in the category to the set of candidate processor cores corresponding to the category; Randomly select a processor core from the set of candidate processor cores corresponding to each of the categories, and obtain the distance from each of the processor cores to the listening filter; The processor core with the smallest distance is determined as the target processor core; the multiple categories are obtained by classifying the processor cores according to the distance between the processor cores and the listening filter; each category corresponds to a distance range; different categories correspond to different distance ranges.
2. The method according to claim 1, characterized in that The determining, based on the correspondence between the address and the processor core recorded by the intercept filter and the multiple classifications, the target processor core where the data corresponding to the query address is stored includes: Determine a candidate processor core that matches the query address based on the correspondence between the address recorded by the snoop filter and the processor core; Selecting a category as a target category from the candidate categories to which the candidate processor cores belong; A processor core is selected from the processor cores of the target classification as the target processor core.
3. The method according to claim 2, characterized in that The step of selecting a category as a target category from the candidate categories to which the candidate processor cores belong includes: The category with the smallest distance range among the candidate categories is used as the target category.
4. The method according to claim 3, characterized in that The candidate categories include: a first category, a second category, and a third category; the distance between the processor core and the snoop filter in the first category is a first distance range; the distance between the processor core and the snoop filter in the second category is a second distance range; the distance between the processor core and the snoop filter in the third category is a third distance range; all distances in the first distance range are smaller than the minimum distance in the second distance range; all distances in the second distance range are smaller than the minimum distance in the third distance range; The step of taking the category with the smallest distance range among the candidate categories as the target category includes: If there is a candidate classification that is the first classification, then the first classification is used as the target classification; If there is no candidate classification for the first classification, and there is a candidate classification for the second classification, the second classification is used as the target classification; If there are no candidate categories for the first category and the second category, and there is a candidate category for the third category, the third category is used as the target category.
5. The method according to claim 2, characterized in that: The step of selecting a processor core from the processor cores of the target classification as the target processor core comprises: If the target category includes a plurality of the processor cores, a processor core is randomly selected from the plurality of the processor cores as the target processor core.
6. The method according to claim 1, characterized in that The on-chip network also includes routing modules arranged in a grid array, and the routing modules correspond to the nodes one by one; the method also includes: Acquire a first routing module connected to the processor core and a second routing module connected to the listening filter; The number of routing modules that differs between the first routing module and the second routing module is calculated, and the number of routing modules is determined as the distance from the processor core to the snooping filter.
7. The method according to claim 1, characterized in that After determining the target processor core where the data corresponding to the query address is stored, the method further includes: The data corresponding to the query address is extracted from the target processor core through the snoop filter and sent to the query node.
8. A device for selecting a processor core, characterized in that: The device comprises: A first acquisition module is used to receive a query request sent by a query node and acquire a query address corresponding to the query request; the network on chip includes a plurality of nodes; the node includes at least one listening filter and a plurality of processor cores; the query node is any node in the network on chip except the listening filter; A determination module is used to determine the target processor core where the data corresponding to the query address is stored based on the correspondence between the address recorded by the intercept filter and the processor core, and multiple classifications, the determination module comprising: A second determination submodule, configured to determine a candidate processor core matching the query address based on a correspondence between the address recorded by the intercept filter and the processor core; an adding submodule, configured to add the processor core in the category to a set of candidate processor cores corresponding to the category if the identifier of the processor core in the category is the same as the identifier of the candidate processor core; A third selection submodule, configured to randomly select a processor core from the set of candidate processor cores corresponding to each of the categories, and obtain a distance from each of the processor cores to the listening filter; The third determination submodule is used to determine the processor core with the smallest distance as the target processor core; the multiple classifications are obtained by classifying the processor cores according to the distance between the processor cores and the listening filter; each of the classifications corresponds to a distance range; different classifications correspond to different distance ranges.
9. An electronic device, characterized in that: include: processor; A memory for storing the processor executable instructions; wherein the processor is configured to execute the instructions to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method as claimed in any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for achieving cache consistency between multiple cores
CN104462007A
Passive data caching implementation method of multi-bare-chip interconnection system
CN116955270A