Business processing method, device and equipment for multiple processing cores
By adopting a business processing method in a multi-core system and using the sharing of operation information between cores, the problem of poor effectiveness of speculative means in the multi-core system is solved, and the effect of improving speculative accuracy and efficiency is achieved.
Patent Information
- Application Number
- CN202311794035.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
In multi-core systems, cache resources and bandwidth resources often become bottlenecks in peak performance, and speculative methods in single-core systems are not effective in multi-core systems.
By introducing a business processing method in the multi-core system, the first processing core obtains its operation information and receives operation information from the second processing core with a faster operation speed. When both determine that the business is partially similar, the first processing core runs its business according to the operation information of the second processing core to improve the accuracy and efficiency of speculation.
This method significantly improves the accuracy and efficiency of speculation in multi-core systems, improves overall business processing performance, and reduces the power consumption and overhead caused by traditional speculation failure.
Smart Images

Figure CN120196427A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic technologies, and particularly to a service processing method, apparatus, and device for multiple processing cores. Background Art
[0002] During the processing of services, the cache resources and bandwidth resources of a multi-core system often become bottlenecks at peak performance. Currently, in a single-core system, speculative means such as front-end prefetching and back-end prefetching are usually adopted to improve the performance of service processing. However, this speculative means will incur additional bandwidth overhead and cache overhead. For example, incorrect front-end prefetching will cause the incorrect path to fetch invalid instructions and generate invalid memory accesses, and back-end prefetching will generate a large number of memory access requests for speculative addresses, thus polluting the cache and causing a bandwidth bottleneck. Therefore, the speculative means adopted by the above single-core system has poor effects when applied to a multi-core system. Summary of the Invention
[0003] This application provides a service processing method, apparatus, and device for multiple processing cores, which are used to improve the accuracy of speculation in a multi-core system, and thus improve the performance of service processing.
[0004] To achieve the above objective, the embodiments of this application adopt the following technical solutions:
[0005] In a first aspect, a service processing method for multiple processing cores is provided. The multiple processing cores may be multiple processing cores in a multi-core system. The method includes: a first processing core obtains first running information, where the first running information is obtained by the first processing core when running a first service, and the first processing core is a following core or a slave core; the first processing core receives second running information from a second processing core, where the second running information is obtained by the second processing core when running a second service, the second processing core is a processing core with a faster running speed among the multiple processing cores, and the second processing core is referred to as a leading core or a master core, and the running speed of the first processing core is less than that of the second processing core; when it is determined according to the first running information and the second running information that there is at least partial identity between the first service and the second service, the first processing core runs the first service according to the second running information.
[0006] In the above technical solution, since the second running information is the running information actually obtained by the second processing core, and there is partial identity between the first service and the second service, that is, there is repetition and similarity between the running of the first service and the running of the second service. In this way, when the first processing core runs the first service according to the second running information, the accuracy and efficiency of speculation can be greatly improved, thereby improving the overall performance of the multiple processing cores in processing services and reducing the power consumption overhead caused by traditional speculation failures.
[0007] In a possible implementation of the first aspect, the first running information and the second running information include at least one of the following: data address information, instruction address information, branch address information, jump information, or value prediction information. In the above possible implementation, when the second running information includes data address information, the first processing core runs the first service according to the second running information, which can greatly reduce the data cache miss rate; when the second running information includes instruction address information, the first processing core runs the first service according to the second running information, which can greatly reduce the instruction cache miss rate; when the second running information includes branch address information, the first processing core runs the first service according to the second running information, which can greatly reduce the branch prediction failure rate; when the second running information includes jump information or value prediction information, the first processing core runs the first service according to the second running information, which can greatly improve the accuracy and efficiency of speculation, thereby improving the overall performance of the service.
[0008] In a possible implementation of the first aspect, the first running information includes a plurality of first address information, and the second running information includes a plurality of second address information; the first processing core determines that there is at least partial identity between the first service and the second service according to the first running information and the second running information, including: when the number of identical address information existing in the plurality of first address information and the plurality of second address information is greater than a preset threshold, the first processing core determines that there is at least partial identity between the first service and the second service; or, when the number of identical address offset values existing in the address offset values of the plurality of second address information relative to the plurality of first address information is greater than a preset threshold, the first processing core determines that there is at least partial identity between the first service and the second service, where the address offset value of the plurality of second address information relative to the plurality of first address information includes the address offset value between each second address information and each first address information. In the above possible implementation, when determining whether there is at least partial identity between the first service and the second service, the first processing core can be determined according to the number of identical address information in the plurality of first address information included in the first running information and the plurality of second address information included in the second running information, or according to the number of identical address offset values existing in the address offset values of the plurality of second address information relative to the plurality of first address information, so as to cover more scenarios of dynamic address allocation, thereby improving the coverage of benefits.
[0009] In a possible implementation of the first aspect, the method further includes: the first processing core receives security information from the second processing core, and verifies the security of the second processing core according to the security information. Optionally, the security information includes at least one of the following: address space identifier (ASID), virtual machine identifier (VMID), or privilege level information. Optionally, the security information can be synchronously transmitted together with the data packet of the second running information to reduce the complexity of transmitting the security information and the second running information; or, the security information is asynchronously transmitted through a data packet in a specific format, that is, the security information and the second running information are transmitted through different data packets, so as to reduce the length of the data packet through asynchronous transmission. In the above possible implementation, by performing a security check on the second processing core, the first processing core can improve the security of speculation.
[0010] In a possible implementation of the first aspect, the first processing core receives the second running information from the second processing core, including: the first processing core receives the second running information from the second processing core through a bus. Wherein, the bus can be a multiplexed bus or a newly added bus.
[0011] In a possible implementation of the first aspect, the operating frequency of the first processing core is different from that of the second processing core. The first processing core receives the first running information from the second processing core through a bus, including: the asynchronous interface of the first processing core receives the second running information from the second processing core through the bus, and performs asynchronous logic processing on the second running information. In the above possible implementation, when the frequency of the first processing core is different from that of the second processing core, synchronous processing can be performed to ensure that the first processing core can still run the first service according to the second running information.
[0012] In a possible implementation of the first aspect, the first processing core includes a first cache. Before the first processing core runs the first service according to the second running information, the method further includes: the first processing core writes the second running information into the first cache as prefetch information for the first service. In the above possible implementation, the accuracy and efficiency of speculation can be greatly improved, thereby improving the overall performance of multiple processing cores in processing services and reducing the power consumption overhead caused by traditional speculation failures.
[0013] In a possible implementation of the first aspect, the second processing core includes a memory access control module and a second cache, and the method further includes: the memory access control module sends second running information to the second cache, where the second running information includes address information accessed by the memory access control module during the process of the second processing core running the second service; the second cache receives and filters duplicate addresses in the second running information, and sends the second running information after filtering the duplicate addresses to the first processing core. Optionally, the second cache includes a cache queue, and specifically, the cache queue receives and filters duplicate addresses in the second running information, and sends the second running information after filtering the duplicate addresses to the first processing core. In the above possible implementation, by receiving and filtering the second running information of the first processing core sent by the second cache, the transmission efficiency of the second running information can be improved; in addition, the action of receiving and filtering duplicate addresses in the second running information can be specifically executed by the cache queue in the second cache, and the cache queue is a multi-in single-out queue, and can also achieve the effect of balancing the bandwidth difference between input and output.
[0014] In a second aspect, a service processing device is provided, and the device includes a first processing core and a second processing core; the first processing core is configured to: obtain first running information, where the first running information is obtained by the first processing core running the first service; receive second running information from the second processing core, where the second running information is obtained by the second processing core running the second service; when it is determined according to the first running information and the second running information that there is at least partial identity between the first service and the second service, run the first service according to the second running information; the second processing core is configured to: provide the second running information to the first processing core.
[0015] In a possible implementation of the second aspect, the first running information and the second running information include at least one of the following: data address information, instruction address information, branch address information, jump information, or value prediction information.
[0016] In a possible implementation of the second aspect, the first running information includes a plurality of first address information, the second running information includes a plurality of second address information, and the first processing core is further configured to: when the number of identical address information existing in the plurality of first address information and the plurality of second address information is greater than a preset threshold, determine that there is at least partial identity between the first service and the second service; or, when the number of identical address offset values existing in the address offset values of the plurality of second address information relative to the plurality of first address information is greater than a preset threshold, the first processing core determines that there is at least partial identity between the first service and the second service, where the address offset values of the plurality of second address information relative to the plurality of first address information include the address offset values between each second address information and each first address information.
[0017] In a possible implementation of the second aspect, the first processing core is further configured to: receive security information from the second processing core, and verify the security of the second processing core according to the security information. Optionally, the security information includes at least one of the following: an address space identifier (ASID), a virtual machine identifier (VMID), and privilege level information.
[0018] In a possible implementation of the second aspect, the apparatus further includes: a bus, configured to transmit second operation information from the second processing core to the first processing core.
[0019] In a possible implementation of the second aspect, the operating frequency of the first processing core is different from that of the second processing core; the first processing core includes: an asynchronous interface, configured to receive second operation information from the second processing core through the bus and perform asynchronous logic processing on the second operation information.
[0020] In a possible implementation of the second aspect, the first processing core includes a first cache; the first processing core is further configured to write the second operation information as prefetch information for a first service into the first cache.
[0021] In a possible implementation of the second aspect, the second processing core includes a memory access control module and a second cache; the memory access control module is configured to send the second operation information to the second cache, where the second operation information includes address information accessed by the memory access control module during the second processing core running a second service; the second cache is configured to filter duplicate addresses in the second operation information.
[0022] In a third aspect, there is provided an electronic device, which includes a memory and at least one processor. The memory is configured to store computer instructions, and the at least one processor includes a plurality of processing cores and is configured to execute the computer instructions so that the electronic device implements the service processing method provided in the first aspect or any possible implementation of the first aspect.
[0023] In yet another aspect of the present application, there is provided a computer-readable storage medium, in which a computer program or instruction is stored. When the computer program or instruction is run, the service processing method provided in the first aspect or any possible implementation of the first aspect is implemented.
[0024] In yet another aspect of the present application, there is provided a computer program product, which includes: a computer program (which may also be referred to as code or instruction). When the computer program is run, the computer is caused to execute the service processing method provided in the first aspect or any possible implementation of the first aspect.
[0025] Understandably, for any of the service processing devices, electronic devices, computer-readable storage media, and computer program products provided above, the beneficial effects they can achieve can be correspondingly referred to the beneficial effects in the service processing method provided above, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 FIG. 6 is a schematic diagram of the performance when applying speculative means to a single-core system and a multi-core system according to an embodiment of the present application;
[0027] Figure 2 FIG. 10 is a schematic diagram of a plurality of processing cores processing homogeneous services according to an embodiment of the present application;
[0028] Figure 3 FIG. 14 is a schematic structural diagram of a multi-core system according to an embodiment of the present application;
[0029] Figure 4 FIG. 18 is a schematic flowchart of a service processing method for a plurality of processing cores according to an embodiment of the present application;
[0030] Figure 5 FIG. 22 is a schematic diagram of a plurality of processing cores processing services according to an embodiment of the present application;
[0031] Figure 6 FIG. 26 is a schematic diagram of a service learning mode according to an embodiment of the present application;
[0032] Figure 7 FIG. 30 is a schematic diagram of a data structure according to an embodiment of the present application;
[0033] Figure 8 FIG. 34 is a schematic diagram of another data structure according to an embodiment of the present application;
[0034] Figure 9 FIG. 38 is a schematic diagram of a service processing device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The following will elaborate on the fabrication and use of each embodiment in detail. However, it should be understood that many applicable inventive concepts provided by the present application can be implemented in a variety of specific environments. The specific embodiments discussed merely illustrate the specific ways to implement and use the present application and this technology, and do not limit the scope of the present application.
[0036] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by those of ordinary skill in the art.
[0037] Each circuit or other component may be described as or referred to as "configured to" perform one or more tasks. In such cases, "configured to" is used to imply structure by indicating that the circuit / component includes the structure (e.g., circuitry) that performs the one or more tasks during operation. Thus, even when the specified circuit / component is not currently operational (e.g., not turned on), the circuit / component can still be referred to as configured to perform the task. Circuits / components used in conjunction with the phrase "configured to" include hardware, such as circuitry that performs the operations, etc.
[0038] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. In the present application, "at least one" means one or more, and "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or similar expressions thereof refer to any combination of these items, including any combination of single item (s) or plural item (s). For example, at least one (item) of a, b, or c may represent: a, b, c, a and b, a and c, b and c, a, b, and c; where a, b, and c may be single or multiple.
[0039] In the embodiments of the present application, terms such as "first" and "second" are used to distinguish objects with similar names, functions, or roles. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order. The term "coupled" is used to indicate an electrical connection, including being directly connected through a wire or connection terminal or indirectly connected through other devices. Therefore, "coupled" should be regarded as a generalized electronic communication connection.
[0040] It should be noted that in the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0041] During the processing of a service, the cache resources and bandwidth resources of a multi-core system often become bottlenecks at peak performance. In a single-core system, speculative means such as front-end prefetching and back-end prefetching are usually adopted to improve the performance of service processing. However, this speculative means will incur additional bandwidth overhead and cache overhead. For example, front-end prefetching errors will cause the wrong path to fetch invalid instructions and generate invalid memory accesses, and back-end prefetching will generate a large number of memory access requests with speculative addresses, thus polluting the cache and causing bandwidth bottlenecks. The speculative means adopted by the above single-core system has poor effects when applied to a multi-core system. To solve this problem, the following embodiments provide corresponding technical means.
[0042] When speculative means are adopted in a single-core system and a multi-core system, the corresponding performance is usually evaluated by the following indicators: coverage, accuracy, and timeliness. Coverage: It refers to the proportion of all the data used by a processing core (or called a core, also called a processor core) that is prefetched into the cache. The higher the coverage, the better. Accuracy: It refers to the proportion of the prefetched data accessed by the processing core in all the prefetched data. The higher the accuracy, the better. If the accuracy is too low, the useless prefetched data will pollute the cache and occupy the memory bandwidth. Timeliness: It refers to the timing when the data to be accessed is prefetched into the cache. If the data is just prefetched into the cache when it is accessed, the timeliness is good. If the data has not been prefetched into the cache or has been prefetched into the cache too early when it is accessed, the timeliness is relatively poor. Figure 1 The following shows the illustration of the above three indicators when speculative means are applied to a single-core system and a multi-core system. Among them, when speculative means are adopted in a multi-core system, the above accuracy and coverage often cannot be satisfied simultaneously. Generally, only by reducing the performance gain can the accuracy meet the multi-core requirements.
[0043] Based on this, the embodiments of the present application provide a service processing method for multiple processing cores. When there is at least partial identity (or called homogeneous services or homogeneous tasks) in the services processed by multiple processing cores, the method can obtain the running information (or called speculative auxiliary information) obtained from the operation of a certain processing core (for example, a core with higher performance) shared among the cores, and based on the shared information, run the services of other cores, so as to greatly improve the performance of other cores, thereby improving the accuracy of speculation and further improving the performance of service processing. Exemplarily, as Figure 2 shown Figure 2Multiple arrows in [the figure] represent multiple processing cores, and the length of each arrow represents the operating efficiency of the corresponding processing core. Among them, in the first stage, multiple processing cores start to run homogeneous services and their corresponding operating efficiencies are close. In the second stage, the operating efficiency of the first processing core is relatively high. At this time, the first processing core shares operating information with other processing cores. After other cores run services according to the shared operating information in the third stage, the finally achieved operating efficiency is close to that of the first processing core, thereby improving the overall performance of the system. At least part of the services involved above are the same or homogeneous, including at least part of the same data used in different task execution operations. For example, there is at least part of the same data address information, instruction address information, branch address information, jump information, or value prediction information involved in the tasks.
[0044] The service processing method provided by the embodiments of this application can be applied to a multi-core system with multiple processing cores. The structure of this multi-core system will be introduced and described below.
[0045] Figure 3 It is a schematic structural diagram of a multi-core system provided by the embodiments of this application. The multi-core system includes multiple processing cores, and the multiple processing cores can be used to deploy homogeneous services. For example, the multiple processing cores can be applied to a server or a device for performing high performance computing (HPC). As Figure 3 shown, the multiple processing cores can be coupled into various different structures, and the various different structures include but are not limited to: a ring structure, a mesh structure, a cross structure, or a combination of the above structures, etc. A to P in the figure represent different processing cores. In one example, the multiple processing cores are coupled into a ring structure through a ring bus. For example, processing cores A to G are coupled into a ring structure. In another example, the multiple processing cores are coupled into a mesh structure through a mesh bus. For example, processing cores A to P are coupled into a mesh structure. In yet another example, the multiple processing cores are coupled into a crossbar structure through a crossbar bus. For example, processing cores A to E are coupled into a crossbar structure.
[0046] In this multi-core system, any one of the multiple processing cores and the connected processing cores can be directly coupled, or coupled through a bus, or coupled through a router, etc. The embodiments of this application do not make specific limitations in this regard. Only the case where the multiple processing cores are coupled through a bus is taken as an example for illustration in the above examples, and the above examples do not constitute a limitation to the embodiments of this application.
[0047] In addition, the multiple processing cores may include multiple different processing cores of the same processor, or the multiple processing cores include processing cores of multiple different processors. Optionally, the structures and the sizes of the respective hardware resources of any two of the multiple processing cores may be the same. For example, the cache sizes of the multiple processing cores are the same.
[0048] Furthermore, the multi-core system may further include a memory coupled to the multiple processing cores. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. Exemplarily, the RAM may be a static RAM (SRAM), a dynamic random access memory (DRAM), a synchronous DRAM (SDRAM), and a double data rate synchronous DRAM (DDR SDRAM), etc. The memory is not shown in the figure.
[0049] The above multi-core system can be an electronic device, or a system on a chip (SoC) applied to an electronic device, or a chipset including multiple chips, or a module including the SoC or the chipset. The electronic device can be a terminal device or a server. Optionally, the electronic device includes, but is not limited to: mobile phone, tablet computer, laptop computer, desktop computer, palmtop computer, ultra-mobile personal computer (UMPC), mobile internet device (MID), netbook, camera, camera, wearable device (such as smart watch and smart bracelet, etc.), vehicle-mounted device (such as car, bicycle, electric vehicle, airplane, ship, train, high-speed train, etc.), virtual reality (VR) device, augmented reality (AR) device, wireless terminal in industrial control, smart home device (such as refrigerator, TV, air conditioner, electricity meter, etc.), smart robot, workshop device, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, or wireless terminal in smart home, flight device (such as smart robot, hot air balloon, drone, airplane), etc.
[0050] Figure 4 FIG. 4 is a schematic flowchart of a service processing method for multiple processing cores provided by an embodiment of the present application. The method can be applied to the multi-core system provided above. The multiple processing cores include a first processing core and a second processing core. The method includes the following steps.
[0051] S201a: The first processing core obtains first operation information, where the first operation information is obtained by the first processing core running a first service.
[0052] S201b: The second processing core obtains second operation information, where the second operation information is obtained by the second processing core running a second service.
[0053] Among them, the second processing core can be a processing core with a relatively fast operating speed among the multiple processing cores, and the second processing core can be referred to as the leading core or the master core. The first processing core can be any one of the multiple processing cores other than the second processing core, and the operating speed of the first processing core is less than that of the second processing core. The first processing core can be referred to as the following core or the slave core.
[0054] In addition, the first service and the second service can be exactly the same service, or partially the same service (or referred to as similar or analogous services). The partially same service can mean that there are some same data and / or instructions that two tasks need to access during operation. For example, the first service and the second service can be two parallel computing services in the same high-performance computing. The operation of the above first service and second service has a large degree of repetition and similarity. For example, the executed instructions, data, or branches, etc. have repetition and similarity.
[0055] Optionally, the second processing core as the leading core can be preset or specified. Exemplarily, the leading core is determined according to the physical positions of the multiple processing cores. For example, the first core in the physical position is set as the leading core; or, the leading core is the processing core corresponding to the logical core determined by the basic input output system (BIOS) after the multi-core system is powered on; or, the leading core is specified by the software running on the multi-core system. The embodiments of the present application do not specifically limit the method of specifying the leading core.
[0056] In a possible embodiment, when the first processing core receives an execution instruction for the first service, the first processing core runs the first service and obtains first running information according to the running result of the already-run part of the first service; similarly, when the second processing core receives an execution instruction for the second service, the second processing core runs the second service and obtains second running information according to the running result of the already-run part of the second service. Among them, the operating speed of the second processing core is greater than that of the first processing core, or it can be said that the running progress of the second processing core is greater than that of the first processing core. In this way, the second processing core sharing the obtained second running information with the first processing core can improve the speculative performance of the first processing core, and further improve the performance of service processing.
[0057] Optionally, the first running information and the second running information include at least one of the following obtained during the execution of the service: data address information, instruction address information, branch address information, jump information, or value prediction information. Among them, the data address information (also known as memory access information) is the address of the actually accessed data in the memory. The instruction address information is the address of the actually accessed instruction in the memory. The branch address information is the address of the actually executed branch in the memory. The jump information may include a jump direction and a jump destination. The jump direction is used to indicate the actual jump direction of the jump instruction (i.e., jump or not jump), and the jump destination is the destination address after the actual jump. The value prediction information refers to the value obtained through prediction when the memory access is not completed, and the value may be the result of intermediate execution. The memory may be on the same or different chips as one or more of the above processing cores, and may be a volatile or non-volatile memory, which is not limited in this embodiment.
[0058] Exemplarily, the first running information includes the data address information obtained by the first processing core during the execution of the first service, and the second running information includes the data address information obtained by the second processing core during the execution of the second service; or, the first running information includes the instruction address information obtained by the first processing core during the execution of the first service, and the second running information includes the instruction address information obtained by the second processing core during the execution of the second service; or, the first running information includes the branch address information obtained by the first processing core during the execution of the first service, and the second running information includes the branch address information obtained by the second processing core during the execution of the second service.
[0059] S202a: The second processing core sends the second running information to the first processing core.
[0060] In a possible embodiment, during the execution of the second service, the second processing core may send the second running information obtained by running the second service to the first processing core one or more times. Exemplarily, the second processing core sends the second running information obtained by running the second service to the first processing core periodically or aperiodically. For example, during the execution of the second service, the second processing core periodically sends the second running information obtained by running the second service in each cycle to the first processing core; or, the second processing core sends the second running information obtained by running the second service to the first processing core when the storage state of the local cache meets certain conditions.
[0061] Optionally, as Figure 5As shown in the figure, the second processing core includes a memory access control module and a cache (referred to as the second cache in this article). This access control module can also be called a load store unit (LSU), and this second cache can be a prefetch input buffer (inbuf). Further, the second processing core may also include a level 1 cache (L1), a hardware prefetch (HWP), a prefetch to level 1 cache (PFL1) interface, a prefetch to level 2 cache (PFL2) interface, a prefetch to level 3 cache (PFL3) interface, and an asynchronous interface, etc. In the figure, an example is given where the second processing core includes an LSU, an inbuf, an L1, a PFL1 interface, a PFL2 interface, a PFL3 interface, and an asynchronous interface, and the LSU and L1 are represented as LSU+L1. Optionally, as Figure 5 shown, other processing cores (slave cores), such as the first processing core, may have a similar structure to the second processing core.
[0062] In a possible example, in combination with Figure 5 shown, after the second processing core obtains the second running information, the memory access control module (for example, LSU) may send the second running information to the second cache (for example, inbuf). The second running information includes the address information accessed by the memory access control module during the process of the second processing core executing the second service. When the second cache receives the second running information, the second cache may filter the duplicate addresses in the second running information and then send it to the first processing core. For example, the second cache sends the second running information through the PFL3 interface. Exemplarily, the second cache includes a cache queue. Specifically, the cache queue may receive the second running information, filter the duplicate addresses in the second running information, and then send it through the PFL3 interface. Among them, the cache queue may be a multi-in and single-out cache queue, which can be used not only to filter duplicate addresses but also to balance the bandwidth difference between input and output to improve the transmission efficiency of the second running information.
[0063] Exemplarily, when the second running information sent by the second processing core to the first processing core may include multiple address information (referred to as the second address information in this article), the multiple second address information may be the memory access request addresses where the level 1 cache misses, or a subset or the whole set of any regular level 1 memory access addresses, etc. The embodiments of the present application do not make specific limitations on this.
[0064] Further, the second processing core may also send security information to the first processing core, and the security information can be used to verify the security of the second processing core. The security information may be identification information and / or privilege information assigned to the second processing core, etc. Exemplarily, the security information includes at least one of the following: address space identifier (ASID), virtual machine identifier (VMID), privilege level information.
[0065] The above security information may be synchronously transmitted with the data packet of the second running information to reduce the complexity of the transmission of the security information and the second running information; or, the security information is asynchronously transmitted through a data packet in a specific format, that is, the security information and the second running information are transmitted through different data packets, so as to reduce the length of the data packet through asynchronous transmission. Exemplarily, the second processor core may send the security information in advance so that the first processor core can complete the security verification. For example, the second processor core may send the security information when the device is powered on.
[0066] S202b: The first processing core receives the second running information from the second processing core.
[0067] Optionally, the multiple processing cores are coupled by a bus. The bus may be the original bus in the multi-core system (i.e., the bus in the multi-core system is reused), or a bus independently set in this application (for example, a newly added bus). The embodiments of this application do not make specific limitations in this regard. The second processing core may send the second running information to the first processing core through the bus; correspondingly, the first processing core receives the second running information from the second processing core through the bus.
[0068] Optionally, the second running information may be transmitted in a compressed format or may not be transmitted in a compressed format. The embodiments of this application do not make specific limitations in this regard.
[0069] In a possible embodiment, the frequency of the first processing core is different from the frequency of the second processing core. Optionally, the frequency of the first processing core is less than the frequency of the second processing core. At this time, the asynchronous interface of the first processing core may perform asynchronous logic processing during the process of receiving the second running information, so that the frequency of the second running information is the same as the frequency of the first processing core.
[0070] Optionally, as Figure 5As shown, the structure of the first processing core is similar to that of the second processing core. The first processing core may include: a memory access control module and a first cache, and the first cache may be a prefetch input cache inbuf. Further, the first processing core may also include other modules such as a level-1 cache L1, a hardware prefetch unit HWP, a PFL1 interface, a PFL2 interface, a PFL3 interface, and an asynchronous interface. In addition, the PFL3 interface of the first processing core and the PFL3 interface of the second processing core, as well as the asynchronous interface of the first processing core and the asynchronous interface of the second processing core, may be coupled via a bus.
[0071] In a possible example, in combination with Figure 5 As shown, the first processing core receives second operation information from the second processing core via a bus, which may include: the asynchronous interface of the first processing core receives the second operation information from the second processing core via the bus and performs asynchronous logic processing on the second operation information to achieve timing synchronization.
[0072] It can be understood that the multiple processing cores may include multiple slave cores, that is, in addition to the first processing core, the multiple processing cores may also include other slave cores similar to the first processing core. The second processing core (i.e., the master core) may send the second operation information to each slave core. Figure 5 In
[0073] Further, when the second processing core also sends security information to the first processing core, the first processing core may receive the security information and verify the security of the second processing core according to the security information. When it is determined that the second processing core is secure according to the security information, the first processing core may execute step S203 below. Alternatively, the security information transmission and security verification may also be performed in advance.
[0074] S203: When it is determined that there is at least partial identity between the first service and the second service according to the first operation information and the second operation information, the first processing core runs the first service according to the second operation information. Optionally, when the first service and the second service are different, the first processing core does not run the first service according to the second operation information, that is, at this time, the two processing cores are not running homogeneous tasks.
[0075] Among them, determining that the first service and the second service have at least partial identity can also be referred to as determining to enter the homogeneous (cluster learning optimization, CLO) mode or the service sharing and learning mode, that is, the first processing core can run or process the first service through the second running information shared by the second processing core, or it can be said that the first processing core learns the way the second processing core runs or processes the second service to process the first service. Optionally, the opening or closing of the above homogeneous mode can be controlled by the first processing core, or a software routing can be designed to be opened or closed by the upper-layer operating system. The embodiments of the present application do not make specific limitations on this.
[0076] In one implementation, the first running information includes a plurality of first address information, and the second running information includes a plurality of second address information. The plurality of first address information and the plurality of second address information can be data address information, or instruction address information, or branch address information, or include two of the above three address information. Exemplarily, the first running information includes a plurality of first data address information, and the second running information includes a plurality of second data address information; or, the first running information includes a plurality of first instruction address information, and the second running information includes a plurality of second instruction address information; or, the first running information includes a plurality of first branch address information, and the second running information includes a plurality of second branch address information; the first running information includes a plurality of first data address information and a plurality of first instruction address information, and the second running information includes a plurality of second data address information and a plurality of second instruction address information; or, the first running information includes a plurality of first instruction address information and a plurality of first branch address information, and the second running information includes a plurality of second instruction address information and a plurality of second branch address information.
[0077] Exemplarily, according to the different information included in the second running information, the above service sharing and learning can include the sharing and learning of different information. For example, as Figure 6 shown, when the plurality of second address information included in the second running information is data address information, the service sharing and learning includes data cache sharing and learning; when the plurality of second address information included in the second running information is instruction address information, the service sharing and learning includes instruction cache sharing and learning; when the plurality of second address information included in the second running information is branch address information, the service sharing and learning includes branch prediction sharing and learning. Figure 6 In [description], the processing core 0 among the plurality of processing cores is taken as the leading core, and other processing cores (such as processing core 1 and processing core 2, etc.) are taken as slave cores for illustration, and the main pipelines of each processing core are shown.
[0078] In a possible embodiment, determining that at least part of the first service and the second service are the same according to the first operation information and the second operation information includes: when the number of identical address information among the multiple first address information and the multiple second address information is greater than a first preset threshold, determining that at least part of the first service and the second service are the same.
[0079] Exemplarily, the multiple second address information included in the second operation information may be arranged according to a certain data structure. When the first processing core determines that a certain first address information among the multiple first address information included in the first operation information is the same as a certain second address information, the first processing core records the identical first address information in the data structure, counts the identical address information, and determines that at least part of the first service and the second service are the same when the number of identical address information is greater than the first preset threshold. For example, the data structure may be a homogeneous training table (CLO train table, CTT), as Figure 5 shown. The CTT may be a naturally wound storage structure that is refreshed by continuous writing. Specifically, as Figure 7 shown, the CTT may include n (n is a positive integer) rows and two columns. Each row in the first column includes a second address information, and each row in the second column can be used to store a first address information. When the first processing core finds a first address information identical to a certain second address information among the multiple first address information, the first processing core may store the first address information in the corresponding row and color the row. Figure 7 The n second address information are represented as VA1 to VAn, the n first address information are represented as TA1 to TAn, and the colored row is represented by filling.
[0080] In another possible embodiment, when the number of identical address offset values among the address offset values corresponding to the multiple first address information and the multiple second address information is greater than a second preset threshold, the first processing core determines that at least part of the first service and the second service are the same. Wherein, the address offset values corresponding to the multiple first address information and the multiple second address information include the address offset values between the multiple second address information and each first address information.
[0081] Exemplarily, the multiple second address information included in the second operation information may be arranged according to the homogeneous training table CTT, as Figure 5 shown. For each first address information among the multiple first address information, the first processing core sequentially determines the address offset value between the first address information and each second address information, records each address offset value and the count corresponding to the address offset value (the count can be obtained through a counter) through the CTT, and colors the row where the count of the address offset value is greater than the second preset threshold. Specifically, as Figure 8As shown, the CTT may include n (n is a positive integer) rows and three columns. Each row in the first column includes a second address information. Each row in the second column can be used to store an address offset value. Each row in the third column can be used to record the count of the address offset value corresponding to the row. When the count of the address offset value of a certain row is greater than a second preset threshold, that row is colored. For example, the row with the maximum count is colored, that is, a fixed offset value is learned. Figure 8 Among them, the n second address information are represented as VA1 to VAn, the n address offset values are represented as OF1 to OFn, the corresponding counts are represented as CT1 to CTn, and the colored row is represented by filling.
[0082] The above first preset threshold and second preset threshold can be set in advance, and the embodiments of the present application do not limit the specific values.
[0083] In addition, in a possible embodiment, when the first processing core determines that there is at least partial identity between the first service and the second service, the first processing core runs the first service according to the second running information, which may include: the first processing core takes the second running information received from the second processing core subsequently as prefetch information, and continues to run the first service according to this prefetch information. Exemplarily, when the second running information includes data address information, the first processing core runs the first service according to this data address information; when the second running information includes instruction address information, the first processing core runs the first service according to this instruction address information; when the second running information includes branch address information, the first processing core runs the first service according to this branch address information; when the second running information includes jump information, the first processing core runs the first service according to this jump information; when the second running information includes value prediction information, the first processing core runs the first service according to this value prediction information.
[0084] Optionally, before the first processing core runs the first service according to the second running information, the second running information can also be written into the first cache as prefetch information for the first service, and instructions and / or data are obtained from the memory according to this prefetch information, so as to run the first service according to the obtained instructions and / or data.
[0085] Exemplarily, as Figure 5 shown, after the first processing core enters the homogenization mode according to the CTT recognition, the second running information shared by the second processing core is sent to the main pipeline of the hardware prefetch unit (HWP). After the main pipeline of the hardware prefetch unit goes through processing such as caching and address translation, it is sent to the cache with the target of level 1, level 2, and level 3 according to the policy setting, that is, the data and / or instructions to be used in the future are filled into the cache in advance, and then the first service is run based on the data and / or instructions in the cache to improve the performance of the first processing core running the first service.
[0086] In an embodiment of the present application, the first processing core obtains the first running information of the first service, receives the second running information of the second service running on the second processing core, and when it is determined according to the first running information and the second running information that there is a partial similarity between the first service and the second service, runs the first service according to the second running information. Since the second running information is the running information actually obtained by the second processing core, and there is a partial similarity between the first service and the second service, that is, there is repetition and similarity between the running of the first service and the running of the second service. In this way, when the first processing core runs the first service according to the second running information, the accuracy and efficiency of speculation can be greatly improved, thereby improving the overall performance of multiple processing cores in processing services and reducing the power consumption overhead caused by traditional speculation failures. Specifically, when the first processing core enters the data cache sharing learning, the first processing core runs the first service according to the second running information actually sent by the second processing core, which can greatly reduce the data cache miss rate; when the first processing core enters the instruction cache sharing learning, the first processing core runs the first service according to the second running information actually sent by the second processing core, which can greatly reduce the instruction cache miss rate; when the first processing core enters the branch prediction sharing learning, the first processing core runs the first service according to the second running information actually sent by the second processing core, which can greatly reduce the branch prediction failure rate.
[0087] In addition, the first processing core runs the first service according to the second running information when it is determined that there is at least a partial similarity between the first service and the second service, and does not run the first service according to the second running information when it is determined that the first service and the second service are not similar. In this way, the first processing core can automatically determine whether to run the first service according to the second running information (that is, automatically determine whether it belongs to the profitable scenario), thereby avoiding mis-triggering and negative benefits in non-target scenarios.
[0088] In addition, when determining whether there is at least a partial similarity between the first service and the second service, the first processing core can directly determine according to the number of identical address information among the multiple first address information included in the first running information and the multiple second address information included in the second running information, or determine according to the number of identical address offset values among the address offset values corresponding to the multiple first address information and the multiple second address information, so as to cover more scenarios of dynamic address allocation, thereby improving the coverage of benefits.
[0089] Furthermore, the above embodiment is described by taking the running speed of the second processing core being higher than that of the first processing core as an example. In fact, this setting is not used to limit the application scenarios of the embodiment. In practical applications, multiple processing cores can have the same or different capabilities or processing speeds, and it is not limited which one has stronger capabilities. As long as the running information obtained after one processing core executes a task can be shared with another processing core to execute the technical solution mentioned in this embodiment. For example, still takingFigure 5 For example, the functions of the master core and the slave core can be swapped. The first processing core (slave core) can send its running information to the second processing core (master core) so that the second processing core can execute similar Figure 4 functions as those implemented in the process of the above embodiment. For example, due to software scheduling or user selection, etc., the first processing core may first execute at least part of the functions of a task and obtain running information, and this running information can be shared with the second processing core for continuing to execute similar homogeneous tasks in order to achieve a similar effect. It can be understood that the master core and the slave core mentioned in this embodiment, as well as the differences between these two cores, are only an applicable scenario and are not used to limit the technical solution.
[0090] Based on this, an embodiment of the present application further provides a service processing device, which can be applied to a multi-core system. As Figure 9 shown, the device includes: a first processing core and a second processing core. In the embodiment of the present application, the first processing core can be used to execute S201, S202b, S203 in the above method embodiment, receive security information from the second processing core and verify the security of the second processing core according to this security information, and / or other steps described herein, etc.; the second processing core can be used to execute step S201b in the above method embodiment, the step of sending security information to the first processing core, and / or other steps described herein, etc.
[0091] It can be understood that all relevant contents about the specific structures of the first processing core and the second processing core, as well as each step involved in the above method embodiment, can be cited into the embodiment of this service processing device, and the embodiment of the present application will not elaborate herein.
[0092] In the embodiment of the present application, since the second running information is the running information actually obtained by the second processing core, and there are some similarities between the first service and the second service, that is, the running of the first service and the running of the second service have repeatability and similarity. In this way, when the first processing core runs the first service according to the second running information, it can greatly improve the accuracy and efficiency of speculation, thereby improving the overall performance of multiple processing cores in processing services and reducing the power consumption overhead caused by traditional speculation failures.
[0093] On the other hand, the present application further provides an electronic device, which includes a memory and at least one processor. The memory is used to store computer instructions, and the at least one processor includes multiple processing cores and is used to execute the computer instructions so that the electronic device can implement any one of the service processing methods for multiple processing cores provided above. Optionally, the at least one processor includes the service processing device provided above.
[0094] It can be understood that all relevant content of each step involved in the above method embodiments can be cited in the embodiments of the service processing device and the embodiments of the electronic device, and the embodiments of the present application will not be elaborated herein.
[0095] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.
[0096] The units described as separate components may or may not be physically separated. The components shown as units may be a physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0097] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. The readable storage medium can include: various media such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs that can store program codes. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product.
[0098] In another embodiment of the present application, a readable storage medium is further provided. Computer-executable instructions are stored in the readable storage medium. When a device (which can be a single-chip microcomputer, a chip, etc.) or a processor executes the steps in the above method embodiments.
[0099] In yet another embodiment of the present application, a computer program product is further provided. The computer program product includes computer instructions, and the computer instructions are stored in a readable storage medium; at least one processor of the device can read the computer instructions from the readable storage medium, and at least one processor executes the computer instructions to enable the device to execute the steps in the above method embodiments.
[0100] Finally, it should be noted that the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A service processing method for multiple processing cores, characterized in that The method includes: The first processing core obtains first running information, where the first running information is obtained by the first processing core running a first service; The first processing core receives second running information from a second processing core, where the second running information is obtained by the second processing core running a second service; When it is determined according to the first running information and the second running information that there is at least partial identity between the first service and the second service, the first processing core runs the first service according to the second running information.
2. The method according to claim 1, wherein The first running information and the second running information include at least one of the following: data address information, instruction address information, branch address information, jump information, or value prediction information.
3. The method according to claim 1 or 2, characterized in that The first running information includes a plurality of first address information, and the second running information includes a plurality of second address information; The determining that there is at least partial identity between the first service and the second service according to the first running information and the second running information includes: When the number of identical address information existing in the plurality of first address information and the plurality of second address information is greater than a preset threshold, the first processing core determines that there is at least partial identity between the first service and the second service; or, When the number of identical address offset values existing in the address offset values of the plurality of second address information relative to the plurality of first address information is greater than a preset threshold, the first processing core determines that there is at least partial identity between the first service and the second service, where the address offset values of the plurality of second address information relative to the plurality of first address information include: the address offset value of each second address information relative to each first address information.
4. The method according to any one of claims 1 to 3, characterized in that The method further includes: The first processing core receives security information from the second processing core and verifies the security of the second processing core according to the security information.
5. The method according to claim 4, wherein The security information includes at least one of the following: address space identifier ASID, virtual machine identifier VMID, or privilege level information.
6. The method according to any one of claims 1-5, characterized in that, The first processing core receiving second running information from a second processing core includes: The first processing core receives the second running information from the second processing core through a bus.
7. The method according to claim 6, wherein The running frequency of the first processing core is different from that of the second processing core. The first processing core receiving first running information from the second processing core through a bus includes: The asynchronous interface of the first processing core receives the second running information from the second processing core through a bus and performs asynchronous logic processing on the second running information.
8. The method according to any one of claims 1 to 7, characterized in that The first processing core includes a first cache. Before the first processing core runs the first service according to the second running information, the method further includes: The first processing core writes the second running information into the first cache as prefetch information for the first service.
9. The method according to any one of claims 1-8, characterized in that, The second processing core includes a memory access control module and a second cache. The method further includes: The memory access control module sends the second running information to the second cache, and the second running information includes the address information accessed by the memory access control module during the process of the second processing core running the second service; The second cache filters duplicate addresses in the second running information and sends the second running information after filtering the duplicate addresses to the first processing core.
10. A service processing device, characterized in that, The device includes: A first processing core, configured to: Obtain first running information, where the first running information is obtained by the first processing core running a first service; Receive second running information from a second processing core, where the second running information is obtained by the second processing core running a second service; When it is determined that at least part of the first service and the second service are the same according to the first running information and the second running information, run the first service according to the second running information; and The second processing core is configured to provide the second running information to the first processing core.
11. The device according to claim 10, wherein, The first running information and the second running information include at least one of the following: data address information, instruction address information, branch address information, jump information, or value prediction information.
12. The device according to claim 10 or 11, characterized in that, The first running information includes a plurality of first address information, and the second running information includes a plurality of second address information; The first processing core is further configured to: When the number of identical address information existing in the plurality of first address information and the plurality of second address information is greater than a preset threshold, determine that at least part of the first service and the second service are the same; or, When the number of identical address offset values existing in the address offset values of the plurality of second address information relative to the plurality of first address information is greater than a preset threshold, determine that at least part of the first service and the second service are the same, where the address offset values of the plurality of second address information relative to the plurality of first address information include: the address offset value of each second address information relative to each first address information.
13. The device according to any one of claims 10 to 12, characterized in that, The first processing core is further configured to: Receive security information from the second processing core and verify the security of the second processing core according to the security information.
14. The device according to claim 13, characterized in that, The security information includes at least one of the following: address space identifier ASID, virtual machine identifier VMID, or privilege level information.
15. The device according to any one of claims 10 to 14, characterized in that, The device further includes: a bus, configured to transmit the second running information from the second processing core to the first processing core.
16. The device according to claim 15, wherein, The operating frequency of the first processing core is different from that of the second processing core; The first processing core includes: an asynchronous interface, configured to receive the second running information from the second processing core through the bus and perform asynchronous logic processing on the second running information.
17. The device according to any one of claims 10-16, characterized in that, The first processing core includes a first cache; The first processing core is further configured to write the second running information as prefetch information of the first service into the first cache.
18. The device according to any one of claims 10 to 17, characterized in that The second processing core includes: A memory access control module, configured to send the second running information to a second cache, where the second running information includes address information accessed by the memory access control module during the process of the second processing core running the second service; The second cache is configured to filter duplicate addresses in the second running information and send the second running information after filtering the duplicate addresses to the first processing core.
19. An electronic device, characterized in that, The electronic device includes a memory and at least one processor. The memory is used to store computer instructions. The at least one processor includes a plurality of processing cores and is used to execute the computer instructions so that the electronic device implements the service processing method for a plurality of processing cores as described in any one of claims 1-9.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions. When the computer instructions run on a device, the device is caused to execute the service processing method for a plurality of processing cores as described in any one of claims 1-9.
Citation Information
Cited By
Service processing method and apparatus for plurality of processing cores, and device
WO2025130916A1