Service processing apparatus and method, and device and storage medium

By introducing the coupling of a global predictor with multiple processing cores in the central processing unit (CPU) subsystem, the historical operation information and the operation information generated by the operation service are cached, and the problem of low efficiency of the CPU subsystem during business processing is solved, and the effect of reducing the cache missing rate and improving the prediction accuracy is achieved.

WO2025130918A1PCT designated stage expired Publication Date: 2025-06-26HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/140283
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-18
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The central processing unit (CPU) subsystem has problems with low efficiency when performing business processing, which are manifested as high cache missing rate, low branch prediction accuracy, and low instruction number/clock cycle (IPC).

Method used

A service processing device that uses a global predictor coupled with multiple processing cores is used to cache the historical operation information of multiple processing cores and the operation information generated by the operation service through the global predictor, and dynamically update the cache information to reduce the cache missing rate and improve the prediction accuracy.

Benefits of technology

The cache missing rate of multiple processing cores when processing services is reduced, prediction accuracy and IPC are improved, and the operation efficiency of the service is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024140283_26062025_PF_FP_ABST
    Figure CN2024140283_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present application relate to the technical field of electronics. Provided are a service processing apparatus and method, and a device and a storage medium, which are used for reducing the cache miss rate when a plurality of processing cores process services, improving the prediction accuracy and the IPC, and further improving the operation efficiency of the services. The service processing apparatus comprises: a global predictor, and a plurality of processing cores coupled to the global predictor, wherein the global predictor is configured to cache historical operation information of at least one of the plurality of processing cores; a first processing core among the plurality of processing cores is configured to acquire first historical operation information from among the historical operation information, and operate a first service on the basis of the first historical operation information, so as to obtain first operation information; the first processing core is any one of the plurality of processing cores; and the global predictor is further configured to cache the first operation information.
Need to check novelty before this filing date? Find Prior Art

Description

Business processing device, method, equipment and storage medium

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 22, 2023, with application number 202311795739.1 and application name “A business processing device, method, equipment and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of electronic technology, and in particular to a business processing device, method, equipment and storage medium. Background Art

[0003] Currently, business processing is typically performed using a central processing unit (CPU) subsystem, which includes multiple processing cores. These cores can be used to execute tasks such as benchmark performance testing, business performance testing, and similar business deployment or business suite deployment. These tasks require large instructions and data volumes, while the CPU subsystem's multiple processing cores have limited cache resources, resulting in low efficiency when running these tasks. Summary of the Invention

[0004] The present application provides a business processing apparatus, method, device and storage medium for improving the operating efficiency of multiple processing cores when running business.

[0005] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:

[0006] In a first aspect, a business processing device is provided, comprising: a global predictor, and multiple processing cores coupled to the global predictor; the global predictor is used to cache historical operation information of at least one processing core among the multiple processing cores, and the at least one processing core may be part or all of the multiple processing cores; a first processing core among the multiple processing cores is used to obtain first historical operation information among the historical operation information, and operate a first business according to the first historical operation information to obtain first operation information, and the first historical operation information may be operation information related to the first business, and the first processing core is any processing core among the multiple processing cores; the global predictor is further used to cache the first operation information, that is, the global predictor can dynamically update the historical operation information according to the operation information obtained by any processing core among the multiple processing cores when operating the business.

[0007] In the above technical solution, the global predictor caches historical operation information and the operation information generated by any processing core in the multiple processing cores when running the business, so as to update the cached operation information for subsequent use. Any processing core in the multiple processing cores can obtain the required operation information from the global predictor when running the business. In this way, compared with expanding the cache with a larger capacity for the multiple processing cores, this solution can reduce the cache miss rate when the multiple processing cores process the business, improve the prediction accuracy and IPC, and thus improve the operation efficiency of the business.

[0008] In a possible implementation of the first aspect, the historical operation information includes historical operation information of multiple businesses, and the historical operation information of the multiple businesses includes first historical operation information; the global predictor includes: multiple buffers for respectively caching the historical operation information of the multiple businesses, wherein the first buffer among the multiple buffers is used to cache the first historical operation information. For example, the historical operation information of different businesses can be cached in different buffers. In the above possible implementation, by caching the historical operation information of the multiple businesses in multiple buffers, it is possible to prevent any of the multiple processing cores from obtaining historical operation information that does not belong to its own business, avoid the overflow of historical operation information in the global predictor and cause security information leakage, thereby improving the security of the historical operation information in the global predictor.

[0009] In a possible implementation of the first aspect, the first processing core is further used to: obtain configuration information for the first business, the configuration information is used to indicate the first buffer (for example, for indicating the size of the first buffer, and / or indicating the address space corresponding to the first buffer), the configuration information can be sent by the software running on the multiple processing cores, or can be pre-configured in the first processing core; configure the first buffer for the first business according to the configuration information, wherein configuring the first buffer includes enabling, disabling or clearing the first buffer, for example, enabling the first buffer through the configuration information when it is needed, and disabling or clearing the first buffer through the configuration information after use. In the above possible implementation, by configuring a corresponding buffer in the global predictor for a business of any of the multiple processing cores through the configuration information, the resources in the global predictor can be allocated to the key business running on the processing core, thereby avoiding resource bottlenecks and further improving the operating efficiency of the business; in addition, by enabling, disabling or clearing the buffer in the global predictor, the utilization rate of the resources in the global predictor can also be improved.

[0010] In one possible implementation of the first aspect, the size (or dimension) of each of the multiple buffers is determined based on the service to which the historical operation information cached in the buffer belongs. For example, the size of the first buffer is determined based on the first service. In this possible implementation, determining the size of the buffer used to cache the historical operation information of a particular service based on the service can improve the accuracy of the configured buffer size, thereby avoiding resource waste or resource shortage.

[0011] In one possible implementation of the first aspect, the multiple processing cores further include: a second processing core configured to, when the first service is switched from the first processing core to the second processing core, obtain first historical operation information from the first buffer and operate the first service based on the first historical operation information. In this possible implementation, when the first service is switched from the first processing core to the second processing core, the second processing core can still obtain the first historical operation information from the first buffer that caches the first service, thereby avoiding the problem of a large number of cold misses on the second processing core when the first service is switched.

[0012] In one possible implementation of the first aspect, the historical operation information or the first operation information includes at least one of the following: branch address pair information, jump information, instruction traces, data traces, cache hot and cold information, scheduling recommendation information, page table information, hardware prefetching pattern information, or difficult-to-predict address information. This possible implementation can significantly improve the accuracy and efficiency of speculation based on the at least one item of information, thereby improving the operational efficiency of the business.

[0013] In one possible implementation of the first aspect, the historical operation information or the first operation information further includes a branch block index, which is used to index the at least one item. In this case, when a branch block corresponding to the branch block index is accessed, a search of the global predictor can be triggered. This possible implementation can significantly improve the accuracy and efficiency of branch block speculation in the business, thereby improving the operational efficiency of the business.

[0014] In one possible implementation of the first aspect, the first processing core includes a buffer configured to obtain at least a portion of the first historical operation information from the global predictor and cache the at least a portion. In this possible implementation, by providing a relatively small buffer in the first processing core and dynamically obtaining operation information required for the first service from the first historical operation information of the global predictor through the buffer, the impact of transmission delay of the first historical operation information on the operation efficiency of the first service is avoided, thereby improving the operation efficiency of the first service.

[0015] In one possible implementation of the first aspect, the first historical operation information is historical operation information of another service that is identical to or of the same type as the first service. In this possible implementation, when the first historical operation information is historical operation information of another service that is identical to or of the same type as the first service, and the operation information of the first service has significant duplication and similarity with the first historical operation information, the first processing core can significantly improve the operating efficiency of the first service by operating the first service based on the first historical operation information.

[0016] In a second aspect, a business processing method is provided, which is applied to a business processing device, the device including a global predictor and multiple processing cores coupled to the global predictor; the method includes: the global predictor caching historical operation information of at least one processing core among the multiple processing cores; a first processing core among the multiple processing cores obtaining first historical operation information in the historical operation information, and running a first business according to the first historical operation information to obtain first operation information, the first processing core being any processing core among the multiple processing cores; the global predictor caching the first operation information.

[0017] In a possible implementation of the second aspect, the historical operation information includes historical operation information of multiple businesses, and the historical operation information of the multiple businesses includes first historical operation information; the global predictor caches the historical operation information of at least one processing core among the multiple processing cores, including: multiple buffers of the global predictor respectively cache the historical operation information of the multiple businesses, wherein the first buffer among the multiple buffers is used to cache the first historical operation information.

[0018] In a possible implementation of the second aspect, the method also includes: the first processing core obtains configuration information for the first business, and the configuration information is used to indicate the first buffer; the first processing core configures the first buffer for the first business according to the configuration information, wherein configuring the first buffer includes enabling, disabling or clearing the first buffer.

[0019] In a possible implementation manner of the second aspect, the size of the first buffer is determined according to the first service.

[0020] In a possible implementation of the second aspect, the multiple processing cores also include a second processing core, and the method also includes: when the first business is switched from the first processing core to the second processing core, the second processing core obtains the first historical operation information from the first buffer and runs the first business according to the first historical operation information.

[0021] In a possible implementation of the second aspect, the historical operation information or the first operation information includes at least one of the following: branch address pair information, jump information, instruction trace, data trace, cache hot and cold information, scheduling recommendation information, page table information, hardware prefetching regularity information or difficult-to-predict address information.

[0022] In a possible implementation manner of the second aspect, the historical operation information or the first operation information further includes a branch block index, where the branch block index is used to index the at least one item.

[0023] In a possible implementation of the second aspect, the first processing core obtains first historical operation information in the historical operation information, including: a buffer of the first processing core obtains at least part of the first historical operation information from the global predictor and caches the at least part.

[0024] In a possible implementation manner of the second aspect, the first historical operation information is historical operation information of other services that are the same as or belong to the same type as the first service.

[0025] In a third aspect, an electronic device is provided, which includes a memory and a business processing device, the business processing device includes a global predictor and multiple processing cores, the memory is used to store computer instructions, and the business processing device is used to execute the computer instructions so that the electronic device implements the business processing method provided by the second aspect or any possible implementation of the second aspect.

[0026] In another aspect of the present application, a computer-readable storage medium is provided, in which a computer program or instruction is stored. When the computer program or instruction is executed, the business processing method provided by the second aspect or any possible implementation of the second aspect is implemented.

[0027] In another aspect of the present application, a computer program product is provided, which includes: a computer program (also referred to as code, or instructions), which, when executed, enables a computer to execute a business processing method as provided in the second aspect or any possible implementation of the second aspect.

[0028] It can be understood that the beneficial effects that can be achieved by any of the business processing methods, electronic devices, computer-readable storage media and computer program products provided above can correspond to the beneficial effects in the business processing device provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] FIG1 is a schematic diagram of the structure of a multi-core system provided in an embodiment of the present application;

[0030] FIG2 is a schematic diagram of the structure of a service processing device provided in an embodiment of the present application;

[0031] FIG3 is a schematic structural diagram of a LBTB provided in an embodiment of the present application;

[0032] FIG4 is a schematic structural diagram of another LBTB provided in an embodiment of the present application;

[0033] FIG5 is a schematic diagram of a processing core acquiring historical operation information according to an embodiment of the present application;

[0034] FIG6 is a schematic structural diagram of an sBTB provided in an embodiment of the present application;

[0035] FIG7 is a schematic diagram of a method for multiple processing cores to share a global prediction according to an embodiment of the present application;

[0036] FIG8 is a schematic diagram of configuring a global predictor provided in an embodiment of the present application;

[0037] FIG9 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0038] The following will discuss in detail the making and use of various embodiments. However, it should be understood that many applicable inventive concepts provided herein can be implemented in a variety of specific contexts. The specific embodiments discussed are merely illustrative of specific ways to implement and use the present application and technology and do not limit the scope of this application.

[0039] Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art.

[0040] Various circuits or other components may be described or referred to as being "configured to" perform one or more tasks. In this case, "configured to" is used to imply structure by indicating that the circuit / component includes structure (e.g., circuitry) that performs the one or more tasks during operation. Thus, even when a specified circuit / component is not currently operational (e.g., not turned on), the circuit / component may be referred to as being configured to perform the task. Circuits / components used with the phrase "configured to" include hardware, such as circuitry that performs an operation, etc.

[0041] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. In the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can represent: a, b, c, a and b, a and c, b and c, a, b and c; where a, b and c can be single or multiple.

[0042] The embodiments of this application use terms such as "first" and "second" to distinguish objects with similar names, functions, or effects. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or order of execution. The term "coupled" is used to indicate an electrical connection, including direct connection via wires or connectors or indirect connection via other devices. Therefore, "coupling" should be considered a broadly defined electronic communication connection.

[0043] It should be noted that, in this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0044] Currently, central processing unit (CPU) subsystems often suffer from inefficiencies when processing services. These issues manifest themselves in high cache miss rates, low branch prediction accuracy, and low instructions per clock cycle (IPC). For example, the instructions and data required for services running on the CPU subsystem are large, while the capacity of resources within the CPU subsystem, such as the cache, branch target buffer (BTB), and branch direction predictor, is limited, leading to frequent cache misses. Another example is that software within the CPU subsystem is unaware of opportunistic practices such as front-end and back-end prefetching, rendering prefetch information ineffective for supporting other predictions and resulting in information waste. Furthermore, applications running on the CPU subsystem may have a significantly greater number of threads than the number of processing cores (also known as cores) included in the CPU subsystem. This can easily lead to a large number of thread switches (both within the same core and between different cores). When a thread switch occurs, the data stored in resources such as the cache, BTB, and branch direction predictor may not be required by the current thread, resulting in a large number of cold misses for the current thread.

[0045] Based on this, an embodiment of the present application provides a business processing device, which includes a global predictor (GP) and multiple processing cores coupled to the global predictor. The global predictor can be used to cache historical operation information, and any one of the multiple processing cores can be used to obtain the operation information required in the historical operation information and run the business based on the obtained operation information. The global predictor is also used to cache the operation information obtained by running the business. In this way, by caching the historical operation information and the operation information generated by any one of the processing cores running the business, the multiple processing cores obtain the operation information they need from the global predictor when running the business, which can reduce the cache miss rate of the multiple processing cores when processing the business, improve the prediction accuracy and IPC (instruction number / clock cycle), and thus improve the operation efficiency of the business.

[0046] The business processing device provided in the embodiment of the present application can be applied to a multi-core system, which can be called a processor subsystem (for example, a CPU subsystem). The structure of the multi-core system is introduced below.

[0047] Figure 1 is a structural diagram of a multi-core system provided in an embodiment of the present application. The multi-core system includes a global predictor, and a plurality of processing cores coupled to the global predictor, and the plurality of processing cores can share the global predictor through software configuration. The plurality of processing cores can be used to deploy different services, the same services or partially the same services, and the partially the same services may refer to the existence of partially identical data and / or instructions that two services need to access during operation. In an embodiment of the present application, the plurality of processing cores can be used to perform benchmark performance testing, service performance testing, similar service deployment or service suite deployment and other services. Exemplarily, the business scenarios of the multi-core system include but are not limited to: server cluster benchmark performance testing, server cluster bidding service performance testing, server cluster similar service deployment (for example, high performance computing (HPC)), server cluster service suite deployment, terminal field benchmark performance testing, service switching, service migration, etc.

[0048] The number of multiple processing cores included in the multi-core system can be configured according to actual needs. For example, the number of the multiple processing cores can be 2, 4, 6 or 8; in addition, the multiple processing cores can be a homogeneous structure (i.e., including multiple processing cores) or a heterogeneous structure (i.e., including different processing cores). In one example, as shown in (a) of Figure 1, the multi-core system includes: a global predictor, and 4 processing cores coupled to the global predictor, the 4 processing cores are respectively represented as core 0 to core 3, and the 4 processing cores are a homogeneous structure. In another example, as shown in (b) of Figure 1, the multi-core system includes: a global predictor, and 8 processing cores coupled to the global predictor, the 8 processing cores are respectively represented as core 0 to core 3, and the 8 processing cores are a homogeneous structure. In another example, as shown in (c) in Figure 1, the multi-core system includes: a global predictor, and 8 processing cores coupled to the global predictor. The 8 processing cores can be a heterogeneous structure, and the 8 processing cores can include super large core 0, large core 1 to large core 3, and small core 4 to small core 7.

[0049] In the multi-core system, the multiple processing cores may include multiple different processing cores of the same processor, or the multiple processing cores may include processing cores of multiple different processors, that is, the multiple processing cores may form one or more processors. The processors include, but are not limited to, CPUs, general-purpose processors, graphics processing units (GPUs), image signal processors (ISPs), digital signal processors (DSPs), network processing units (NPUs), artificial intelligence (AI) processors, and the like. Optionally, the structures of any two of the multiple processing cores and the sizes of their respective hardware resources (e.g., cache resources) may be the same or different, and this embodiment of the application does not impose any specific restrictions on this.

[0050] Furthermore, the multi-core system may also include a memory coupled to the multiple processing cores, which may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. Exemplarily, the RAM may be a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), and a double data rate synchronous dynamic random access memory (DDR SDRAM). The memory is not shown in the figure.

[0051] The multi-core system can be an electronic device, or a system on chip (SoC) or a chipset comprising multiple chips applied to an electronic device, or a module comprising the SoC or chipset. The electronic device can be used as a terminal device or a server. Optionally, the electronic device includes but is not limited to: mobile phones, tablet computers, laptops, desktop computers, PDAs, ultra-mobile personal computers (umPCs), mobile internet devices (MIDs), netbooks, camcorders, cameras, wearable devices (such as smart watches and smart bracelets, etc.), vehicle-mounted equipment (such as cars, bicycles, electric vehicles, airplanes, ships, trains, high-speed railways, etc.), virtual reality (VR) equipment, augmented reality (AR) equipment, wireless terminals in industrial control, smart home devices (such as refrigerators, televisions, air conditioners, electric meters, etc.), intelligent robots, workshop equipment, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, or wireless terminals in smart homes, flying equipment (such as intelligent robots, hot air balloons, drones, airplanes), etc.

[0052] FIG2 is a schematic diagram of the structure of a service processing device provided in an embodiment of the present application. The service processing device includes a global predictor 10 and multiple processing cores 20 coupled to the global predictor 10. The multiple processing cores 20 share the global predictor 10, for example, by configuring the multiple processing cores 20 to share different resources of the global predictor 10 through software. FIG2 illustrates an example in which the multiple processing cores 20 include processing cores 21 through 2n, where n is an integer greater than 1.

[0053] The global predictor 10 is used to cache historical operation information of at least one processing core among the multiple processing cores 20. The at least one processing core includes one or more processing cores. When the at least one processing core includes multiple processing cores, the at least one processing core may be part of the multiple processing cores 20 or all of the multiple processing cores 20. The historical operation information cached in the global predictor 10 may be obtained by the at least one processing core running a business. For example, each of the at least one processing core may run one or more businesses and transmit the operation information obtained from the operation to the global predictor 10. The global predictor 10 caches the operation information transmitted by the at least one processing core as the historical operation information.

[0054] For any one of the multiple processing cores 20 (hereinafter referred to as the first processing core 21 for the convenience of description), it is used to: obtain the first historical operation information in the historical operation information, and run the first business according to the first historical operation information to obtain the first operation information. The first historical operation information can be part or all of the historical operation information. Optionally, the first historical operation information is the historical operation information of other businesses that are the same as or belong to the same type as the first business. For example, the first historical operation information can be the historical operation information of the first business, that is, the operation information obtained by the previous operation of the first business; or, the first business and the second business are two businesses in the same high-performance computing, and the first historical operation information is the historical operation information of the second business. The following description takes the first historical operation information being the historical operation information of the first business as an example.

[0055] The global predictor 10 is also used to: cache the first operation information. For example, when the first processing core 21 obtains the first operation information, it transmits the first operation information to the global predictor 10, and the global predictor 10 receives and caches the first operation information. After the global predictor 10 caches the first operation information, when any processing core of the multiple processing cores 20 (for example, the first processing core 21 or other processing cores) obtains historical operation information from the global predictor 10, the historical operation information includes the first operation information. That is, the historical operation information is dynamically updated, and the global predictor 10 can dynamically update the historical operation information based on the operation information obtained by the multiple processing cores 20 when running the business.

[0056] Optionally, the historical operation information or the first operation information includes at least one of the following: branch address pair information, jump information, instruction trace, data trace, cache hot and cold information, scheduling recommendation information, page table information, hardware prefetching regularity information or difficult-to-predict address information. The branch address pair information refers to the address of the executed branch in the memory, which includes a source address and a target address, and the source address and the target address form an address pair. The jump information may include a jump direction and a jump destination, where the jump direction indicates the jump direction of the jump instruction (i.e., jump or not jump), and the jump destination refers to the destination address after the jump. The data trace indicates the trace of the accessed data. The instruction trace indicates the trace of the accessed instructions. The cache hot and cold information indicates the hotness of the data and / or instructions in the cache (e.g., the first, second, and third level caches). The scheduling recommendation information indicates information recommended during the scheduling process, such as a scheduling policy recommended based on quality of service (QoS). The page table information is used to indicate the correspondence between the virtual address and the physical address corresponding to the accessed data and / or instruction. The regular information of hardware prefetching is used to indicate the regularity of the hardware prefetching data and / or instructions, which may include front-end prefetching and / or back-end prefetching. The difficult-to-predict address information is used to indicate the address information that is difficult for the hardware prefetcher (hardware prefetch, HWP) to predict. The memory targeted by the above address may be located on the same or different chip as the one or more processing cores mentioned above, and may be a volatile or non-volatile memory, which is not limited in this embodiment.

[0057] It can be understood that, in addition to the above information, the above historical operation information or the first operation information may also include other information required during the business operation process (such as value prediction information, which may refer to the numerical value obtained by prediction when the memory access is not completed). The embodiment of the present application does not impose specific restrictions on this.

[0058] Optionally, the historical operation information or the first operation information further includes a branch block index, which is used to index at least one of the above-mentioned information. In this case, when the branch block corresponding to the branch block index is accessed, a search of the global predictor 10 can be triggered. The branch block index can also be referred to as a branch block program counter (PC) index (block PC index). For ease of understanding, the historical operation information and the first operation information are used as an example to illustrate the branch block index and at least one item corresponding to the branch block index included in the historical operation information and the first operation information.

[0059] In one example, assuming that the historical operation information includes historical operation information of multiple businesses, the historical operation information of each business may include at least one branch block index, and historical operation sub-information corresponding to each branch block index, and each historical operation sub-information may include at least one of the following: branch address pair information, jump information, instruction trace, data trace, cache hot and cold information, scheduling recommendation information, page table information, hardware prefetching regularity information or difficult-to-predict address information.

[0060] Each of the aforementioned multiple businesses may include at least one branch block, and each branch block (or subroutine or program segment) in the at least one branch block may correspond to a branch block index. After the branch block is executed, a corresponding historical execution sub-information may be obtained. The branch block index of a branch block may be used to index the historical execution sub-information obtained after the branch block is executed. Optionally, the branch block index may be an instruction fetch address, and the instruction fetch address may be a source address corresponding to the branch block.

[0061] Exemplarily, the historical operation information may be shown in the following Table 1. In Table 1, it is assumed that the multiple services include M services, and the M services are represented as services 1 to M, at least one branch block index of service 1 is represented as ID1-1, ID1-2, ..., at least one branch block index of service 2 is represented as ID2-1, ID2-2, ..., and at least one branch block index of service M is represented as IDM-1, IDM-2, ....

[0062] Table 1

[0063] It can be understood that the above Table 1 uses the example that the information types included in the historical operation sub-information corresponding to different branch block indexes are the same. In actual applications, the information types included in the historical operation sub-information corresponding to different branch block indexes may also be different. The embodiment of the present application does not impose specific restrictions on this.

[0064] In addition, the global predictor 10 may also include other information of each of the multiple services. For example, the other information may include a service identifier and / or a service context, etc. This embodiment of the present application does not impose any specific limitation on this.

[0065] Furthermore, the global predictor 10 includes multiple buffers. The multiple buffers are used to cache the historical operation information of the multiple services, respectively. Each of the multiple buffers can be used to cache the historical operation information obtained by a processing core, i.e., the historical operation information obtained by different processing cores can be cached in different buffers; or to cache the historical operation information of a service, i.e., the historical operation information of different services can be cached in different buffers.

[0066] Optionally, the multiple buffers include a first buffer, and the first buffer is used to cache historical operation information of the first business, and the historical operation information of the first business may include first historical operation information. In a possible example, for any one of the multiple processing cores 20 except the first processing core 21 (referred to as the second processing core 22 herein), when the first business switches from the first thread of the first processing core 21 to the second thread of the second processing core 22, the second processing core 22 can be used to obtain the first historical operation information from the first buffer, and run the first business according to the first historical operation information. In this example, when the first business switches from the first processing core 21 to the second processing core 22, the second processing core 22 can still obtain the first historical operation information from the first buffer that caches the first business, thereby avoiding the problem of a large number of cold misses in the second processing core 22 when the first business switches.

[0067] Optionally, the size (or dimension) of each buffer in the plurality of buffers may be determined based on the business to which the historical operation information cached in the buffer belongs. Exemplarily, the size of the first buffer is determined based on the first business. That is, the size of the buffer used to cache the historical operation information of a business may be dynamically determined based on the size of the business. For example, in actual applications, the size of the corresponding buffer for each business may be dynamically determined by software (e.g., an operating system) running on the plurality of processing cores 20. This embodiment of the present application does not impose any specific restrictions on this.

[0068] In a possible embodiment, the global predictor 10 includes a multi-way buffer, and each buffer in the above-mentioned multiple cache areas may include at least one buffer in the multi-way buffer. Optionally, each buffer in the multi-way buffer can be used to cache the historical operation information of a service, and the historical operation information of different services is cached in different buffers. The historical operation information of the same service can occupy one or more buffers. In addition, the multi-way buffers are independent of each other, and each buffer can be accessed separately, and the access corresponding to the buffers of different channels does not affect each other. Exemplarily, the multi-way buffer includes a first buffer, and the first buffer is used to cache the historical operation information of the first service.

[0069] When the historical operation information of a service includes at least one of the above-mentioned multiple information items, such as branch address pair information, jump information, and instruction trace, each buffer in the global predictor 10 can cache any of the at least one item using the following structure. For ease of description, the following example uses the first processing core 21 running the first service, and the first buffer is used to cache the historical operation information of the first service, where the historical operation information of the first service includes branch address pair information. The first buffer can also be called a link branch target buffer (LBTB).

[0070] In one example, as shown in Figure 3, the LBTB may include a history queue, a storage circuit, a search queue, and a fill queue. The history queue, the search queue, and the fill queue are all coupled to the storage circuit.

[0071] The history queue is used to obtain the correct branch address pair (i.e., including the source address and the target address) from the pipeline corresponding to the first business when a branch prediction target address failure occurs during the operation of the first business (represented as a prediction failure in the figure, and the prediction failure may mean that the target address corresponding to the branch is not found in the cache of the first processing core 21 (e.g., the first buffer and the second buffer below)), and output the branch address pair to the storage circuit. Optionally, in order to reduce the number of updates to the LBTB to reduce power consumption and access conflicts, the history queue can be used to output the preset number of branch address pairs as a group according to a certain storage format to the storage circuit when a preset number of branch address pairs (e.g., 4 address pairs) are accumulated. In Figure 3, the preset number of branch address pairs includes brn add0-tgt add0, brn add1-tgt add1, brn add2-tgt add2, brn add3-tgt add3, and the storage format also includes a link address link add as an example. The link add is used to indicate the branch address pair of the next group.

[0072] The storage circuit is configured to receive and cache branch address pairs output by the history queue. When caching a group of branch address pairs in the aforementioned storage format, the storage circuit may index the group of branch addresses using the source address of the first branch address pair to write the group of branch addresses into the corresponding storage. Optionally, the storage circuit may include multiple connected RAMs, thereby improving search efficiency when searching the storage circuit.

[0073] The search queue is used to obtain the instruction fetch address of the branch block and output the instruction fetch address to the storage circuit. The search queue can receive and filter duplicate address information; in addition, the search queue can be a multi-input and one-output queue, which can achieve the effect of balancing the bandwidth difference between input and output. Optionally, the search queue can obtain the instruction fetch address when the first processing core 21 meets certain conditions. For example, the conditions may include but are not limited to: querying and missing in the main branch target buffer (mBTB), or querying and hitting in the stream branch target buffer (sBTB). For a detailed description of mBTB and sBTB, please refer to the description below, and the embodiments of the present application will not be repeated here.

[0074] The storage circuit is further configured to, upon receiving an instruction fetch address of a branch block of the first processing core 21, obtain a corresponding branch address pair based on the instruction fetch address and output the branch address pair to the backfill queue. Optionally, when caching a group of branch address pairs in the above-mentioned storage format, the storage circuit may disassemble the obtained group of branch addresses, output a preset number of the disassembled branch address pairs to the backfill queue, and output a link add in the group of branch address pairs to the search queue, so that the search queue performs the next search based on the link add.

[0075] The backfill queue is used to backfill the branch address pairs output by the storage circuit into the first processing core 21. Optionally, the backfill queue can backfill a preset number of branch address pairs output by the storage circuit into the buffer of the first processing core 21, for example, into the sBTB of the first processing core 21.

[0076] In another example, as shown in FIG4 in conjunction with FIG3 , for each processing core in the multiple processing cores 20, the global predictor 10 may include a history queue, a storage circuit, a search queue, and a backfill queue corresponding to each processing core. The different storage circuits in the global predictor 10 may be integrated together as a storage module; alternatively, the global predictor includes a storage module, which includes a storage circuit corresponding to each processing core, and the storage circuit includes a multi-way cache. FIG4 illustrates the multiple processing cores 20 as comprising four processing cores, denoted as cores 0 through 3. Core 0 corresponds to history queue 0, search queue 0, backfill queue 0, and two-way caches w0 and w1. Core 1 corresponds to history queue 1, search queue 1, backfill queue 1, and two-way caches w2 and w3. Core 2 corresponds to history queue 2, search queue 2, backfill queue 2, and two-way caches w4 and w5. Core 3 corresponds to history queue 3, search queue 3, backfill queue 3, and four-way caches w6 through w9.

[0077] Furthermore, the following takes the example of the first processing core 21 obtaining the first historical operation information from the global predictor 10 to introduce and illustrate the relevant process of any processing core among the multiple processing cores 20 obtaining the historical operation information from the global predictor 10 .

[0078] In one possible embodiment, the first processing core 21 includes a first buffer, which is used to obtain at least a portion of the first historical operation information from the global predictor 10 and cache the at least a portion. The first processing core 21 may also include a second buffer, which is different from the first buffer. The second buffer may refer to a buffer set in the first processing core 21 for caching prefetch information. The second buffer is not used to cache the historical operation information obtained from the global predictor 10. In the embodiment of the present application, the first buffer may also be referred to as a candidate buffer (e.g., sBTB), and the second buffer may be referred to as a primary buffer (e.g., mBTB).

[0079] Optionally, the first buffer may be specifically configured to obtain and cache at least a portion of the first historical operation information from the global predictor 10 when a certain condition is met. For example, the condition may include: a query in the second buffer resulting in a miss (i.e., the second buffer does not have the required information), or a query in the first buffer resulting in a hit (i.e., the first buffer has the required information), etc. This embodiment of the present application does not impose any specific limitations on this.

[0080] When the first historical execution information includes at least one of the above-mentioned multiple pieces of information, such as branch address pair information, jump information, and instruction trace, the number of first buffers of the first processing core 21 can be at least one, that is, the first processing core 21 includes at least one first buffer, and each of the at least one first buffer can be used to cache one of the above-mentioned multiple pieces of historical execution information. For ease of description, the following example is used to illustrate that the first historical execution information includes branch address pair information, and the first processing core 21 includes a main branch target buffer mBTB (corresponding to the second buffer) and a streaming branch target buffer sBTB (corresponding to the first buffer).

[0081] Exemplarily, as shown in FIG5 , the first processing core 21 includes an mBTB and an sBTB. The mBTB caches a first branch address pair set corresponding to a first service prefetched by the first processing core 21, and the sBTB caches a second branch address pair set obtained from the LBTB of the global predictor 10. During the process of the first processing core 21 executing the first service, when the first processing core 21 obtains an instruction fetch address FIVA through instruction fetching, the first processing core 21 may first query whether the first branch address pair set cached in the mBTB contains the target branch address corresponding to the FIVA; if so (i.e., a hit), execution continues according to the target branch address; if not (i.e., a miss), the first processing core 21 queries whether the second branch address pair set cached in the sBTB contains the target branch address corresponding to the FIVA; if so (i.e., a hit), execution continues according to the queried target branch address. In addition, when a query in the mBTB fails, or a query in the sBTB hits, it indicates that the historical operation information is accurate and available as prefetch information. The sBTB can obtain more branch address pairs corresponding to the first business from the LBTB of the global predictor 10. For example, the sBTB sends the next instruction fetch address to the global predictor 10 to obtain the more branch address pairs through the next instruction fetch address for use when the first processing core 21 runs the first business. Optionally, when the first business ends, or when other businesses need to use the sBTB, the first processing core 21 can also clear (or flush) the sBTB. For example, the first processing core 21 can clear the sBTB through the valid bit.

[0082] It is understandable that the conditions for the sBTB to obtain the branch address pair from the LBTB of the global predictor 10 may also include other conditions, such as obtaining it when a certain time period is reached, or obtaining it when all pipeline jumps of the first processing core 21 occur, etc. The embodiment of the present application does not impose specific restrictions on this.

[0083] In addition, Figure 6 shows a possible structural diagram of the sBTB described above. As shown in Figure 6, the sBTB may include a storage circuit and a valid bit array. The storage circuit may be used to cache historical operation information from the global predictor 10, and the valid bit array may be used to indicate the validity of the cached information in the storage circuit. The storage circuit may include multi-way connected RAM, which can improve search efficiency when searching the storage circuit. Optionally, the sBTB may also include a matching and selection circuit, which may be used to match the business belonging to the FIVA with the business currently cached by the sBTB and output when the target branch address hit by the sBTB is valid. In addition, the output end of the mBTB and the output end of the sBTB may be coupled via a selector to select the output result of the mBTB when the mBTB hits, and to select the output result of the sBTB when the sBTB hits.

[0084] Optionally, when the first processing core 21 queries the first buffer and finds a hit, the first processing core 21 may reload the hit operation information as speculative information into modules or components required by the pipeline. For example, branch address pairs may be loaded back into the branch pipeline, hot and cold and criticality information of the instruction cache and data cache may be loaded back into the instruction prefetch and replacement component, historical page table operation information may be loaded back into the page table prefetch component, and regularity information and unpredictable address information of hardware prefetch may be loaded back into the hardware prefetch component.

[0085] Furthermore, in a possible example, as shown in FIG7 , when the historical execution information cached in the global predictor includes multiple pieces of information such as branch address pair information, jump information, and instruction traces, the second buffer for caching the multiple pieces of information in any one of the multiple processing cores 21 (e.g., core 0, core 1, and core 2) (e.g., core 0) may include a branch target buffer (BTB), an instruction cache (Icache), a data cache (Dcache), and a translation lookaside buffer (TLB). The processing core may also include other components such as an execution pipeline. FIG7 illustrates an example in which the processing core caches the multiple pieces of information obtained from the global predictor 10 via a first buffer, and the first buffer is a stream buffer.

[0086] Accordingly, in one example, as shown in FIG7 , the global predictor 10 includes multiple buffers (e.g., N buffers, represented as GP w0 to GP wN), each of which can be used to cache the historical operation information of a service. The historical operation information of any service can include multiple branch block indexes, as well as the above-mentioned multiple items of information corresponding to each branch block index. FIG7 exemplifies the example in which the branch block indexes in the historical operation information of service 1 include block-ID1 and block-ID2. When the same service switches between different processing cores, the processing core to which the service is switched can share the historical operation information of the service in the global predictor 10.

[0087] After introducing the related structures of the global predictor 10 and the multiple processing cores 20 , the process of allocating resources in the global predictor 10 to any processing core in the multiple processing cores 20 is described in detail below.

[0088] Optionally, the resources in the global predictor 10 used by any of the multiple processing cores 20 can be obtained through configuration. When any of the multiple processing cores 20 needs to use the global predictor 10 to cache historical operation information of a certain service, the global predictor 10 can be configured for that processing core, for example, by configuring a corresponding buffer in the global predictor 10 for that processing core. Furthermore, services that can use the global predictor 10 can also be configured for the multiple processing cores 20, that is, certain key services can be configured for any of the multiple processing cores 20. Only the historical operation information of the configured key services can be cached in the global predictor. The following describes the process of configuring the global predictor using the first processing core 21 as an example.

[0089] In a possible embodiment, the first processing core 21 is further used to: obtain configuration information for the first business, the configuration information is used to indicate the first buffer; configure the first buffer for the first business according to the configuration information. Optionally, configuring the first buffer may include enabling, disabling or clearing the first buffer. The configuration information may be sent to the first processing core 21 by software running on the multiple processing cores 20, or may be pre-configured in the first processing core 21, and the embodiment of the present application does not impose specific restrictions on this. The configuration information can be used to indicate the size of the first buffer, and / or to indicate the address space corresponding to the first buffer. Furthermore, the configuration information can also be used to indicate the first business. For example, the configuration information may include a business identifier of the first business.

[0090] For ease of understanding, as shown in FIG8 , the following describes an example of how to enable, disable, and clear the global predictor 10 using configuration information when the first processing cores 21 are core 0, core 1, and core 2, respectively. FIG8 illustrates this example by taking core 0, core 1, and core 2 as examples, each of which includes an HWP and an SBTB.

[0091] In one example, assuming that the first processing core 21 is core 0, enabling the global predictor 10 through configuration information includes: S11. When service a needs to be run on core 0, the software running on the multiple processing cores 20 can send first configuration information for service a to core 0; S12. When core 0 receives the first configuration information, the HWP of core 0 enables the global predictor 10 according to the first configuration information, for example, sending service information and enable information to the global predictor 10 to enable the global predictor 10 to cache historical operation information of service a; S13. Core 0 can also enable the sBTB when the information currently cached in the sBTB matches service a; S14. The global predictor 10 starts working.

[0092] In another example, assuming that the first processing core 21 is core 1, shutting down the global predictor 10 through configuration information includes: S21. When it is necessary to shut down the global predictor 10 used by the service b running on core 1, the software running on the multiple processing cores 20 can send second configuration information for the service b to core 1; S22. When receiving the second configuration information, the HWP of core 1 shuts down the global predictor 10 according to the second configuration information, for example, sending service information and shutdown information to the global predictor 10; S23. The global predictor 10 stops providing services for the service b on core 1, that is, stops providing historical operation information of the service to core 1.

[0093] In another example, assuming that the first processing core 21 is core 2, clearing the global predictor 10 using configuration information includes: S31. When it is necessary to clear the global predictor 10 used by service C running on core 2, the software running on the multiple processing cores 20 may send third configuration information for service C to core 2; S32. Upon receiving the third configuration information, the HWP of core 2 clears the global predictor 10 according to the third configuration information, for example, by sending thread information and clearing information to the global predictor 10; S33. The global predictor 10 clears the information corresponding to the thread, that is, clears the historical running information of service C. Optionally, the global predictor 10 may be cleared using a row-by-row clearing method.

[0094] In an embodiment of the present application, the global predictor 10 caches historical operation information and the operation information generated by any processing core when running a business, so as to update the cached operation information for subsequent use. When running a business, the multiple processing cores 20 obtain the operation information they need from the global predictor 10. This is compared to expanding the cache of a larger capacity for the multiple processing cores, thereby reducing the cache miss rate of the multiple processing cores 20 when processing the business, improving the prediction accuracy and IPC, and thus improving the operation efficiency of the business. For business switching scenarios, since the global predictor 10 can be shared by the multiple processing cores 20, when a business is switched from one processing core to another, the other processing core can still obtain the historical operation information of the business from the global predictor 10 for running the business, thereby avoiding the problem of a large number of cold misses on the other processing core. In addition, the resources in the global predictor 10 used by any of the multiple processing cores 20 are obtained through software configuration, and the multiple processing cores 20 can also be configured with services that can use the global predictor 10. This can avoid information overflow in the global predictor 10 causing security information leakage. At the same time, from a software perspective, key services can be identified and the resources in the global predictor 10 can be allocated to the key services running on the processing cores, thereby avoiding resource bottlenecks and further improving the operating efficiency of the services.

[0095] Based on this, embodiments of the present application also provide a service processing method, which can be applied to a service processing device including a global predictor and multiple processing cores coupled to the global predictor. A description of the service processing device can be found in the above description. The method includes: the global predictor caching historical operation information of at least one processing core among the multiple processing cores; a first processing core among the multiple processing cores obtaining first historical operation information from the historical operation information, and executing a first service based on the first historical operation information to obtain first operation information; and the global predictor caching the first operation information.

[0096] Optionally, the historical operation information includes historical operation information of multiple services, and the historical operation information of the multiple services includes first historical operation information.

[0097] In one possible embodiment, the global predictor includes multiple buffers, each of which caches historical operation information of the multiple services, wherein a first buffer among the multiple buffers caches first historical operation information. Accordingly, the method may further include: the first processing core obtaining configuration information for the first service, the configuration information being used to indicate the first buffer; and the first processing core configuring the first buffer for the first service based on the configuration information. Configuring the first buffer includes enabling, disabling, or clearing the first buffer.

[0098] It can be understood that all the contents in the above-mentioned device embodiment can be referred to in the embodiment corresponding to the business processing method, and the embodiments of the present application will not be repeated here.

[0099] In an embodiment of the present application, a global predictor is used to cache historical operation information and operation information generated by any processing core when running a business, so as to update the cached operation information for subsequent use. Any processing core among the multiple processing cores can obtain the required operation information from the global predictor when running a business. In this way, compared with expanding a larger capacity cache for the multiple processing cores, this solution can reduce the cache miss rate when the multiple processing cores process the business, improve the prediction accuracy and IPC, and thus improve the operation efficiency of the business.

[0100] In another aspect of the present application, an electronic device is provided, as shown in FIG9 . The electronic device includes a memory and a service processing device, the service processing device including a global predictor and multiple processing cores. The memory is configured to store computer instructions, and the service processing device is configured to execute the computer instructions, so that the electronic device implements any of the service processing methods provided above. A detailed description of the memory can be found in the previous embodiment.

[0101] It can be understood that all relevant contents of each step involved in the above method embodiment can be referred to the embodiment of the business processing method and the embodiment of the electronic device, and the embodiment of the present application will not be repeated here.

[0102] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical functional division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another device, or ignoring or not implementing certain features.

[0103] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0104] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. The readable storage medium may include: a USB flash drive, a mobile hard drive, a read-only memory, a random access memory, a magnetic disk, or an optical disk, etc., which can store program code. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product.

[0105] In another embodiment of the present application, a readable storage medium is also provided, which stores computer execution instructions. When a device (which may be a single-chip microcomputer, chip, etc.) or a processor executes the steps in the above method embodiment.

[0106] In another embodiment of the present application, a computer program product is provided, which includes computer instructions stored in a readable storage medium; at least one processor of the device can read the computer instructions from the readable storage medium, and at least one processor executes the computer instructions so that the device performs the steps in the above method embodiment.

[0107] Finally, it should be noted that the above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A service processing device, characterized in that: include: A global predictor, and a plurality of processing cores coupled to the global predictor; wherein, The global predictor is used to cache historical operation information of at least one processing core among the multiple processing cores; a first processing core among the multiple processing cores, configured to obtain first historical operation information among the historical operation information, and to run a first service according to the first historical operation information to obtain first operation information, wherein the first processing core is any processing core among the multiple processing cores; The global predictor is further used to cache the first operation information.

2. The device according to claim 1, characterized in that The historical operation information includes historical operation information of multiple services, and the historical operation information of the multiple services includes the first historical operation information; The global predictor includes: a plurality of buffers, used to cache the historical operation information of the plurality of services respectively, wherein a first buffer among the plurality of buffers is used to cache the first historical operation information.

3. The device according to claim 2, characterized in that The first processing core is further configured to: Acquire configuration information for the first service, where the configuration information is used to indicate the first buffer; The first buffer is configured for the first service according to the configuration information, wherein configuring the first buffer includes enabling, disabling or clearing the first buffer.

4. The device according to claim 3, characterized in that The size of the first buffer is determined according to the first service.

5. The device according to any one of claims 2 to 4, characterized in that: The plurality of processing cores also include: The second processing core is used to obtain the first historical operation information from the first buffer and run the first service according to the first historical operation information when the first service is switched from the first processing core to the second processing core.

6. The device according to any one of claims 1 to 5, characterized in that: The historical operation information or the first operation information includes at least one of the following: branch address pair information, jump information, instruction trace, data trace, cache hot and cold information, scheduling recommendation information, page table information, hardware prefetching regularity information or difficult-to-predict address information.

7. The device according to claim 6, characterized in that The historical operation information or the first operation information further includes a branch block index, and the branch block index is used to index the at least one item.

8. The device according to any one of claims 1 to 7, characterized in that: The first processing core comprises: A buffer is used to obtain at least part of the first historical operation information from the global predictor and cache the at least part.

9. The device according to any one of claims 1 to 8, characterized in that: The first historical operation information is historical operation information of other services that are the same as or belong to the same type as the first service.

10. A business processing method, characterized in that: Applied in a service processing device, the device includes a global predictor and a plurality of processing cores coupled to the global predictor; the method includes: The global predictor caches historical operation information of at least one processing core among the plurality of processing cores; A first processing core among the multiple processing cores acquires first historical operation information among the historical operation information, and runs a first service according to the first historical operation information to obtain first operation information, wherein the first processing core is any processing core among the multiple processing cores; The global predictor caches the first running information.

11. The method according to claim 10, characterized in that The historical operation information includes historical operation information of multiple services, and the historical operation information of the multiple services includes the first historical operation information; the global predictor caches historical operation information of at least one processing core of the multiple processing cores, including: The multiple buffers of the global predictor cache the historical operation information of the multiple services respectively, wherein a first buffer among the multiple buffers is used to cache the first historical operation information.

12. The method according to claim 11, characterized in that The method further comprises: The first processing core obtains configuration information for the first service, where the configuration information is used to indicate the first buffer; The first processing core configures the first buffer for the first service according to the configuration information, wherein configuring the first buffer includes enabling, disabling or clearing the first buffer.

13. The method according to claim 12, characterized in that The size of the first buffer is determined according to the first service.

14. The method according to any one of claims 11 to 13, characterized in that: The plurality of processing cores further include a second processing core, and the method further includes: When the first service is switched from the first processing core to the second processing core, the second processing core obtains the first historical operation information from the first buffer and runs the first service according to the first historical operation information.

15. The method according to any one of claims 10 to 14, characterized in that: The historical operation information or the first operation information includes at least one of the following: branch address pair information, jump information, instruction trace, data trace, cache hot and cold information, scheduling recommendation information, page table information, hardware prefetching regularity information or difficult-to-predict address information.

16. The method according to claim 15, characterized in that The historical operation information or the first operation information further includes a branch block index, and the branch block index is used to index the at least one item.

17. The method according to any one of claims 10 to 16, characterized in that: The first processing core obtains first historical operation information in the historical operation information, including: The buffer of the first processing core obtains at least part of the first historical operation information from the global predictor and caches the at least part.

18. The method according to any one of claims 10 to 17, characterized in that: The first historical operation information is historical operation information of other services that are the same as or belong to the same type as the first service.

19. An electronic device, characterized in that: The electronic device includes a memory and a business processing device, the business processing device includes a global predictor and multiple processing cores, the memory is used to store computer instructions, and the business processing device is used to execute the computer instructions so that the electronic device implements the business processing method as described in any one of claims 10-18.

20. A readable storage medium, characterized in that: The readable storage medium stores computer instructions, and when the computer instructions are executed on a device, the device executes the business processing method according to any one of claims 10 to 18.

Citation Information

Patent Citations

  • Business processing device and method, equipment and storage medium

    CN120196428A

  • Data processing method, processor and electronic equipment

    CN112231243A

  • Instruction processing method applied to multi-core processor and multi-core processor

    CN114518900A

  • Event address register history buffers for supporting profile-guided and dynamic optimizations

    US20090287903A1