Business processing device and method, equipment and storage medium

By introducing a global predictor into the CPU subsystem, the historical operation information and operation information of multiple processing cores is cached and updated, the problem of low efficiency of the CPU subsystem during business processing is solved, and the effect of reducing the cache missing rate, improving prediction accuracy and IPC is achieved.

CN120196428APending Publication Date: 2025-06-24HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311795739.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The central processing unit (CPU) subsystem has problems with low efficiency when performing business processing, which are manifested as high cache missing rate, low branch prediction accuracy, and low instruction number/clock cycle (IPC).

Method used

A global predictor is used to cache the historical operation information of multiple processing cores and update the operation information dynamically to reduce cache missing rates, improve prediction accuracy and IPC.

Benefits of technology

Through the global predictor, the cache missing rate of multiple processing cores when processing services is reduced, the prediction accuracy and IPC are improved, and the operation efficiency of the service is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196428A_ABST
    Figure CN120196428A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a service processing device and method, equipment and a storage medium, relates to the technical field of electronics, and is used for reducing the cache missing rate when a plurality of processing cores process services, improving the prediction accuracy and IPC and further improving the operation efficiency of the services. The service processing device comprises a global predictor and a plurality of processing cores coupled to the global predictor, the global predictor is used for caching historical operation information of at least one processing core in the plurality of processing cores; a first processing core in the plurality of processing cores is used for acquiring first historical operation information in the historical operation information and operating a first service according to the first historical operation information to obtain the first operation information, and the first processing core is any one of the plurality of processing cores; and the global predictor is also used for caching the first operation information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic technologies, and in particular, to a service processing apparatus, method, device, and storage medium. Background Art

[0002] Currently, a central processing unit (CPU) subsystem is generally used for service processing. The CPU subsystem includes multiple processing cores, and the multiple processing cores can be used to execute services such as benchmark performance testing, service performance testing, similar service deployment, or service suite deployment. The scale of instructions and data of the above services is large, while the cache resources of the multiple processing cores in the CPU subsystem are limited. Therefore, the CPU subsystem has a problem of low efficiency when running the above services. Summary of the Invention

[0003] This application provides a service processing apparatus, method, device, and storage medium for improving the operation efficiency of multiple processing cores when running services.

[0004] To achieve the above objective, the embodiments of this application adopt the following technical solutions:

[0005] In a first aspect, a service processing apparatus is provided, including: a global predictor, and multiple processing cores coupled to the global predictor; the global predictor is configured to cache historical operation information of at least one processing core among the multiple processing cores, and the at least one processing core may be part or all of the multiple processing cores; a first processing core among the multiple processing cores is configured to obtain first historical operation information in the historical operation information, and run a first service according to the first historical operation information to obtain first operation information. The first historical operation information may be operation information related to the first service, and the first processing core is any one of the multiple processing cores; the global predictor is further configured to cache the first operation information, that is, the global predictor can dynamically update the historical operation information according to the operation information obtained by any one of the multiple processing cores when running a service.

[0006] In the above technical solution, by caching the historical operation information by the global predictor and caching the operation information generated by any one of the multiple processing cores when running a service, the cached operation information is updated for subsequent use. Any one of the multiple processing cores can obtain the required operation information from the global predictor when running a service. In this way, compared with expanding a larger cache capacity for the multiple processing cores, this solution can reduce the cache miss rate, improve the prediction accuracy and IPC when the multiple processing cores process services, and thus improve the operation efficiency of the service.

[0007] In a possible implementation of the first aspect, the historical operation information includes the historical operation information of multiple services, and the historical operation information of the multiple services includes the first historical operation information. The global predictor includes: multiple buffers for respectively caching the historical operation information of the multiple services, where the first buffer in the multiple buffers is used to cache the first historical operation information. For example, the historical operation information of different services can be correspondingly cached in different buffers. In the above possible implementation, by respectively caching the historical operation information of the multiple services in multiple buffers, any processing core in the multiple processing cores can be prevented from obtaining the historical operation information that does not belong to its own service, and the overflow of the historical operation information in the global predictor can be avoided, thus preventing security information leakage and further improving the security of the historical operation information in the global predictor.

[0008] In a possible implementation of the first aspect, the first processing core is further configured to: obtain configuration information for the first service, where the configuration information is used to indicate the first buffer (for example, used to indicate the size of the first buffer and / or the address space corresponding to the first buffer). The configuration information can be sent by software running on the multiple processing cores or pre-configured in the first processing core; configure the first buffer for the first service according to the configuration information, where configuring the first buffer includes enabling, closing, or clearing the first buffer. For example, the first buffer can be enabled through the configuration information when it is needed, and closed or cleared through the configuration information after use. In the above possible implementation, by using the configuration information to configure a corresponding buffer in the global predictor for a certain service of any processing core in the multiple processing cores, the resources in the global predictor can be allocated to the critical services running on the processing cores, thus avoiding resource bottlenecks and further improving the operation efficiency of the services; in addition, by enabling, closing, or clearing the buffers in the global predictor, the utilization rate of the resources in the global predictor can also be improved.

[0009] In a possible implementation of the first aspect, the size (or dimension) of each buffer in the multiple buffers is determined according to the service to which the historical operation information cached in the buffer belongs. For example, the size of the first buffer is determined according to the first service. In the above possible implementation, determining the size of the buffer used to cache the historical operation information of a certain service according to the service can improve the accuracy of the configured buffer size, thus avoiding problems of resource waste or resource shortage.

[0010] In a possible implementation of the first aspect, the plurality of processing cores further includes: a second processing core, configured to: when a first service is switched from a first processing core to the second processing core, obtain first historical operation information from a first buffer, and operate the first service according to the first historical operation information. In the above possible implementation, when the first service is switched from the first processing core to the second processing core, the second processing core can still obtain the first historical operation information from the first buffer that caches the first service, thereby avoiding a large number of cold misses in the second processing core when the first service is switched.

[0011] In a possible implementation of the first aspect, the historical operation information or the first operation information includes at least one of the following: branch address pair information, jump information, instruction trace, data trace, cache cold and hot information, scheduling recommendation information, page table information, regular information of hardware prefetching, or difficult-to-predict address information. In the above possible implementation, the speculation accuracy rate and speculation efficiency of at least one of the above information can be greatly improved, thereby improving the operation efficiency of the service.

[0012] In a possible implementation of the first aspect, the historical operation information or the first operation information further includes a branch block index, and the branch block index is used to index the at least one; at this time, when the branch block corresponding to the branch block index is accessed, a search of the global predictor can be triggered. In the above possible implementation, the speculation accuracy rate and speculation efficiency of the branch block in the service can be greatly improved, thereby improving the operation efficiency of the service.

[0013] In a possible implementation of the first aspect, the first processing core includes: a buffer, configured to obtain at least a part of the first historical operation information from the global predictor and cache the at least a part. In the above possible implementation, by setting a buffer with a smaller capacity in the first processing core, and dynamically obtaining the operation information required for the operation of the first service from the first historical operation information of the global predictor through the buffer, the influence of the transmission delay of the first historical operation information on the operation efficiency of the first service is avoided, thereby improving the operation efficiency of the first service.

[0014] In a possible implementation of the first aspect, the first historical operation information is the historical operation information of other services that are the same as or belong to the same type as the first service. In the above possible implementation, when the first historical operation information is the historical operation information of other services that are the same as or belong to the same type as the first service, there is a large degree of repetition and similarity between the operation information of the first service and the first historical operation information. At this time, when the first processing core operates the first service according to the first historical operation information, the operation efficiency of the first service can be greatly improved.

[0015] Second aspect, a service processing method is provided, which is applied to a service processing device. The device includes a global predictor and a plurality of processing cores coupled to the global predictor. The method includes: the global predictor caches historical operation information of at least one of the plurality of processing cores; a first processing core among the plurality of processing cores obtains first historical operation information from the historical operation information and runs a first service according to the first historical operation information to obtain first operation information, where the first processing core is any one of the plurality of processing cores; the global predictor caches the first operation information.

[0016] In a possible implementation manner of the second aspect, the historical operation information includes historical operation information of a plurality of services, and the historical operation information of the plurality of services includes the first historical operation information. The global predictor caches the historical operation information of at least one of the plurality of processing cores, including: a plurality of buffers of the global predictor respectively cache the historical operation information of the plurality of services, where a first buffer among the plurality of buffers is used to cache the first historical operation information.

[0017] In a possible implementation manner of the second aspect, the method further includes: the first processing core obtains configuration information for the first service, and the configuration information is used to indicate the first buffer; the first processing core configures the first buffer for the first service according to the configuration information, where configuring the first buffer includes enabling, closing, or clearing the first buffer.

[0018] In a possible implementation manner of the second aspect, the size of the first buffer is determined according to the first service.

[0019] In a possible implementation manner of the second aspect, the plurality of processing cores further includes a second processing core, and the method further includes: when the first service is switched from the first processing core to the second processing core, the second processing core obtains the first historical operation information from the first buffer and runs the first service according to the first historical operation information.

[0020] In a possible implementation manner of the second aspect, the historical operation information or the first operation information includes at least one of the following: branch address pair information, jump information, instruction trace, data trace, cache hot and cold information, scheduling recommendation information, page table information, regular information of hardware prefetching, or difficult-to-predict address information.

[0021] In a possible implementation manner of the second aspect, the historical operation information or the first operation information further includes a branch block index, and the branch block index is used to index the at least one item.

[0022] In a possible implementation of the second aspect, the first processing core obtains the first historical operation information in the historical operation information, including: the buffer of the first processing core obtains at least a part of the first historical operation information from the global predictor and caches the at least a part.

[0023] In a possible implementation of the second aspect, the first historical operation information is the historical operation information of other services that are the same as or belong to the same type as the first service.

[0024] In a third aspect, an electronic device is provided. The electronic device includes a memory and a service processing device. The service processing device includes a global predictor and a plurality of processing cores. The memory is used to store computer instructions, and the service processing device is used to execute the computer instructions so that the electronic device implements the service processing method provided in the second aspect or any possible implementation of the second aspect.

[0025] In yet another aspect of the present application, a computer-readable storage medium is provided. A computer program or instruction is stored in the computer-readable storage medium. When the computer program or instruction is run, the service processing method provided in the second aspect or any possible implementation of the second aspect is implemented.

[0026] In yet another aspect of the present application, a computer program product is provided. The computer program product includes: a computer program (which can also be referred to as code or instruction). When the computer program is run, the computer is caused to execute the service processing method provided in the second aspect or any possible implementation of the second aspect.

[0027] It can be understood that for any of the service processing methods, electronic devices, computer-readable storage media, and computer program products provided above, the beneficial effects that can be achieved can be correspondingly referred to the beneficial effects in the service processing device provided above, and will not be elaborated here. Description of the Drawings

[0028] Figure 1 It is a schematic structural diagram of a multi-core system provided by an embodiment of the present application;

[0029] Figure 2 It is a schematic structural diagram of a service processing device provided by an embodiment of the present application;

[0030] Figure 3 It is a schematic structural diagram of an LBTB provided by an embodiment of the present application;

[0031] Figure 4 It is a schematic structural diagram of another LBTB provided by an embodiment of the present application;

[0032] Figure 5A schematic diagram for a processing core to obtain historical operation information provided by an embodiment of the present application;

[0033] Figure 6 A schematic diagram of the structure of an sBTB provided by an embodiment of the present application;

[0034] Figure 7 A schematic diagram of multiple processing cores sharing global prediction provided by an embodiment of the present application;

[0035] Figure 8 A schematic diagram of configuring a global predictor provided by an embodiment of the present application;

[0036] Figure 9 A schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0037] The fabrication and use of each embodiment will be discussed in detail below. However, it should be understood that many applicable inventive concepts provided by the present application can be implemented in a variety of specific environments. The specific embodiments discussed merely illustrate the specific ways to implement and use the present application and the present technology, and do not limit the scope of the present application.

[0038] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by those of ordinary skill in the art.

[0039] Each circuit or other component may be described as or referred to as "configured to" perform one or more tasks. In this case, "configured to" is used to imply a structure by indicating that the circuit / component includes a structure (such as circuitry) that performs one or more tasks during operation. Thus, even when the specified circuit / component is currently inoperable (e.g., not turned on), the circuit / component can still be referred to as being configured to perform the task. A circuit / component used in conjunction with the phrase "configured to" includes hardware, such as circuitry that performs the operation, etc.

[0040] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. In the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single item(s) or plural item(s). For example, at least one of a, b, or c may represent: a, b, c, a and b, a and c, b and c, a, b, and c; where a, b, and c may be single or multiple.

[0041] In the embodiments of the present application, terms such as "first" and "second" are used to distinguish objects with similar names, functions, or roles. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order. The term "coupled" is used to indicate an electrical connection, including being directly connected through a wire or a connection terminal or being indirectly connected through other devices. Therefore, "coupled" should be regarded as a general electronic communication connection.

[0042] It should be noted that in the present application, words such as "exemplary" or "for example" are used to give examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0043] Currently, the central processing unit (CPU) subsystem usually has the problem of low efficiency when performing business processing. This problem is specifically manifested as: a high cache miss rate, a low branch prediction accuracy rate, and a low instruction per clock (IPC). For example, the scale of instructions and data of the business running on the CPU subsystem is large, while the capacities of resources such as the cache, branch target buffer (BTB), and branch direction predictor in the CPU subsystem are limited, which easily causes more cache misses. For another example, the software in the CPU subsystem is not aware of speculative means such as front-end prefetching and back-end prefetching, resulting in the prefetch information being unable to assist other predictions, thus causing information waste. For yet another example, the number of threads of the application program running on the CPU subsystem is much larger than the number of processing cores (or called cores, also called processor cores) included in the CPU subsystem, which easily causes a large number of thread switches (including switches within the same processing core and switches between different processing cores). When a thread switch occurs, the data stored in resources such as the cache, BTB, and branch direction predictor is not the data required by the current thread, so the current thread will have a large number of cold miss problems.

[0044] Based on this, an embodiment of the present application provides a service processing device. The service processing device includes a global predictor (GP) and multiple processing cores coupled to the global predictor. The global predictor can be used to cache historical operation information. Any one of the multiple processing cores can be used to obtain the required operation information from the historical operation information and operate the service according to the obtained operation information. The global predictor is also used to cache the operation information obtained by running the service. In this way, by caching the historical operation information in the global predictor and the operation information generated by any processing core when operating the service, each of the multiple processing cores can obtain the required operation information from the global predictor when operating the service, which can reduce the cache miss rate of the multiple processing cores when processing the service, improve the prediction accuracy and IPC (instructions per clock cycle), and thus improve the operation efficiency of the service.

[0045] The service processing device provided by the embodiment of the present application can be applied to a multi-core system, which can be referred to as a processor subsystem (for example, a CPU subsystem). The structure of the multi-core system will be introduced and described below.

[0046] Figure 1 FIG. 7 is a schematic structural diagram of a multi-core system provided by an embodiment of the present application. The multi-core system includes a global predictor and multiple processing cores coupled to the global predictor. The multiple processing cores can share the global predictor through software configuration. The multiple processing cores can be used to deploy different services, the same service, or partially the same services. The partially the same services may refer to that there are some same data and / or instructions that two services need to access during operation. In the embodiment of the present application, the multiple processing cores can be used to execute services such as benchmark performance testing, service performance testing, similar service deployment, or service suite deployment. Exemplarily, the service scenarios of the multi-core system include but are not limited to: server cluster benchmark performance testing, server cluster bidding service performance testing, server cluster similar service deployment (for example, high performance computing (HPC)), server cluster service suite deployment, terminal field benchmark performance testing, service switching, service migration, etc.

[0047] The number of the multiple processing cores included in the multi-core system can be configured according to actual needs. For example, the number of the multiple processing cores can be 2, 4, 6, or 8, etc. In addition, the multiple processing cores can be a homogeneous structure (that is, including multiple identical processing cores) or a heterogeneous structure (that is, including different processing cores). In one example, as shown in (a) of FIG. 12, the multi-core system includes: a global predictor and 4 processing cores coupled to the global predictor. The 4 processing cores are respectively denoted as core 0 to core 3, and the 4 processing cores are of a homogeneous structure. In another example, as shown in Figure 1 shown in (a) of FIG. 12, the multi-core system includes: a global predictor and 4 processing cores coupled to the global predictor. The 4 processing cores are respectively denoted as core 0 to core 3, and the 4 processing cores are of a homogeneous structure. In another example, as shown inFigure 1 As shown in (b) therein, the multi-core system includes: a global predictor, and eight processing cores coupled to the global predictor. The eight processing cores are respectively denoted as core 0 to core 7, and the eight processing cores have a homogeneous structure. In another example, as Figure 1 shown in (c) therein, the multi-core system includes: a global predictor, and eight processing cores coupled to the global predictor. The eight processing cores may have a heterogeneous structure, and the eight processing cores may include a super large core 0, large cores 1 to 3, and small cores 4 to 7.

[0048] In the multi-core system, the multiple processing cores may include multiple different processing cores of the same processor, or the multiple processing cores include processing cores of multiple different processors, that is, the multiple processing cores may form one or more processors. Among them, the processor includes but is not limited to: CPU, general-purpose processor, graphics processing unit (GPU), image signal processor (ISP), digital signal processor (DSP), network processing unit (NPU), artificial intelligence (AI) processor, etc. Optionally, the structures and the sizes of the respective hardware resources (such as cache resources) of any two of the multiple processing cores may be the same or different, and the embodiments of the present application do not make specific limitations thereon.

[0049] Further, the multi-core system may further include a memory coupled to the plurality of processing cores, and the memory may be a volatile memory, a non-volatile memory, or a memory including both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. Exemplarily, the RAM may be a static RAM (SRAM), a dynamic random access memory (DRAM), a synchronous DRAM (SDRAM), or a double data rate synchronous DRAM (DDR SDRAM), etc. The memory is not shown in the figure.

[0050] The above multi-core system can be an electronic device, or a system on a chip (SoC) applied to an electronic device, or a chipset including multiple chips, or a module including the SoC or the chipset. The electronic device can be a terminal device or a server. Optionally, the electronic device includes but is not limited to: mobile phone, tablet computer, laptop computer, desktop computer, palmtop computer, ultra-mobile personal computer (umPC), mobile internet device (MID), netbook, camera, camera, wearable device (such as smart watch and smart bracelet, etc.), vehicle-mounted device (such as, car, bicycle, electric vehicle, airplane, ship, train, high-speed rail, etc.), virtual reality (VR) device, augmented reality (AR) device, wireless terminal in industrial control, smart home device (such as, refrigerator, TV, air conditioner, electric meter, etc.), intelligent robot, workshop device, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, or wireless terminal in smart home, flight device (such as, intelligent robot, hot air balloon, drone, airplane), etc.

[0051] Figure 2 FIG. is a schematic structural diagram of a service processing device provided by an embodiment of the present application. The service processing device includes a global predictor 10 and a plurality of processing cores 20 coupled to the global predictor 10. The plurality of processing cores 20 share the global predictor 10, for example, by software configuration, the plurality of processing cores 20 share different resources of the global predictor 10. Figure 2 In this example, the plurality of processing cores 20 include processing core 21 to processing core 2n, where n is an integer greater than 1.

[0052] The global predictor 10 is used to: cache the historical operation information of at least one processing core among the multiple processing cores 20. The at least one processing core includes one or more processing cores. Wherein, when the at least one processing core includes multiple processing cores, the at least one processing core may be a part of the multiple processing cores 20 or all of the multiple processing cores 20. The historical operation information cached in the global predictor 10 may be obtained from the operation of the at least one processing core. For example, each processing core in the at least one processing core may run one or more services, and transmit the obtained operation information to the global predictor 10, and the global predictor 10 caches the operation information transmitted by the at least one processing core as the historical operation information.

[0053] For any one of the multiple processing cores 20 (for convenience of description, hereinafter referred to as the first processing core 21), it is used to: obtain the first historical operation information in the historical operation information, and run the first service according to the first historical operation information to obtain the first operation information. The first historical operation information may be part or all of the historical operation information. Optionally, the first historical operation information is the historical operation information of the same service as the first service or other services of the same type. For example, the first historical operation information may be the historical operation information of the first service, that is, the operation information obtained from the previous operation of the first service; or, the first service and the second service are two services in the same high-performance computing, and the first historical operation information is the historical operation information of the second service. Hereinafter, the case where the first historical operation information is the historical operation information of the first service is taken as an example for description.

[0054] The global predictor 10 is further used to: cache the first operation information. For example, when the first processing core 21 obtains the first operation information, it transmits the first operation information to the global predictor 10, and the global predictor 10 receives and caches the first operation information. Wherein, after the global predictor 10 caches the first operation information, when any one of the multiple processing cores 20 (for example, the first processing core 21 or other processing cores) obtains the historical operation information from the global predictor 10, the historical operation information includes the first operation information. That is to say, the historical operation information is dynamically updated, and the global predictor 10 can dynamically update the historical operation information according to the operation information obtained from the operation of the multiple processing cores 20.

[0055] Optionally, the historical operation information or the first operation information includes at least one of the following: branch address pair information, jump information, instruction trace, data trace, cache hot and cold information, scheduling recommendation information, page table information, regular information of hardware prefetching, or hard-to-predict address information. Among them, the branch address pair information refers to the address of the executed branch in the memory, and the address includes a source address and a target address, and the source address and the target address form an address pair. The jump information may include a jump direction and a jump destination. The jump direction is used to indicate the jump direction of the jump instruction (i.e., jump or not jump), and the jump destination refers to the destination address after the jump. The data trace is used to indicate the trace of the accessed data. The instruction trace is used to indicate the trace of the accessed instruction. The cache hot and cold information is used to indicate the hot and cold degree of the data and / or instructions in the cache (such as the first-level, second-level, and third-level caches). The scheduling recommendation information is used to indicate the information recommended during the scheduling process, such as the scheduling policy recommended based on quality of service (QoS). The page table information is used to indicate the correspondence between the virtual address and the physical address corresponding to the accessed data and / or instructions. The regular information of hardware prefetching is used to indicate the rule when the hardware prefetches data and / or instructions, and the prefetching may include front-end prefetching and / or back-end prefetching. The hard-to-predict address information is used to indicate the address information that is hard to predict by the hardware prefetch (HWP). The memory targeted by the above address may be on the same or different chips as the above one or more processing cores, and may be a volatile or non-volatile memory, and this embodiment does not make a limitation.

[0056] It can be understood that, in addition to the above information, the above historical operation information or the first operation information may further include other information required during the business operation process (such as value prediction information, which may refer to the value obtained by prediction when the memory access is not completed), and the embodiments of the present application do not make specific limitations on this.

[0057] Optionally, the above historical operation information or the first operation information further includes a branch block index, and the branch block index is used to index at least one of the above information. At this time, when the branch block corresponding to the branch block index is accessed, the search of the global predictor 10 can be triggered. Among them, the branch block index may also be referred to as a branch block program counter (PC) index (block PC index). For ease of understanding, the following takes the historical operation information as an example to illustrate the branch block index included in the historical operation information and the first operation information and at least one item corresponding to the branch block index.

[0058] In one example, it is assumed that the historical operation information includes the historical operation information of multiple services. The historical operation information of each service may include at least one branch block index and the corresponding historical operation sub-information for each branch block index. Each historical operation sub-information may include at least one of the following: branch address pair information, jump information, instruction trace, data trace, cache hot and cold information, scheduling recommendation information, page table information, regular information of hardware prefetching, or difficult-to-predict address information.

[0059] Among them, each service in the above multiple services may include at least one branch block. Each branch block (or called subroutine or program segment) in the at least one branch block may correspond to a branch block index. After running, the branch block can correspond to a historical operation sub-information. The branch block index of a branch block can be used to index the historical operation sub-information obtained after the branch block runs. Optionally, the branch block index may be the instruction fetch address, and the instruction fetch address may be the source address corresponding to the branch block.

[0060] Exemplarily, the historical operation information may be as shown in Table 1 below. In Table 1, it is assumed that the multiple services include M services, and the M services are represented as Service 1 to Service M. The at least one branch block index of Service 1 is represented as ID1-1, ID1-2, …, the at least one branch block index of Service 2 is represented as ID2-1, ID2-2, …, and the at least one branch block index of Service M is represented as IDM-1, IDM-2, …

[0061] Table 1

[0062]

[0063] It can be understood that the information types included in the historical operation sub-information corresponding to different branch block indexes are the same in Table 1 above for illustration. In actual applications, the information types included in the historical operation sub-information corresponding to different branch block indexes may also be different, and the embodiments of the present application do not make specific limitations on this.

[0064] In addition, the global predictor 10 may further include other information of each service in the multiple services. For example, the other information may include service identifier and / or service context, etc., and the embodiments of the present application do not make specific limitations on this.

[0065] Further, the global predictor 10 includes a plurality of buffers. The plurality of buffers are used to: cache the historical operation information of the plurality of services respectively. Among them, each buffer in the plurality of buffers can be used to: cache the historical operation information obtained by the operation of one processing core, that is, the historical operation information obtained by the operation of different processing cores can be cached in different buffers; or, used to cache the historical operation information of one service, that is, the historical operation information of different services can be cached in different buffers.

[0066] Optionally, the plurality of buffers include a first buffer, and the first buffer is used to cache the historical operation information of the first service. The historical operation information of the first service may include first historical operation information. In a possible example, for any one of the plurality of processing cores 20 other than the first processing core 21 (referred to as the second processing core 22 in this article), when the first service switches from the first thread of the first processing core 21 to the second thread of the second processing core 22, the second processing core 22 can be used to obtain the first historical operation information from the first buffer and run the first service according to the first historical operation information. In this example, when the first service switches from the first processing core 21 to the second processing core 22, the second processing core 22 can still obtain the first historical operation information from the first buffer that caches the first service, so that when the first service switches, the problem of a large number of cold misses in the second processing core 22 is avoided.

[0067] Optionally, the size (or dimension) of each buffer in the plurality of buffers can be determined according to the service to which the cached historical operation information belongs. Exemplarily, the size of the first buffer is determined according to the first service. That is, the size of the buffer used to cache the historical operation information of one service can be dynamically determined according to the size of the service. For example, in practical applications, specifically, the software running on the plurality of processing cores 20 (such as the operating system) can dynamically determine the corresponding buffer size for each service, and the embodiments of the present application do not make specific limitations on this.

[0068] In a possible embodiment, the global predictor 10 includes a multi-way buffer, and each buffer in the above-mentioned plurality of buffers can include at least one way of buffer in the multi-way buffer. Optionally, each way of buffer in the multi-way buffer can be used to cache the historical operation information of one service, the historical operation information of different services is cached in different buffers, and the historical operation information of the same service can occupy one or more ways of buffer. In addition, the multi-way buffers are independent of each other, each way of buffer can be accessed separately, and the accesses corresponding to different ways of buffer do not affect each other. Exemplarily, the multi-way buffer includes a first way of buffer, and the first way of buffer is used to cache the historical operation information of the first service.

[0069] When the historical operation information of a certain service includes at least one of the above multiple pieces of information such as branch address pair information, jump information, and instruction trace, for any one of the at least one piece of information, each way buffer in the global predictor 10 can be cached through the following structure. For ease of description, in the following text, it is described by taking the first processing core 21 running the first service and the first way buffer being used to cache the historical operation information of the first service, and the historical operation information of the first service including branch address pair information as an example. Among them, the first way buffer can also be called a link branch target buffer (LBTB).

[0070] In one example, as Figure 3 shown, the LBTB may include: a history queue, a storage circuit, a search queue, and a fill queue. The history queue, the search queue, and the fill queue are all coupled to the storage circuit.

[0071] The history queue is used for: when a branch prediction target address fails during the operation of the first service (represented as a prediction failure in the figure, and the prediction failure may refer to that the target address corresponding to the branch cannot be found in the cache of the first processing core 21 (for example, the first buffer and the second buffer in the following text)), obtaining the correct branch address pair (that is, including the source address and the target address) from the pipeline corresponding to the first service, and outputting the branch address pair to the storage circuit. Optionally, in order to reduce the number of updates to the LBTB to reduce power consumption overhead and access conflicts, the history queue can be used to output a preset number of branch address pairs (for example, 4 address pairs) as a group to the storage circuit according to a certain storage format when the preset number of branch address pairs is accumulated. Figure 3 Taking the preset number of branch address pairs including brn add0-tgtadd0, brn add1-tgt add1, brn add2-tgt add2, brn add3-tgt add3, and the storage format also including a link address link add as an example, the link add is used to indicate the next group of branch address pairs.

[0072] The storage circuit is used for: receiving and caching the branch address pairs output by the history queue. Among them, when caching a group of branch address pairs in the above storage format, the storage circuit can index with the source address of the first branch address pair and write the group into the corresponding storage. Optionally, the storage circuit may include a multi-way connected RAM, so that the search efficiency can be improved when searching the storage circuit.

[0073] The search queue is used to: obtain the instruction fetch address of the branch block and output the instruction fetch address to the storage circuit. Among them, the search queue can receive and filter duplicate address information; in addition, the search queue can be a multi-input and single-output queue, so as to achieve the effect of balancing the bandwidth difference between input and output. Optionally, the search queue can obtain the instruction fetch address when the first processing core 21 meets certain conditions. For example, the conditions can include but are not limited to: querying in the main branch target buffer (mBTB) and missing, or querying in the stream branch target buffer (sBTB) and hitting, etc. For the detailed description of mBTB and sBTB, reference can be specifically made to the description in the following text, and the embodiments of the present application will not be elaborated here.

[0074] The storage circuit is further used to: when receiving the instruction fetch address of the branch block of the first processing core 21, obtain the corresponding branch address pair according to the instruction fetch address and output the branch address pair to the fill queue. Optionally, when caching a group of branch address pairs in the above storage format, the storage circuit can disassemble the obtained group of branch addresses and output a preset number of disassembled branch address pairs to the fill queue, and output the link add in the group of branch address pairs to the search queue, so that the search queue can perform the next search based on the link add.

[0075] The fill queue is used to: fill the branch address pair output by the storage circuit back into the first processing core 21. Optionally, the fill queue can fill a preset number of branch address pairs output by the storage circuit into the buffer of the first processing core 21, for example, fill them into the sBTB of the first processing core 21.

[0076] In another example, in combination with Figure 3 , as Figure 4 shown, for each processing core in the multiple processing cores 20, the global predictor 10 can include a history queue, a storage circuit, a search queue, and a fill queue corresponding to each processing core. Among them, different storage circuits in the global predictor 10 can be integrated together as a storage module; or, the global predictor includes a storage module, the storage module includes a storage circuit corresponding to each processing core, and the storage circuit includes a multi-way cache. Figure 4Among them, the multiple processing cores 20 include 4 processing cores, which are respectively denoted as core 0 to core 3. The history queue 0, search queue 0, backfill queue 0, and two-way caches w0 and w1 corresponding to core 0, the history queue 1, search queue 1, backfill queue 1, and two-way caches w2 and w3 corresponding to core 1, the history queue 2, search queue 2, backfill queue 2, and two-way caches w4 and w5 corresponding to core 2, and the history queue 3, search queue 3, backfill queue 3, and four-way caches w6 to w9 corresponding to core 3.

[0077] Further, taking the first processing core 21 obtaining the first historical running information from the global predictor 10 as an example, the related process of any one of the multiple processing cores 20 obtaining the historical running information from the global predictor 10 will be introduced and described.

[0078] In a possible embodiment, the first processing core 21 includes a first buffer, and the first buffer is used to obtain at least part of the first historical running information from the global predictor 10 and cache the at least part. Among them, the first processing core 21 may further include a second buffer, which is different from the first buffer. The second buffer may refer to a buffer set in the first processing core 21 for caching prefetch information, and the second buffer is not used to cache the historical running information obtained from the global predictor 10. In the embodiments of the present application, the first buffer may also be referred to as a candidate buffer (such as sBTB), and the second buffer may be referred to as a main buffer (such as mBTB).

[0079] Optionally, the first buffer is specifically used to obtain at least part of the first historical running information from the global predictor 10 and cache the at least part when certain conditions are met. For example, the conditions may include: querying in the second buffer and missing (that is, the second buffer does not have the required information), or querying in the first buffer and hitting (the first buffer has the required information), etc. The embodiments of the present application do not make specific limitations on this.

[0080] Among them, when the first historical running information includes at least one of the above multiple pieces of information such as branch address pair information, jump information, and instruction trace, the number of the first buffers of the first processing core 21 may be at least one, that is, the first processing core 21 includes at least one first buffer, and each of the at least one first buffers can be used to cache one of the above multiple pieces of historical running information. For ease of description, in the following, it is taken as an example that the first historical running information includes branch address pair information, and the first processing core 21 includes a main branch target buffer mBTB (corresponding to the second buffer) and a streaming branch target buffer sBTB (corresponding to the first buffer).

[0081] Exemplarily, such as Figure 5As shown in the figure, the first processing core 21 includes an mBTB and an sBTB. The mBTB caches a set of first branch address pairs corresponding to the first service prefetched by the first processing core 21, and the sBTB caches a set of second branch address pairs obtained from the LBTB of the global predictor 10. During the operation of the first service by the first processing core 21, when the first processing core 21 obtains the fetch instruction address FIVA through instruction fetching, the first processing core 21 can first query whether there is a target branch address corresponding to the FIVA in the set of first branch address pairs cached in the mBTB; if it exists (i.e., a hit), it continues to execute according to the target branch address; if it does not exist (i.e., a miss), it queries whether there is a target branch address corresponding to the FIVA in the set of second branch address pairs cached in the sBTB, and if it exists (i.e., a hit), it continues to execute according to the queried target branch address. In addition, when the query in the mBTB misses, or when the query in the sBTB hits, it indicates that the historical operation information as prefetch information is accurate and available. The sBTB can obtain more branch address pairs corresponding to the first service from the LBTB of the global predictor 10. For example, the sBTB sends the next fetch instruction address to the global predictor 10 to obtain the more branch address pairs through the next fetch instruction address for the first processing core 21 to use when running the first service. Optionally, when the first service ends, or when other services need to use the sBTB, the first processing core 21 can also clear (or flush) the sBTB. For example, the first processing core 21 can clear the sBTB through a valid bit.

[0082] It can be understood that the conditions for the sBTB to obtain branch address pairs from the LBTB of the global predictor 10 can also include other conditions, such as obtaining at a certain time duration, or obtaining when all pipeline jumps of the first processing core 21 occur, etc. The embodiments of the present application do not make specific limitations on this.

[0083] In addition, Figure 6 shows a possible structural schematic diagram of the above sBTB. As Figure 6As shown, the sBTB may include a storage circuit and a valid bit array. The storage circuit can be used to cache historical operation information from the global predictor 10, and the valid bit array can be used to indicate the validity of the information cached in the storage circuit. Among them, the storage circuit may include a multi-way connected RAM, so that the search efficiency can be improved when searching the storage circuit. Optionally, the sBTB may further include a matching and selection circuit, which can be used to output when the service to which the FIVA belongs matches the service currently cached in the sBTB and the target branch address hit in the sBTB is valid. In addition, the output end of the mBTB and the output end of the sBTB can also be coupled through a selector, which is used to select the output result of the mBTB when the mBTB hits, and select the output result of the sBTB when the sBTB hits.

[0084] Optionally, when the first processing core 21 queries the first buffer and hits, the first processing core 21 may reload the hit operation information as speculative information into the modules or components required by the pipeline. Exemplarily, load the branch address pair back into the branch pipeline, load the cold / hot and criticality information of the instruction cache and data cache back into the instruction prefetch and replacement components, load the historical operation information of the page table back into the page table prefetch component, load the regular information of hardware prefetch and difficult-to-predict address information back into the hardware prefetch component, etc.

[0085] Further, in a possible example, as Figure 7 shown, when the historical operation information cached in the global predictor includes multiple items of information such as branch address pair information, jump information, and instruction trace, the second buffer in any one of the multiple processing cores 21 (for example, core 0, core 1, and core 2) (for example, core 0) for caching the multiple items of information may include a branch target buffer (BTB), an instruction cache (Icache), a data cache (Dcache), a translation lookaside buffer (TLB), etc.; the processing core may further include other devices such as an execution pipeline. Figure 7 This example illustrates that the processing core caches the multiple items of information obtained from the global predictor 10 through the first buffer, and the first buffer is a stream buffer.

[0086] Correspondingly, in an example, as Figure 7As shown, the global predictor 10 includes a multiplex buffer (e.g., an N-way buffer denoted as GP w0 to GP wN), and each buffer can be used to cache the historical running information of a service. The historical running information of any service can include multiple branch block indexes and the above-mentioned multiple pieces of information corresponding to each branch block index. Figure 7 Taking the branch block indexes in the historical running information of service 1 in Figure 7 as an example, which include block-ID1 and block-ID2. When the same service switches between different processing cores, the processing core where the service is located after the switch can share the historical running information of the service in the global predictor 10.

[0087] After introducing the relevant structures of the global predictor 10 and the multiple processing cores 20, the process of allocating resources in the global predictor 10 to any one of the multiple processing cores 20 will be introduced in detail below.

[0088] Optionally, the resources in the global predictor 10 used by any one of the multiple processing cores 20 can be obtained through configuration. When any one of the multiple processing cores 20 needs to use the global predictor 10 to cache the historical running information of a certain service, the global predictor 10 can be configured for the processing core. For example, a corresponding buffer can be configured for the processing core in the global predictor 10. In addition, the services that can use the global predictor 10 can also be configured for the multiple processing cores 20, that is, some key services can be configured for any one of the multiple processing cores 20, and only the historical running information of the configured key services can be cached in the global predictor. The process of configuring the global predictor will be introduced below taking the first processing core 21 as an example.

[0089] In a possible embodiment, the first processing core 21 is further configured to: obtain configuration information for the first service, where the configuration information is used to indicate the first buffer; configure the first buffer for the first service according to the configuration information. Optionally, configuring the first buffer may include enabling, closing, or clearing the first buffer. The configuration information may be sent to the first processing core 21 by software running on the multiple processing cores 20, or may be pre-configured in the first processing core 21. The embodiments of the present application do not make specific limitations on this. The configuration information can be used to indicate the size of the first buffer and / or the address space corresponding to the first buffer. Further, the configuration information can also be used to indicate the first service. For example, the configuration information may include the service identifier of the first service.

[0090] For ease of understanding, as Figure 8 shown, the process of enabling, closing, and clearing the global predictor 10 through configuration information will be exemplified below when the first processing core 21 is the core 0, core 1, and core 2 respectively. Figure 8Taking the example where core 0, core 1, and core 2 all include HWP and SBTB for illustration.

[0091] In one example, assuming that the first processing core 21 is core 0, enabling the global predictor 10 through configuration information includes: S11. When it is necessary to run service a on core 0, the software running on the multiple processing cores 20 can send the first configuration information for service a to core 0; S12. When core 0 receives the first configuration information, the HWP of core 0 enables the global predictor 10 according to the first configuration information. For example, it sends service information and enabling information to the global predictor 10 to enable the global predictor 10 to cache the historical running information of service a; S13. Core 0 can also enable sBTB when the information currently cached in sBTB matches service a; S14. The global predictor 10 starts to work.

[0092] In another example, assuming that the first processing core 21 is core 1, disabling the global predictor 10 through configuration information includes: S21. When it is necessary to disable the use of the global predictor 10 for service b running on core 1, the software running on the multiple processing cores 20 can send the second configuration information for service b to core 1; S22. When receiving the second configuration information, the HWP of core 1 disables the global predictor 10 according to the second configuration information. For example, it sends service information and disabling information to the global predictor 10; S23. The global predictor 10 stops providing services for service b on core 1, that is, it stops providing the historical running information of the service to core 1.

[0093] In yet another example, assuming that the first processing core 21 is core 2, clearing the global predictor 10 through configuration information includes: S31. When it is necessary to clear the use of the global predictor 10 for service c running on core 2, the software running on the multiple processing cores 20 can send the third configuration information for service c to core 2; S32. When receiving the third configuration information, the HWP of core 2 clears the global predictor 10 according to the third configuration information. For example, it sends thread information and clearing information to the global predictor 10; S33. The global predictor 10 clears the information corresponding to the thread, that is, it clears the historical running information of service c. Optionally, the global predictor 10 can use the method of clearing line by line for clearing.

[0094] In the embodiment of the present application, the global predictor 10 caches the historical operation information and the operation information generated by any processing core when running a service, so as to update the cached operation information for subsequent use. When the multiple processing cores 20 run a service, they obtain the required operation information from the global predictor 10 respectively. In this way, it is equivalent to expanding a larger-capacity cache for the multiple processing cores, thereby reducing the cache miss rate of the multiple processing cores 20 when processing services, improving the prediction accuracy and IPC, and further improving the operation efficiency of the service. For the service switching scenario, since the global predictor 10 can be shared by the multiple processing cores 20, when the service is switched from one processing core to another processing core, the other processing core can still obtain the historical operation information of the service from the global predictor 10 to run the service, thus avoiding the problem of a large number of cold misses in the other processing core. In addition, the resources in the global predictor 10 used by any one of the multiple processing cores 20 are obtained through software configuration, and the services that can use the global predictor 10 can also be configured for the multiple processing cores 20. In this way, it can avoid the leakage of security information caused by the overflow of information in the global predictor 10, and at the same time, key services can be identified from the software perspective, and the resources in the global predictor 10 can be configured for the key services running on the processing core, thereby avoiding resource bottlenecks and further improving the operation efficiency of the service.

[0095] Based on this, the embodiment of the present application further provides a service processing method, which can be applied to a service processing device including a global predictor and a plurality of processing cores coupled to the global predictor. For the description of the service processing device, reference can be made to the above description. The method includes: the global predictor caches the historical operation information of at least one of the multiple processing cores; the first processing core among the multiple processing cores obtains the first historical operation information in the historical operation information, and runs the first service according to the first historical operation information to obtain the first operation information; the global predictor caches the first operation information.

[0096] Optionally, the historical operation information includes the historical operation information of multiple services, and the historical operation information of the multiple services includes the first historical operation information.

[0097] In a possible embodiment, the global predictor includes a plurality of buffers, and the plurality of buffers respectively cache the historical operation information of the plurality of services, wherein the first buffer among the plurality of buffers caches the first historical operation information. Correspondingly, the method may further include: the first processing core obtains the configuration information for the first service, and the configuration information is used to indicate the first buffer; the first processing core configures the first buffer for the first service according to the configuration information. Wherein, configuring the first buffer includes enabling, closing or clearing the first buffer.

[0098] It can be understood that all the content in the above device embodiments can be cited in the corresponding embodiments of the service processing method, and will not be elaborated herein in the embodiments of the present application.

[0099] In the embodiments of the present application, the global predictor caches the historical operation information and the operation information generated by any processing core when running the service, so as to update the cached operation information for subsequent use. Any one of the multiple processing cores can obtain the required operation information from the global predictor when running the service. In this way, compared with expanding a larger capacity cache for the multiple processing cores, this solution can reduce the cache miss rate when the multiple processing cores process the service, improve the prediction accuracy and IPC, and thus improve the operation efficiency of the service.

[0100] In another aspect of the present application, an electronic device is further provided, as Figure 9 shown. The electronic device includes a memory and a service processing device. The service processing device includes a global predictor and multiple processing cores. The memory is used to store computer instructions, and the service processing device is used to execute the computer instructions so that the electronic device implements any one of the service processing methods provided above. For the specific introduction of the memory, reference can be made to the previous embodiments.

[0101] It can be understood that all the relevant content of each step involved in the above method embodiments can be cited in the embodiments of the service processing method and the embodiments of the electronic device, and will not be elaborated herein in the embodiments of the present application.

[0102] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0103] The units described as separate components may or may not be physically separated. The components shown as units may be a physical unit or multiple physical units, that is, they can be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0104] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium, which may include various media capable of storing program codes, such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs. Based on such understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product.

[0105] In another embodiment of the present application, a readable storage medium is further provided. Computer-executable instructions are stored in the readable storage medium. When a device (which may be a single-chip microcomputer, a chip, etc.) or a processor executes the steps in the above method embodiment.

[0106] In yet another embodiment of the present application, a computer program product is further provided. The computer program product includes computer instructions, and the computer instructions are stored in a readable storage medium; at least one processor of the device can read the computer instructions from the readable storage medium, and at least one processor executes the computer instructions to enable the device to perform the steps in the above method embodiment.

[0107] Finally, it should be noted that the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A service processing device, characterized in that, Including: A global predictor and a plurality of processing cores coupled to the global predictor; wherein, The global predictor is configured to cache historical operation information of at least one of the plurality of processing cores; A first processing core among the plurality of processing cores is configured to obtain first historical operation information from the historical operation information and run a first service according to the first historical operation information to obtain first operation information, and the first processing core is any one of the plurality of processing cores; The global predictor is further configured to cache the first operation information.

2. The device according to claim 1, wherein The historical operation information includes historical operation information of a plurality of services, and the historical operation information of the plurality of services includes the first historical operation information; The global predictor includes: a plurality of buffers configured to respectively cache the historical operation information of the plurality of services, wherein a first buffer among the plurality of buffers is configured to cache the first historical operation information.

3. The device according to claim 2, characterized in that The first processing core is further configured to: Obtain configuration information for the first service, and the configuration information is used to indicate the first buffer; Configure the first buffer for the first service according to the configuration information, and configuring the first buffer includes enabling, closing, or clearing the first buffer.

4. The device according to claim 3, characterized in that, The size of the first buffer is determined according to the first service.

5. The device according to any one of claims 2-4, characterized in that The plurality of processing cores further includes: A second processing core configured to, when the first service is switched from the first processing core to the second processing core, obtain the first historical operation information from the first buffer and run the first service according to the first historical operation information.

6. The device according to any one of claims 1-5, characterized in that, The historical operation information or the first operation information includes at least one of the following: branch address pair information, jump information, instruction trace, data trace, cache cold / hot information, scheduling recommendation information, page table information, regular information of hardware prefetching, or difficult-to-predict address information.

7. The device according to claim 6, wherein The historical operation information or the first operation information further includes a branch block index, and the branch block index is used to index the at least one item.

8. The device according to any one of claims 1-7, characterized in that, The first processing core includes: A buffer configured to obtain at least a part of the first historical operation information from the global predictor and cache the at least a part.

9. The device according to any one of claims 1-8, characterized in that, The first historical operation information is historical operation information of a service that is the same as or belongs to the same type as the first service.

10. A service processing method, characterized in that, Applied to a service processing device, the device includes a global predictor and a plurality of processing cores coupled to the global predictor; the method includes: The global predictor caches historical operation information of at least one of the plurality of processing cores; A first processing core among the plurality of processing cores obtains first historical operation information from the historical operation information and runs a first service according to the first historical operation information to obtain first operation information, and the first processing core is any one of the plurality of processing cores; The global predictor caches the first operation information.

11. The method according to claim 10, characterized in that, The historical operation information includes historical operation information of a plurality of services, and the historical operation information of the plurality of services includes the first historical operation information; the global predictor caching historical operation information of at least one of the plurality of processing cores includes: The multiple buffers of the global predictor respectively cache the historical operation information of the multiple services, wherein the first buffer among the multiple buffers is used to cache the first historical operation information.

12. The method according to claim 11, wherein The method further includes: The first processing core obtains configuration information for the first service, and the configuration information is used to indicate the first buffer; The first processing core configures the first buffer for the first service according to the configuration information, wherein configuring the first buffer includes enabling, closing, or clearing the first buffer.

13. The method according to claim 12, characterized in that, The size of the first buffer is determined according to the first service.

14. The method according to any one of claims 11-13, characterized in that, The multiple processing cores further include a second processing core, and the method further includes: When the first service is switched from the first processing core to the second processing core, the second processing core obtains the first historical operation information from the first buffer and runs the first service according to the first historical operation information.

15. The method according to any one of claims 10 - 14, characterized in that, The historical operation information or the first operation information includes at least one of the following: branch address pair information, jump information, instruction trace, data trace, cache hot and cold information, scheduling recommendation information, page table information, regular information of hardware prefetch, or difficult-to-predict address information.

16. The method according to claim 15, characterized in that, The historical operation information or the first operation information further includes a branch block index, and the branch block index is used to index the at least one item.

17. The method according to any one of claims 10 - 16, characterized in that, The first processing core obtains the first historical operation information in the historical operation information, including: The buffer of the first processing core obtains at least a part of the first historical operation information from the global predictor and caches the at least a part.

18. The method according to any one of claims 10-17, characterized in that, The first historical operation information is the historical operation information of other services that are the same as or belong to the same type as the first service.

19. An electronic device, characterized in that, The electronic device includes a memory and a service processing device, the service processing device includes a global predictor and multiple processing cores, the memory is used to store computer instructions, and the service processing device is used to execute the computer instructions so that the electronic device implements the service processing method according to any one of claims 10-18.

20. A readable storage medium, characterized in that, The readable storage medium stores computer instructions, and when the computer instructions run on the device, the device is caused to execute the service processing method according to any one of claims 10-18.

Citation Information

Cited By

  • Service processing apparatus and method, and device and storage medium

    WO2025130918A1