Service processing method and apparatus for plurality of processing cores, and device

By introducing the concepts of leadership core and follow-up core in the multi-core system, the operation information of leadership cores improves the speculative accuracy and efficiency of follow-up cores, the problem of cache and bandwidth resource bottlenecks in multi-core systems is solved and the overall business processing performance is improved.

WO2025130916A1PCT designated stage expired Publication Date: 2025-06-26HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/140271
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-18
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In multi-core systems, cache resources and bandwidth resources often become bottlenecks in peak performance, and speculative methods of single-core systems are not effective in multi-core systems.

Method used

By introducing the concepts of leadership and follow-up cores in a multi-core system, the leadership core operates faster, and the follow-up core receives operational information from the leadership core and operates business based on this information, thereby improving the accuracy and efficiency of speculation.

Benefits of technology

It improves the accuracy and efficiency of speculation in multi-core systems, improves the overall performance of business processing, and reduces the power consumption and overhead caused by traditional speculation failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024140271_26062025_PF_FP_ABST
    Figure CN2024140271_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A service processing method and apparatus for a plurality of processing cores, and a device, which relate to the technical field of electronics, and are used for improving the accuracy of speculation in a multi-core system, thereby improving the performance of service processing. The method comprises: a first processing core acquiring first running information, wherein the first running information is obtained by means of running of a first service by the first processing core; the first processing core receiving second running information from a second processing core, wherein the second running information is obtained by means of running of a second service by a second processing core; and when it is determined on the basis of the first running information and the second running information that the first service and the second service are at least partially the same, the first processing core running the first service on the basis of the second running information. Since second running information is running information that is obtained by means of actual running performed by a second processing core, and a first service and a second service are partially the same, a first processing core runs the first service on the basis of the second running information, so that the speculation accuracy and speculation efficiency can be greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

A business processing method, device and equipment for multiple processing cores

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 22, 2023, with application number 202311794035.2 and application name “A business processing method, device and equipment for multiple processing cores”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of electronic technology, and in particular to a method, apparatus, and device for processing services using multiple processing cores. Background Art

[0003] During business processing, the cache and bandwidth resources of multi-core systems often become bottlenecks during peak performance. Currently, speculative methods such as front-end prefetching and back-end prefetching are commonly used in single-core systems to improve business processing performance. However, these speculative methods incur additional bandwidth and cache overhead. For example, front-end prefetch errors can cause invalid instructions to be fetched on the wrong path and generate invalid memory accesses. Back-end prefetching generates a large number of memory access requests for speculative addresses, which pollutes the cache and creates bandwidth bottlenecks. Therefore, the speculative methods used in single-core systems are not very effective when applied to multi-core systems. Summary of the Invention

[0004] The present application provides a business processing method, apparatus and device for multiple processing cores, which are used to improve the accuracy of speculation in a multi-core system and thereby improve the performance of business processing.

[0005] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:

[0006] In a first aspect, a business processing method for multiple processing cores is provided, where the multiple processing cores may be multiple processing cores in a multi-core system, the method comprising: a first processing core acquiring first operating information, where the first operating information is obtained by the first processing core running a first business, and the first processing core is a follower core or a slave core; the first processing core receiving second operating information from a second processing core, where the second operating information is obtained by the second processing core running a second business, and the second processing core is a processing core with a faster operating speed among the multiple processing cores, and the second processing core is called a leading core or a main core, and the operating speed of the first processing core is less than the operating speed of the second processing core; when it is determined based on the first operating information and the second operating information that the first business and the second business are at least partially identical, the first processing core runs the first business based on the second operating information.

[0007] In the above technical solution, since the second operating information is the operating information obtained by the actual operation of the second processing core, and the first business and the second business are partially the same, that is, the operation of the first business and the operation of the second business are repetitive and similar, when the first processing core runs the first business according to the second operating information, it can greatly improve the accuracy and efficiency of speculation, thereby improving the overall performance of multiple processing cores in processing business and reducing the power consumption overhead caused by the failure of traditional speculation.

[0008] In a possible implementation of the first aspect, the first operation information and the second operation information include at least one of the following: data address information, instruction address information, branch address information, jump information, or value prediction information. In the above possible implementation, when the second operation information includes data address information, the first processing core runs the first business according to the second operation information, which can greatly reduce the data cache miss rate; when the second operation information includes instruction address information, the first processing core runs the first business according to the second operation information, which can greatly reduce the instruction cache miss rate; when the second operation information includes branch address information, the first processing core runs the first business according to the second operation information, which can greatly reduce the branch prediction failure rate; when the second operation information includes jump information or value prediction information, the first processing core runs the first business according to the second operation information, which can greatly improve the accuracy and efficiency of speculation, thereby improving the overall performance of the business.

[0009] In a possible implementation of the first aspect, the first operation information includes multiple first address information, and the second operation information includes multiple second address information; the first processing core determines that the first business and the second business are at least partially identical based on the first operation information and the second operation information, including: when the number of identical address information in the multiple first address information and the multiple second address information is greater than a preset threshold, the first processing core determines that the first business and the second business are at least partially identical; or, when the number of identical address offset values ​​in the address offset values ​​of the multiple second address information relative to the multiple first address information is greater than a preset threshold, the first processing core determines that the first business and the second business are at least partially identical, wherein the address offset values ​​of the multiple second address information relative to the multiple first address information include the address offset value between each second address information and each first address information. In the above possible implementation methods, when the first processing core determines whether the first business and the second business are at least partially identical, it can be determined based on the number of identical address information in the multiple first address information included in the first operation information and the multiple second address information included in the second operation information, or based on the number of identical address offset values ​​in the address offset values ​​of the multiple second address information relative to the multiple first address information, thereby covering more dynamic address allocation scenarios and improving the coverage of benefits.

[0010] In a possible implementation of the first aspect, the method further includes: the first processing core receives security information from the second processing core, and verifies the security of the second processing core based on the security information. Optionally, the security information includes at least one of the following: an address space identifier ASID, a virtual machine identifier VMID, or privilege level information. Optionally, the security information can be transmitted synchronously with the data packet of the second operation information to reduce the complexity of the transmission of the security information and the second operation information; or, the security information is transmitted asynchronously through a data packet of a specific format, that is, the security information and the second operation information are transmitted through different data packets to reduce the length of the data packet through asynchronous transmission. In the above possible implementation, the first processing core can improve the security of speculation by performing a security check on the second processing core.

[0011] In a possible implementation of the first aspect, the first processing core receiving the second operation information from the second processing core includes: the first processing core receiving the second operation information from the second processing core via a bus, where the bus may be a multiplexed bus or a newly added bus.

[0012] In a possible implementation of the first aspect, the operating frequency of a first processing core differs from the operating frequency of a second processing core, and the first processing core receives first operating information from the second processing core via a bus, including: an asynchronous interface of the first processing core receives second operating information from the second processing core via the bus, and performs asynchronous logical processing on the second operating information. In this possible implementation, when the frequencies of the first processing core and the second processing core differ, synchronization processing can be performed to ensure that the first processing core can still execute the first service based on the second operating information.

[0013] In one possible implementation of the first aspect, the first processing core includes a first cache. Before the first processing core executes the first service based on the second operation information, the method further includes: the first processing core writing the second operation information into the first cache as prefetch information for the first service. This possible implementation can significantly improve the accuracy and efficiency of speculation, thereby improving the overall performance of service processing by multiple processing cores and reducing power consumption overhead caused by traditional speculation failures.

[0014] In a possible implementation of the first aspect, the second processing core includes a memory access control module and a second cache, and the method further includes: the memory access control module sends second operation information to the second cache, the second operation information including address information accessed by the memory access control module during the second processing core's operation of the second business; the second cache receives and filters duplicate addresses in the second operation information, and sends the second operation information after filtering the duplicate addresses to the first processing core. Optionally, the second cache includes a cache queue, which specifically receives and filters duplicate addresses in the second operation information, and sends the second operation information after filtering the duplicate addresses to the first processing core. In the above possible implementation, the transmission efficiency of the second operation information can be improved by receiving and filtering the second operation information sent by the first processing core through the second cache; in addition, the action of receiving and filtering duplicate addresses in the second operation information can be specifically performed by the cache queue in the second cache, and the cache queue is a multiple-input and one-output queue, which can also achieve the effect of balancing the bandwidth difference between input and output.

[0015] According to a second aspect, a business processing device is provided, which includes a first processing core and a second processing core; the first processing core is used to: obtain first operation information, which is obtained by the first processing core running the first business; receive second operation information from the second processing core, which is obtained by the second processing core running the second business; when it is determined that the first business and the second business are at least partially identical according to the first operation information and the second operation information, run the first business according to the second operation information; the second processing core is used to: provide the second operation information to the first processing core.

[0016] In a possible implementation manner of the second aspect, the first operation information and the second operation information include at least one of the following: data address information, instruction address information, branch address information, jump information, or value prediction information.

[0017] In a possible implementation of the second aspect, the first operation information includes multiple first address information, the second operation information includes multiple second address information, and the first processing core is further used to: when the number of identical address information in the multiple first address information and the multiple second address information is greater than a preset threshold, determine that the first business and the second business are at least partially identical; or, when the number of identical address offset values ​​in the address offset values ​​of the multiple second address information relative to the multiple first address information is greater than a preset threshold, the first processing core determines that the first business and the second business are at least partially identical, wherein the address offset values ​​of the multiple second address information relative to the multiple first address information include the address offset value between each second address information and each first address information.

[0018] In one possible implementation of the second aspect, the first processing core is further configured to receive security information from the second processing core and verify the security of the second processing core based on the security information. Optionally, the security information includes at least one of the following: an address space identifier (ASID), a virtual machine identifier (VMID), and privilege level information.

[0019] In a possible implementation manner of the second aspect, the apparatus further includes: a bus, configured to transmit the second operating information from the second processing core to the first processing core.

[0020] In a possible implementation of the second aspect, the operating frequency of the first processing core is different from the operating frequency of the second processing core; the first processing core includes: an asynchronous interface, used to receive second operating information from the second processing core through a bus, and perform asynchronous logical processing on the second operating information.

[0021] In a possible implementation of the second aspect, the first processing core includes a first cache; and the first processing core is further configured to write the second running information into the first cache as prefetch information of the first service.

[0022] In a possible implementation of the second aspect, the second processing core includes a memory access control module and a second cache; the memory access control module is used to send second operation information to the second cache, and the second operation information includes address information accessed by the memory access control module during the second processing core running the second business; the second cache is used to filter duplicate addresses in the second operation information.

[0023] In a third aspect, an electronic device is provided, which includes a memory and at least one processor, the memory being used to store computer instructions, and the at least one processor including multiple processing cores for executing the computer instructions so that the electronic device implements the business processing method provided in the first aspect or any possible implementation of the first aspect.

[0024] In another aspect of the present application, a computer-readable storage medium is provided, in which a computer program or instruction is stored. When the computer program or instruction is executed, the business processing method provided by the first aspect or any possible implementation of the first aspect is implemented.

[0025] In another aspect of the present application, a computer program product is provided, which includes: a computer program (also referred to as code, or instructions), which, when executed, enables a computer to execute a business processing method as provided in the first aspect or any possible implementation of the first aspect.

[0026] It can be understood that the beneficial effects that can be achieved by any of the business processing devices, electronic devices, computer-readable storage media and computer program products provided above can correspond to the beneficial effects of the business processing methods provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] FIG1 is a performance diagram of a single-core system and a multi-core system when speculative means are applied, provided in an embodiment of the present application;

[0028] FIG2 is a schematic diagram of a method of processing homogeneous services by multiple processing cores according to an embodiment of the present application;

[0029] FIG3 is a schematic diagram of the structure of a multi-core system provided in an embodiment of the present application;

[0030] FIG4 is a flow chart of a service processing method for multiple processing cores provided in an embodiment of the present application;

[0031] FIG5 is a schematic diagram of a plurality of processing cores processing services according to an embodiment of the present application;

[0032] FIG6 is a schematic diagram of a service learning mode provided in an embodiment of the present application;

[0033] FIG7 is a schematic diagram of a data structure provided in an embodiment of the present application;

[0034] FIG8 is a schematic diagram of another data structure provided in an embodiment of the present application;

[0035] FIG9 is a schematic diagram of a service processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] The following will discuss in detail the making and use of various embodiments. However, it should be understood that many applicable inventive concepts provided herein can be implemented in a variety of specific contexts. The specific embodiments discussed are merely illustrative of specific ways to implement and use the present application and technology and do not limit the scope of this application.

[0037] Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art.

[0038] Various circuits or other components may be described or referred to as being "configured to" perform one or more tasks. In this case, "configured to" is used to imply structure by indicating that the circuit / component includes structure (e.g., circuitry) that performs the one or more tasks during operation. Thus, even when a specified circuit / component is not currently operational (e.g., not turned on), the circuit / component may be referred to as being configured to perform the task. Circuits / components used with the phrase "configured to" include hardware, such as circuitry that performs an operation, etc.

[0039] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. In the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can represent: a, b, c, a and b, a and c, b and c, a, b and c; where a, b and c can be single or multiple.

[0040] The embodiments of this application use terms such as "first" and "second" to distinguish objects with similar names, functions, or effects. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or order of execution. The term "coupled" is used to indicate an electrical connection, including direct connection via wires or connectors or indirect connection via other devices. Therefore, "coupling" should be considered a broadly defined electronic communication connection.

[0041] It should be noted that, in this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0042] During the business processing, the cache resources and bandwidth resources of the multi-core system often become bottlenecks at peak performance. In single-core systems, speculative means such as front-end prefetching and back-end prefetching are usually used to improve the performance of business processing. However, this speculative means will generate additional bandwidth overhead and cache overhead. For example, a front-end prefetch error will cause the wrong path to fetch invalid instructions and generate invalid memory access. The back-end prefetch will generate a large number of memory access requests with guessed addresses, thereby polluting the cache and causing a bandwidth bottleneck. The speculative means adopted by the above-mentioned single-core system are not effective when applied to a multi-core system. In order to solve this problem, the following embodiments provide corresponding technical means.

[0043] When speculative means are used in single-core and multi-core systems, the corresponding performance is usually evaluated using the following indicators: coverage, accuracy, and timeliness. Coverage: refers to the proportion of all data used by the processing core (or core, also called processor core) that is prefetched into the cache, where the higher the coverage, the better. Accuracy: refers to the proportion of prefetched data accessed by the processing core to all prefetched data, where the higher the accuracy, the better. If the accuracy is too low, useless prefetched data will pollute the cache and occupy memory bandwidth. Timeliness: refers to the timing when the data to be accessed is prefetched into the cache, where if the data is prefetched into the cache just when it is accessed, the timeliness is good. If the data has not been prefetched into the cache when it is accessed or is prefetched into the cache too early, the timeliness is poor. Figure 1 shows a schematic diagram of the above three indicators when speculative means are applied in single-core and multi-core systems. Among them, when speculative means are used in multi-core systems, the above-mentioned accuracy and coverage often cannot be met simultaneously. Generally, the only way to ensure that the accuracy meets the multi-core requirements is to reduce performance benefits.

[0044] Based on this, an embodiment of the present application provides a business processing method for multiple processing cores. When the businesses processed by multiple processing cores are at least partially the same (or referred to as homogeneous businesses or homogeneous tasks), the method can be based on the operation information (or referred to as speculation auxiliary information) obtained by running the business of a certain processing core (for example, a core with higher performance) shared between the cores, and run the business of other cores based on the shared information, so that the performance of other cores is greatly improved, thereby improving the accuracy of speculation and thus improving the performance of business processing. For example, as shown in Figure 2, the multiple arrows in Figure 2 represent multiple processing cores, and the length of each arrow represents the operation efficiency of the processing core; wherein, in the first stage, the multiple processing cores start to run homogeneous businesses and the corresponding operation efficiencies are close, in the second stage, the operation efficiency of the first processing core is faster, at this time, the first processing core shares the operation information with the other processing cores, and in the third stage, after the other cores run the business according to the shared operation information, the operation efficiency finally achieved is close to the operation efficiency of the first processing core, thereby improving the overall performance of the system. The above-mentioned businesses are at least partially the same or homogeneous, including that the operations performed by different tasks or the data used are at least partially the same. For example, there is at least a portion of the data address information, instruction address information, branch address information, jump information, or value prediction information involved in the tasks that are identical.

[0045] The service processing method provided in the embodiment of the present application can be applied to a multi-core system having multiple processing cores. The structure of the multi-core system is introduced and described below.

[0046] FIG3 is a schematic diagram of the structure of a multi-core system provided in an embodiment of the present application. The multi-core system includes multiple processing cores, which can be used to deploy homogeneous services. For example, the multiple processing cores can be applied to servers or to devices that perform high-performance computing (HPC). As shown in FIG3 , the multiple processing cores can be coupled into a variety of different structures, including but not limited to: a ring structure, a mesh structure, a cross structure, or a combination of the above structures. A to P in the figure represent different processing cores. In one example, the multiple processing cores are coupled into a ring structure via a ring bus, for example, processing core A to processing core G are coupled into a ring structure. In another example, the multiple processing cores are coupled into a mesh structure via a mesh bus, for example, processing core A to processing core P are coupled into a mesh structure. In yet another example, the multiple processing cores are coupled into a crossbar structure via a crossbar bus, for example, processing core A to processing core E are coupled into a crossbar structure.

[0047] In this multi-core system, any one of the multiple processing cores can be coupled to a connected processing core directly, via a bus, or via a router, etc., and this embodiment of the present application does not specifically limit this. The above example only uses the multiple processing cores coupled via a bus as an example, and the above example does not limit the embodiments of the present application.

[0048] Furthermore, the multiple processing cores may include multiple different processing cores of the same processor, or the multiple processing cores may include processing cores of multiple different processors. Optionally, the structures and hardware resource sizes of any two processing cores in the multiple processing cores may be identical, for example, the cache sizes of the multiple processing cores may be identical.

[0049] Furthermore, the multi-core system may also include a memory coupled to the multiple processing cores, which may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. Exemplarily, the RAM may be a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), and a double data rate synchronous dynamic random access memory (DDR SDRAM). The memory is not shown in the figure.

[0050] The multi-core system can be an electronic device, or a system on chip (SoC) or a chipset comprising multiple chips applied to an electronic device, or a module comprising the SoC or chipset. The electronic device can be used as a terminal device or a server. Optionally, the electronic device includes but is not limited to: mobile phones, tablet computers, laptops, desktop computers, PDAs, ultra-mobile personal computers (umPCs), mobile internet devices (MIDs), netbooks, camcorders, cameras, wearable devices (such as smart watches and smart bracelets, etc.), vehicle-mounted equipment (such as cars, bicycles, electric vehicles, airplanes, ships, trains, high-speed railways, etc.), virtual reality (VR) equipment, augmented reality (AR) equipment, wireless terminals in industrial control, smart home devices (such as refrigerators, televisions, air conditioners, electric meters, etc.), intelligent robots, workshop equipment, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, or wireless terminals in smart homes, flying equipment (such as intelligent robots, hot air balloons, drones, airplanes), etc.

[0051] Figure 4 is a flow chart of a business processing method for multiple processing cores provided in an embodiment of the present application. The method can be applied to the multi-core system provided above. The multiple processing cores include a first processing core and a second processing core. The method includes the following steps.

[0052] S201a: The first processing core obtains first operation information, where the first operation information is obtained by the first processing core running a first service.

[0053] S201b: The second processing core obtains second operation information, where the second operation information is obtained by the second processing core running the second service.

[0054] The second processing core may be a processing core with a faster operating speed among the multiple processing cores, and the second processing core may be referred to as a leader core or a master core. The first processing core may be any processing core among the multiple processing cores except the second processing core, and the operating speed of the first processing core is lower than that of the second processing core, and the first processing core may be referred to as a follower core or a slave core.

[0055] In addition, the first business and the second business can be exactly the same business, or partially the same business (or referred to as close or similar business), and the partially identical business can refer to the data and / or instructions that the two tasks need to access during operation are partially the same. For example, the first business and the second business can be two parallel computing businesses in the same high-performance computing. There is a large degree of repetitiveness and similarity in the operation of the above-mentioned first business and the second business, for example, there is repetitiveness and similarity in the instructions, data or branches that are run.

[0056] Optionally, the second processing core can be pre-set or designated as the leader core. Exemplarily, the leader core is determined based on the physical location of the multiple processing cores, such as setting the first core at the physical location as the leader core; or, the leader core is the processing core corresponding to the logical core determined by the basic input and output system (BIOS) after the multi-core system is powered on; or, the leader core is designated by software running on the multi-core system. The embodiments of the present application do not specifically limit the method of designating the leader core.

[0057] In one possible embodiment, when a first processing core receives an execution instruction for a first service, the first processing core executes the first service and obtains first execution information based on the results of the executed portion of the first service. Similarly, when a second processing core receives an execution instruction for a second service, the second processing core executes the second service and obtains second execution information based on the results of the executed portion of the second service. The second processing core executes at a higher speed than the first processing core, or the second processing core's execution progress is higher than the first processing core's. Sharing the second execution information obtained by the second processing core with the first processing core can improve the speculation performance of the first processing core, thereby improving service processing performance.

[0058] Optionally, the first operation information and the second operation information include at least one of the following items obtained during the execution of the business: data address information, instruction address information, branch address information, jump information, or value prediction information. The data address information (or memory access information) is the address of the data actually accessed in the memory. The instruction address information is the address of the instruction actually accessed in the memory. The branch address information is the address of the branch actually executed in the memory. The jump information may include a jump direction and a jump destination. The jump direction is used to indicate the actual jump direction of the jump instruction (i.e., jump or not jump), and the jump destination refers to the destination address after the actual jump. The value prediction information refers to a numerical value obtained by prediction when the memory access is not completed, and the numerical value may be the result of an intermediate execution. The memory may be located on the same or different chip as the one or more processing cores mentioned above, and may be a volatile or non-volatile memory, which is not limited in this embodiment.

[0059] Exemplarily, the first operation information includes data address information obtained by the first processing core during the operation of the first business, and the second operation information includes data address information obtained by the second processing core during the operation of the second business; or, the first operation information includes instruction address information obtained by the first processing core during the operation of the first business, and the second operation information includes instruction address information obtained by the second processing core during the operation of the second business; or, the first operation information includes branch address information obtained by the first processing core during the operation of the first business, and the second operation information includes branch address information obtained by the second processing core during the operation of the second business.

[0060] S202a: The second processing core sends second operation information to the first processing core.

[0061] In one possible embodiment, during the operation of the second service, the second processing core may send second operation information obtained by running the second service to the first processing core once or multiple times. Exemplarily, the second processing core periodically or aperiodically sends the second operation information obtained by running the second service to the first processing core. For example, during the operation of the second service, the second processing core periodically sends the second operation information obtained by running the second service to the first processing core each cycle; or, the second processing core sends the second operation information obtained by running the second service to the first processing core when the storage status of the local cache meets certain conditions.

[0062] Optionally, as shown in FIG5 , the second processing core includes a memory access control module and a cache (referred to herein as a second cache). The access control module may also be referred to as a load store unit (LSU), and the second cache may be a prefetch input buffer (inbuf). Furthermore, the second processing core may also include a level 1 cache (L1), a hardware prefetcher (HWP), a prefetch to level 1 cache (PFL1) interface targeting the first level cache, a prefetch to level 2 cache (PFL2) interface targeting the second level cache, a prefetch to level 3 cache (PFL3) interface targeting the third level cache, and an asynchronous interface. The figure illustrates an example in which the second processing core includes the LSU, inbuf, L1, PFL1 interface, PFL2 interface, PFL3 interface, and an asynchronous interface, and the LSU and L1 are represented as LSU+L1. Optionally, as shown in FIG5 , other processing cores (slave cores), such as the first processing core, may have a structure similar to that of the second processing core.

[0063] In one possible example, as shown in FIG5 , after the second processing core obtains the second operation information, the memory access control module (e.g., LSU) can send the second operation information to the second cache (e.g., inbuf), where the second operation information includes address information accessed by the memory access control module during the second processing core's execution of the second service. When the second cache receives the second operation information, the second cache can filter duplicate addresses in the second operation information before sending it to the first processing core. For example, the second cache sends the second operation information via the PFL3 interface. Exemplarily, the second cache includes a cache queue, which can specifically receive the second operation information, filter duplicate addresses in the second operation information, and send it via the PFL3 interface. The cache queue can be a multi-input and one-output cache queue, which, in addition to filtering duplicate addresses, can also be used to balance the bandwidth difference between input and output to improve the transmission efficiency of the second operation information.

[0064] For example, when the second operating information sent by the second processing core to the first processing core may include multiple address information (referred to as second address information in this article), the multiple second address information may be memory access request addresses that do not hit the first-level cache, or may be a subset or a full set of first-level high-speed memory access addresses of any rules, etc. The embodiments of the present application do not impose specific restrictions on this.

[0065] Furthermore, the second processing core may also send security information to the first processing core, and the security information may be used to verify the security of the second processing core. The security information may be identification information and / or permission information assigned to the second processing core. Exemplarily, the security information includes at least one of the following: an address space identifier (ASID), a virtual machine identifier (VMID), and privilege level information.

[0066] The security information can be transmitted synchronously with the data packet of the second operational information to reduce the complexity of transmitting the security information and the second operational information. Alternatively, the security information can be transmitted asynchronously via a data packet of a specific format, i.e., the security information and the second operational information are transmitted via different data packets, thereby reducing the length of the data packets through asynchronous transmission. For example, the second processor core can send the security information in advance to allow the first processor core to complete the security check. For example, the second processor core can send the security information when the device is powered on.

[0067] S202b: The first processing core receives second operation information from the second processing core.

[0068] Optionally, the multiple processing cores are coupled via a bus, which may be an existing bus in the multi-core system (i.e., a bus in the multi-core system is reused) or a bus separately provided by the present application (e.g., a newly added bus). This embodiment of the present application does not impose any specific limitations on this. The second processing core may send second operation information to the first processing core via the bus; correspondingly, the first processing core may receive the second operation information from the second processing core via the bus.

[0069] Optionally, the second operation information may be transmitted in a compressed format or may not be transmitted in a compressed format, and this embodiment of the present application does not impose any specific limitation on this.

[0070] In one possible embodiment, the frequency of the first processing core is different from the frequency of the second processing core. Optionally, the frequency of the first processing core is lower than the frequency of the second processing core. In this case, the asynchronous interface of the first processing core can perform asynchronous logic processing during the process of receiving the second operating information, so that the frequency of the second operating information is the same as the frequency of the first processing core.

[0071] Optionally, as shown in FIG5 , the structure of the first processing core is similar to that of the second processing core. The first processing core may include a memory access control module and a first cache, which may be a prefetch input cache (inbuf). Furthermore, the first processing core may also include other modules, such as a level 1 cache (L1), a hardware prefetcher (HWP), a PFL1 interface, a PFL2 interface, a PFL3 interface, and an asynchronous interface. Furthermore, the PFL3 interface of the first processing core and the PFL3 interface of the second processing core, as well as the asynchronous interface of the first processing core and the asynchronous interface of the second processing core, may be coupled via a bus.

[0072] In a possible example, as shown in Figure 5, the first processing core receives the second operation information from the second processing core through the bus, which may include: the asynchronous interface of the first processing core receives the second operation information from the second processing core through the bus, and performs asynchronous logical processing on the second operation information to achieve timing synchronization.

[0073] It can be understood that the multiple processing cores may include multiple slave cores, that is, in addition to the first processing core, the multiple processing cores may also include other slave cores similar to the first processing core, and the second processing core (that is, the leader core) may send second operating information to each slave core. Figure 5 is used as an example to illustrate that the multiple slave cores include two processing cores.

[0074] Furthermore, when the second processing core sends security information to the first processing core, the first processing core can receive the security information and verify the security of the second processing core based on the security information. If the second processing core is determined to be secure based on the security information, the first processing core can execute step S203 below. Alternatively, the security information transmission and security verification can also be performed in advance.

[0075] S203: When it is determined based on the first operation information and the second operation information that the first service and the second service are at least partially identical, the first processing core executes the first service based on the second operation information. Optionally, when the first service and the second service are different, the first processing core does not execute the first service based on the second operation information. In other words, the two processing cores are not executing homogeneous tasks.

[0076] Determining that the first and second services are at least partially identical can also be referred to as determining to enter a cluster learning optimization (CLO) mode or a service sharing learning mode, i.e., the first processing core can operate or process the first service using the second operating information shared by the second processing core, or the first processing core can process the first service by learning how the second processing core operates or processes the second service. Optionally, the homogenization mode can be enabled or disabled by the first processing core, or a software path can be designed to be enabled or disabled by an upper-layer operating system, which is not specifically limited in this embodiment of the present application.

[0077] In one implementation, the first operation information includes a plurality of first address information, and the second operation information includes a plurality of second address information. The plurality of first address information and the plurality of second address information can be data address information, or instruction address information, or branch address information, or include two address information indicating two of the three types of address information. Exemplarily, the first operation information includes a plurality of first data address information, and the second operation information includes a plurality of second data address information; or, the first operation information includes a plurality of first instruction address information, and the second operation information includes a plurality of second instruction address information; or, the first operation information includes a plurality of first branch address information, and the second operation information includes a plurality of second branch address information; the first operation information includes a plurality of first data address information and a plurality of first instruction address information, and the second operation information includes a plurality of second data address information and a plurality of second instruction address information; or, the first operation information includes a plurality of first instruction address information and a plurality of first branch address information, and the second operation information includes a plurality of second instruction address information and a plurality of second branch address information.

[0078] Exemplarily, depending on the different information included in the second operation information, the above-mentioned business sharing learning may include sharing learning of different information. For example, as shown in FIG6 , when the multiple second address information included in the second operation information is data address information, the business sharing learning includes data cache sharing learning; when the multiple second address information included in the second operation information is instruction address information, the business sharing learning includes instruction cache sharing learning; when the multiple second address information included in the second operation information is branch address information, the business sharing learning includes branch prediction sharing learning. FIG6 takes the processing core 0 among the multiple processing cores as the leading core and the other processing cores (for example, processing core 1 and processing core 2, etc.) as the slave cores as an example for explanation, and shows the main pipeline of each processing core.

[0079] In a possible embodiment, determining that the first business and the second business are at least partially identical based on the first operation information and the second operation information includes: when the number of identical address information existing in the multiple first address information and the multiple second address information is greater than a first preset threshold, determining that the first business and the second business are at least partially identical.

[0080] Exemplarily, the multiple second address information included in the second operation information can be arranged according to a certain data structure. When the first processing core determines that a certain first address information among the multiple first address information included in the first operation information is the same as a certain second address information, the first processing core records the same first address information in the data structure, and makes statistics on the same address information, and determines that the first business and the second business are at least partially identical when the number of identical address information is greater than a first preset threshold. For example, the data structure can be a homogeneous training table (CLO train table, CTT), as shown in Figure 5. The CTT can be a naturally winding storage structure that is naturally refreshed by continuous writing. Specifically, as shown in Figure 7, the CTT can include n (n is a positive integer) rows and two columns, each row of the first column includes a second address information, and each row of the second column can be used to store a first address information. When the first processing core finds the first address information that is the same as a certain second address information in the multiple first address information, the first processing core can store the first address information in the corresponding row and color the row. In FIG. 7 , n pieces of second address information are represented as VA1 to VAn, n pieces of first address information are represented as TA1 to TAn, and the colored rows are represented by filling.

[0081] In another possible embodiment, when the number of identical address offset values ​​among the address offset values ​​corresponding to the plurality of first address information and the plurality of second address information is greater than a second preset threshold, the first processing core determines that the first service and the second service are at least partially identical. The address offset values ​​corresponding to the plurality of first address information and the plurality of second address information include the address offset value between the plurality of second address information and each first address information.

[0082] Exemplarily, the multiple second address information included in the second operation information can be arranged according to a homogeneous training table (CTT). As shown in FIG5 , for each first address information in the multiple first address information, the first processing core determines the address offset value between the first address information and each second address information, records each address offset value and the corresponding count (which can be obtained by a counter) through the CTT, and colors the rows where the address offset count exceeds a second preset threshold. Specifically, as shown in FIG8 , the CTT can include n (n is a positive integer) rows and three columns. Each row in the first column includes a second address information, each row in the second column can be used to store an address offset value, and each row in the third column can be used to record the count of the address offset value of the corresponding row. When the address offset count of a row exceeds the second preset threshold, the row is colored, for example, the row with the maximum count is colored, thus learning a fixed offset value. In FIG8 , the n second address information are represented as VA1 to VAn, the n address offset values ​​are represented as OF1 to OFn, the corresponding counts are represented as CT1 to CTn, and the colored rows are represented by padding.

[0083] The above-mentioned first preset threshold and second preset threshold can be set in advance, and the embodiment of the present application does not limit the specific values.

[0084] Furthermore, in one possible embodiment, when the first processing core determines that the first service and the second service are at least partially identical, the first processing core running the first service according to the second operation information may include: the first processing core using the second operation information subsequently received from the second processing core as prefetch information, and continuing to run the first service according to the prefetch information. Exemplarily, when the second operation information includes data address information, the first processing core runs the first service according to the data address information; when the second operation information includes instruction address information, the first processing core runs the first service according to the instruction address information; when the second operation information includes branch address information, the first processing core runs the first service according to the branch address information; when the second operation information includes jump information, the first processing core runs the first service according to the jump information; when the second operation information includes value prediction information, the first processing core runs the first service according to the value prediction information.

[0085] Optionally, before the first processing core runs the first business according to the second running information, it can also write the second running information as pre-fetch information of the first business into the first cache, and obtain instructions and / or data from the memory according to the pre-fetch information, thereby running the first business according to the obtained instructions and / or data.

[0086] Exemplarily, as shown in FIG5 , after the first processing core enters the homogeneous mode according to CTT identification, the second operation information shared by the second processing core is sent to the hardware prefetcher (HWP) main pipeline. After the hardware prefetcher main pipeline is processed by caching and address translation, it is sent to the first-level, second-level, and third-level caches according to the policy settings. That is, the data and / or instructions to be used in the future are filled into the cache in advance, and then the first business is run based on the data and / or instructions in the cache to improve the performance of the first processing core in running the first business.

[0087] In an embodiment of the present application, a first processing core obtains first operating information for running a first service, receives second operating information for running a second service by a second processing core, and when it is determined based on the first operating information and the second operating information that the first service and the second service are partially identical, the first service is run according to the second operating information. Because the second operating information is the operating information actually obtained by the second processing core, and the first service and the second service are partially identical, that is, there is duplication and similarity between the operation of the first service and the operation of the second service, when the first processing core runs the first service based on the second operating information, the accuracy and efficiency of speculation can be greatly improved, thereby improving the overall performance of multiple processing cores in processing services and reducing the power consumption overhead caused by traditional speculation failures. Specifically, when the first processing core enters data cache sharing learning, the first processing core runs the first business according to the second operation information of the actual operation sent by the second processing core, which can greatly reduce the data cache miss rate; when the first processing core enters instruction cache sharing learning, the first processing core runs the first business according to the second operation information of the actual operation sent by the second processing core, which can greatly reduce the instruction cache miss rate; when the first processing core enters branch prediction sharing learning, the first processing core runs the first business according to the second operation information of the actual operation sent by the second processing core, which can greatly reduce the branch prediction failure rate.

[0088] In addition, when the first processing core determines that the first business and the second business are at least partially identical, it runs the first business according to the second operation information; when it determines that the first business and the second business are different, it does not run the first business according to the second operation information. In this way, the first processing core can automatically determine whether to run the first business according to the second operation information (that is, automatically determine whether it belongs to a profit scenario), thereby avoiding false triggering and negative profits in non-target scenarios.

[0089] In addition, when the first processing core determines whether the first business and the second business are at least partially identical, it can directly determine it based on the number of identical address information in the multiple first address information included in the first operation information and the multiple second address information included in the second operation information, or determine it based on the number of identical address offset values ​​existing in the address offset values ​​corresponding to the multiple first address information and the multiple second address information, so as to cover more dynamic address allocation scenarios and thereby improve the coverage of the benefits.

[0090] Furthermore, the above embodiment is described by taking the second processing core's operating speed as higher than the first processing core's operating speed as an example. In fact, this setting is not used to limit the application scenario of the embodiment. In actual applications, multiple processing cores can have the same or different capabilities or processing speeds, and there is no limit on which one has stronger capabilities, as long as the operating information obtained after one of the processing cores executes the task can be used to share with another processing core in order to execute the technical solution mentioned in this embodiment. For example, still taking Figure 5 as an example, the functions of the master core and the slave core can be swapped. The first processing core (slave core) can send its operating information to the second processing core (master core) so that the second processing core can execute a similar embodiment process similar to that described in Figure 4 to achieve similar functions. For example, the first processing core may first execute at least part of the function of the task and obtain operating information due to software scheduling or user selection. This operating information can be shared with the second processing core for continuing to execute similar homogeneous tasks to achieve similar effects. It can be understood that the master core and slave core mentioned in this embodiment, and the differences between the two cores, are only applicable scenarios, but are not used to limit the technical solution.

[0091] Based on this, an embodiment of the present application also provides a business processing device, which can be applied to a multi-core system. As shown in Figure 9, the device includes: a first processing core and a second processing core. In an embodiment of the present application, the first processing core can be used to execute S201a, S202b, S203 in the above-mentioned method embodiment, the step of receiving security information from the second processing core and verifying the security of the second processing core based on the security information, and / or other steps described herein; the second processing core can be used to execute step S201b in the above-mentioned method embodiment, the step of sending security information to the first processing core, and / or other steps described herein.

[0092] It can be understood that the specific structure of the first processing core and the second processing core, as well as all relevant contents of the steps involved in the above method embodiment can be referred to the embodiment of the business processing device, and the embodiment of this application will not be repeated here.

[0093] In the embodiment of the present application, since the second operating information is the operating information obtained by the actual operation of the second processing core, and the first business and the second business are partially the same, that is, the operation of the first business and the operation of the second business are repetitive and similar, when the first processing core runs the first business according to the second operating information, it can greatly improve the accuracy and efficiency of speculation, thereby improving the overall performance of multiple processing cores in processing businesses and reducing the power consumption overhead caused by traditional speculation failures.

[0094] In another aspect of the present application, an electronic device is provided, comprising a memory and at least one processor, the memory being configured to store computer instructions, the at least one processor comprising multiple processing cores configured to execute the computer instructions, thereby enabling the electronic device to implement any of the service processing methods for multiple processing cores provided above. Optionally, the at least one processor comprises the service processing apparatus provided above.

[0095] It can be understood that all relevant contents of each step involved in the above method embodiment can be referred to the embodiment of the business processing device and the embodiment of the electronic device, and the embodiments of the present application will not be repeated here.

[0096] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical functional division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another device, or ignoring or not implementing certain features.

[0097] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0098] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. The readable storage medium may include: a USB flash drive, a mobile hard drive, a read-only memory, a random access memory, a magnetic disk, or an optical disk, etc., which can store program code. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product.

[0099] In another embodiment of the present application, a readable storage medium is also provided, which stores computer execution instructions. When a device (which may be a single-chip microcomputer, chip, etc.) or a processor executes the steps in the above method embodiment.

[0100] In another embodiment of the present application, a computer program product is provided, which includes computer instructions stored in a readable storage medium; at least one processor of the device can read the computer instructions from the readable storage medium, and at least one processor executes the computer instructions so that the device performs the steps in the above method embodiment.

[0101] Finally, it should be noted that the above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A service processing method for multiple processing cores, characterized in that: The method comprises: The first processing core acquires first operation information, where the first operation information is obtained by the first processing core running a first service; The first processing core receives second operation information from the second processing core, where the second operation information is obtained by the second processing core running a second service; When it is determined according to the first operation information and the second operation information that the first service and the second service are at least partially identical, the first processing core operates the first service according to the second operation information.

2. The method according to claim 1, characterized in that The first operation information and the second operation information include at least one of the following: data address information, instruction address information, branch address information, jump information, or value prediction information.

3. The method according to claim 1 or 2, characterized in that: The first operation information includes a plurality of first address information, and the second operation information includes a plurality of second address information; The determining, according to the first operation information and the second operation information, that the first service and the second service are at least partially identical includes: When the number of identical address information in the plurality of first address information and the plurality of second address information is greater than a preset threshold, the first processing core determines that the first service and the second service are at least partially identical; or, When the number of identical address offset values ​​existing in the address offset values ​​of the multiple second address information relative to the multiple first address information is greater than a preset threshold, the first processing core determines that the first business and the second business are at least partially identical, wherein the address offset values ​​of the multiple second address information relative to the multiple first address information include: the address offset value of each second address information relative to each first address information.

4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: The first processing core receives security information from the second processing core, and verifies security of the second processing core according to the security information.

5. The method according to claim 4, characterized in that The security information includes at least one of the following: an address space identifier ASID, a virtual machine identifier VMID or privilege level information.

6. The method according to any one of claims 1 to 5, characterized in that: The first processing core receives second operation information from the second processing core, including: The first processing core receives the second operation information from the second processing core through a bus.

7. The method according to claim 6, characterized in that The operating frequency of the first processing core is different from the operating frequency of the second processing core, and the first processing core receives first operating information from the second processing core through a bus, including: The asynchronous interface of the first processing core receives the second operation information from the second processing core through a bus, and performs asynchronous logic processing on the second operation information.

8. The method according to any one of claims 1 to 7, characterized in that: The first processing core includes a first cache, and before the first processing core runs the first service according to the second running information, the method further includes: The first processing core writes the second running information into the first cache as pre-fetch information of the first service.

9. The method according to any one of claims 1 to 8, characterized in that: The second processing core includes a memory access control module and a second cache, and the method further includes: The memory access control module sends the second operation information to the second cache, where the second operation information includes address information accessed by the memory access control module during the process in which the second processing core runs the second service; The second cache filters duplicate addresses in the second operation information, and sends the second operation information after the duplicate addresses are filtered out to the first processing core.

10. A service processing device, characterized in that: The device comprises: The first processing core is used to: Acquire first operation information, where the first operation information is obtained by the first processing core running a first service; receiving second operation information from the second processing core, where the second operation information is obtained by the second processing core running a second service; When it is determined according to the first operation information and the second operation information that the first service and the second service are at least partially identical, operating the first service according to the second operation information; and The second processing core is used to provide the second operation information to the first processing core.

11. The device according to claim 10, characterized in that The first operation information and the second operation information include at least one of the following: data address information, instruction address information, branch address information, jump information, or value prediction information.

12. The device according to claim 10 or 11, characterized in that The first operation information includes a plurality of first address information, and the second operation information includes a plurality of second address information; The first processing core is also used for: When the number of identical address information in the plurality of first address information and the plurality of second address information is greater than a preset threshold, it is determined that the first service and the second service are at least partially identical; or, When the number of identical address offset values ​​existing in the address offset values ​​of the multiple second address information relative to the multiple first address information is greater than a preset threshold, it is determined that the first business and the second business are at least partially identical, wherein the address offset values ​​of the multiple second address information relative to the multiple first address information include: the address offset value of each second address information relative to each first address information.

13. The device according to any one of claims 10 to 12, characterized in that: The first processing core is also used for: The security information from the second processing core is received, and the security of the second processing core is verified according to the security information.

14. The device according to claim 13, characterized in that The security information includes at least one of the following: an address space identifier ASID, a virtual machine identifier VMID or privilege level information.

15. The device according to any one of claims 10 to 14, characterized in that: The device also includes: a bus, configured to transmit the second operation information from the second processing core to the first processing core.

16. The device according to claim 15, characterized in that The operating frequency of the first processing core is different from the operating frequency of the second processing core; The first processing core includes: an asynchronous interface, used for receiving the second operation information from the second processing core through the bus, and performing asynchronous logic processing on the second operation information.

17. The device according to any one of claims 10 to 16, characterized in that: The first processing core includes a first cache; The first processing core is further configured to write the second operation information into the first cache as pre-fetch information of the first service.

18. The device according to any one of claims 10 to 17, characterized in that: The second processing core comprises: a memory access control module, configured to send the second operation information to the second cache, wherein the second operation information includes address information accessed by the memory access control module during the process in which the second processing core runs the second service; The second cache is used to filter duplicate addresses in the second operation information, and send the second operation information after filtering the duplicate addresses to the first processing core.

19. An electronic device, characterized in that: The electronic device includes a memory and at least one processor, the memory is used to store computer instructions, and the at least one processor includes multiple processing cores for executing the computer instructions so that the electronic device implements the business processing method for multiple processing cores as described in any one of claims 1-9.

20. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed on a device, the device executes the service processing method for multiple processing cores according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Business processing method, device and equipment for multiple processing cores

    CN120196427A

  • Task scheduling method in heterogeneous multi-core architecture

    CN104899089A

  • Scheduling method and device of processor, electronic equipment and storage medium

    CN115712337A

  • Multi-core system and dynamic module loading method thereof, medium and processor chip

    CN117234607A

  • Multi-core data interaction circuit

    CN214042316U