Multi-thread service processing device, method and equipment

By using cache and controller in a multi-threaded system to determine the similarity between services, the problem of cache and bandwidth resource bottlenecks in a multi-threaded system is solved, the accuracy and efficiency of speculation are improved, and the overall business processing performance is improved.

CN120196429APending Publication Date: 2025-06-24HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311796294.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In multi-threaded systems, cache resources and bandwidth resources often become bottlenecks in peak performance, and speculation methods in single-threaded systems are not effective in multi-threaded systems.

Method used

By introducing a cache and controller in a multi-threaded system, the operation information of each thread is cached and the similarity between services is determined based on this information, thereby improving the accuracy and efficiency of speculation.

Benefits of technology

It improves the accuracy and efficiency of speculation in multi-threaded systems, improves overall business processing performance, and reduces the power consumption and overhead caused by traditional speculation failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196429A_ABST
    Figure CN120196429A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-thread business processing device, method and equipment, relates to the technical field of electronics, and is used for improving speculation accuracy in a multi-thread system so as to improve business processing performance. The device comprises a buffer used for caching first operation information of a first service operated on a first thread and second operation information of a second service operated on a second thread; and the controller is used for acquiring the first operation information and the second operation information, and determining the first operation information as prefetch information of the second service when the first service and the second service are determined to be at least partially identical according to the first operation information and the second operation information. The first operation information is the operation information obtained by actual operation, and the first service and the second service are partially the same, so that the first operation information is determined as the prefetched information of the second service, and the speculation accuracy and speculation efficiency can be greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic technologies, and in particular, to a multi-threaded service processing apparatus, method, and device. Background Art

[0002] In the process of service processing, the cache resources and bandwidth resources of a multi-threaded system often become bottlenecks at peak performance. Currently, in a single-threaded system, speculative means such as front-end prefetching and back-end prefetching are usually adopted to improve the performance of service processing. However, this speculative means will generate additional bandwidth overhead and cache overhead. For example, front-end prefetching errors will cause the wrong path to fetch invalid instructions and generate invalid memory accesses, and back-end prefetching will generate a large number of memory access requests for speculative addresses, thus polluting the cache and causing bandwidth bottlenecks. Therefore, the speculative means adopted by the above single-threaded system has poor effects when applied to a multi-threaded system. Summary of the Invention

[0003] This application provides a multi-threaded service processing apparatus, method, and device, which are used to improve the accuracy of speculation in a multi-threaded system, and further improve the performance of service processing.

[0004] To achieve the above object, the embodiments of this application adopt the following technical solutions:

[0005] In a first aspect, a multi-threaded service processing apparatus is provided. The apparatus includes: a buffer for buffering first running information of a first service running on a first thread and second running information of a second service running on a second thread, where the first thread is a leading thread or a main thread, the second thread is a following thread or a slave thread, and the running speed of the first thread is greater than the running speed of the second thread; a controller for obtaining the first running information and the second running information, and when it is determined according to the first running information and the second running information that there is at least partial identity between the first service and the second service, determining the first running information as prefetch information of the second service.

[0006] In the above technical solution, since the first running information is the running information actually obtained by the first thread, and there is at least partial identity between the first service and the second service, that is, there is repetition and similarity between the running of the first service and the running of the second service. In this way, when determining the prefetch information of the second service according to the first running information, when the second thread continues to run the second service according to the prefetch information, the accuracy and efficiency of speculation can be greatly improved, thereby improving the overall performance of multi-threaded service processing and reducing the power consumption overhead caused by traditional speculative failures.

[0007] In a possible implementation of the first aspect, the first running information and the second running information include at least one of the following: data address information, instruction address information, branch address information, jump information, or value prediction information. In the above possible implementation, when the first running information includes data address information, determining the data address information as the prefetch information for the second service can greatly reduce the data cache miss rate; when the first running information includes instruction address information, determining the instruction address information as the prefetch information for the second service can greatly reduce the instruction cache miss rate; when the first running information includes branch address information, determining the branch address information as the prefetch information for the second service can greatly reduce the branch prediction failure rate; when the first running information includes jump information or value prediction information, determining the jump information or value prediction information as the prefetch information for the second service can greatly improve the accuracy and efficiency of speculation, thereby improving the overall performance of the service.

[0008] In a possible implementation of the first aspect, the first running information includes multiple first address information, and the second running information includes multiple second address information; the controller is further configured to: when the number of identical address information among the multiple first address information and the multiple second address information is greater than a preset threshold, determine that the first service and the second service have at least partial identity. In the above possible implementation, by determining whether the first service and the second service have at least partial identity according to the number of identical address information included in the first running information and the second running information, the accuracy of determination can be improved and the complexity of determination can be reduced.

[0009] In a possible implementation of the first aspect, the controller is further configured to: obtain the first security information of the first thread and the second security information of the second thread, and determine the security of the communication between the first thread and the second thread according to the first security information and the second security information. Optionally, the first security information and / or the second security information includes at least one of the following: address space identifier ASID, virtual machine identifier VMID, or privilege level information. In the above possible implementation, the controller can improve the security of speculation by performing a security check on the communication between the first thread and the second thread.

[0010] In a possible implementation of the first aspect, the first running information includes multiple first address information, and the controller is further configured to: convert the multiple first address information into multiple third address information corresponding to the second thread; write the storage information in the cache space corresponding to the multiple first address information into the cache space corresponding to the multiple third address information. In the above possible implementation, the accuracy and efficiency of speculation can be greatly improved, thereby improving the overall performance of the device in processing services and reducing the power consumption overhead caused by traditional speculation failures.

[0011] In a possible implementation of the first aspect, the apparatus further includes: a processing unit, configured to: run a first service through a first thread to obtain first running information, and write the first running information into the buffer; run a second service through a second thread to obtain second running information, and write the second running information into the buffer. In the above possible implementation, by caching the running information of the first thread and the second thread in the buffer, when the controller determines that there is at least partial identity between the first service and the second service, the first running information can be determined as the prefetch information of the second service. In this way, the second thread can continue to run the second service according to the prefetch information, which can greatly improve the accuracy and efficiency of speculation, thereby improving the overall performance of multi-threaded service processing and reducing the power consumption overhead caused by traditional speculation failures.

[0012] In a possible implementation of the first aspect, the buffer includes a first buffer queue and a second buffer queue; the first buffer queue is configured to cache the first running information and filter duplicate addresses in the first running information; the second buffer queue is configured to cache the second running information and filter duplicate addresses in the second running information. In the above possible implementation, the first buffer queue and the second buffer queue respectively cache and filter the first running information and the second running information, which can improve the transmission efficiency of the first running information and the second running information; in addition, the first buffer queue and the second buffer queue can be multi-in and one-out queues, and can also achieve the effect of balancing the bandwidth difference between input and output.

[0013] In a second aspect, a multi-threaded service processing method is provided, which is applied to a service processing apparatus including a buffer and a controller. The method includes: the buffer caches the first running information of the first service running on the first thread and the second running information of the second service running on the second thread; when the controller determines that there is at least partial identity between the first service and the second service according to the first running information and the second running information, the controller determines the first running information as the prefetch information of the second service.

[0014] In a possible implementation of the second aspect, the first running information and the second running information include at least one of the following: data address information, instruction address information, branch address information, jump information, or value prediction information.

[0015] In a possible implementation of the second aspect, the first running information includes a plurality of first address information, and the second running information includes a plurality of second address information; the controller determines that there is at least partial identity between the first service and the second service according to the first running information and the second running information, including: when the number of identical address information existing in the plurality of first address information and the plurality of second address information is greater than a preset threshold, the controller determines that there is at least partial identity between the first service and the second service.

[0016] In a possible implementation of the second aspect, the method further includes: the controller obtains first security information of a first thread and second security information of a second thread, and determines the security of communication between the first thread and the second thread according to the first security information and the second security information.

[0017] In a possible implementation of the second aspect, the first security information and / or the second security information includes at least one of the following: an address space identifier ASID, a virtual machine identifier VMID, or privilege level information.

[0018] In a possible implementation of the second aspect, the first running information includes a plurality of first address information, and the controller determines the first running information as prefetch information for a second service, including: converting the plurality of first address information into a plurality of third address information corresponding to the second thread; writing the storage information in the cache space corresponding to the plurality of first address information into the cache space corresponding to the plurality of third address information.

[0019] In a possible implementation of the second aspect, the apparatus further includes a processing unit, and the method further includes: the processing unit runs a first service through the first thread to obtain first running information, and writes the first running information into the buffer; the processing unit runs a second service through the second thread to obtain second running information, and writes the second running information into the buffer.

[0020] In a possible implementation of the second aspect, the buffer includes a first buffer queue and a second buffer queue, and the buffer caches first running information of a first service running on the first thread and second running information of a second service running on the second thread, including: the first buffer queue caches the first running information and filters duplicate addresses in the first running information; the second buffer queue caches the second running information and filters duplicate addresses in the second running information.

[0021] In a third aspect, an electronic device is provided, which includes a memory and a multi-threaded processor. The memory is used to store computer instructions, and the multi-threaded is used to execute the computer instructions so that the electronic device executes the service processing method provided by the second aspect or any possible implementation of the second aspect.

[0022] In another aspect of the present application, a computer-readable storage medium is provided, in which a computer program or instruction is stored. When the computer program or instruction is run, the service processing method provided by the second aspect or any possible implementation of the second aspect is implemented.

[0023] In another aspect of the present application, a computer program product is provided, which includes: a computer program (which can also be referred to as code or instructions), and when the computer program is run, it causes the computer to execute the service processing method provided by the second aspect or any possible implementation manner of the second aspect.

[0024] It can be understood that for any of the multi-threaded service processing methods, electronic devices, computer-readable storage media, and computer program products provided above, the beneficial effects that can be achieved can be correspondingly referred to the beneficial effects in the multi-threaded service processing device provided above, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A schematic diagram of a method for processing homogeneous services by multiple threads provided by an embodiment of the present application;

[0026] Figure 2 A schematic diagram of the structure of a multi-threaded system provided by an embodiment of the present application;

[0027] Figure 3 A schematic diagram of the structure of a multi-threaded service processing device provided by an embodiment of the present application;

[0028] Figure 4 A schematic diagram of a service processing provided by an embodiment of the present application;

[0029] Figure 5 A schematic diagram of a data structure provided by an embodiment of the present application;

[0030] Figure 6 A schematic diagram of determining prefetch information of a second service provided by an embodiment of the present application;

[0031] Figure 7 A schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The production and use of each embodiment will be discussed in detail below. However, it should be understood that many applicable inventive concepts provided by the present application can be implemented in a variety of specific environments. The specific embodiments discussed merely illustrate the specific ways of implementing and using the present application and the present technology, and do not limit the scope of the present application.

[0033] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by those of ordinary skill in the art.

[0034] Each circuit or other component may be described as or referred to as "configured to" perform one or more tasks. In such cases, "configured to" is used to imply structure by indicating that the circuit / component includes the structure (e.g., circuitry) that performs one or more tasks during operation. Thus, even when the specified circuit / component is not currently operational (e.g., not turned on), the circuit / component can still be referred to as being configured to perform the task. Circuits / components used in conjunction with the phrase "configured to" include hardware, such as circuitry that performs the operation, etc.

[0035] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. In the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single item(s) or plural item(s). For example, at least one of a, b, or c may represent: a, b, c, a and b, a and c, b and c, a, b, and c; where a, b, and c may be single or multiple.

[0036] The embodiments of the present application use terms such as "first" and "second" to distinguish objects with similar names, functions, or roles. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order. The term "coupled" is used to indicate an electrical connection, including directly connected through wires or connection terminals or indirectly connected through other devices. Therefore, "coupled" should be regarded as a broad sense of electronic communication connection.

[0037] It should be noted that in the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0038] During the processing of a service, the cache resources and bandwidth resources of a multi-threaded system often become bottlenecks at peak performance. In a single-threaded system, speculative means such as front-end prefetching and back-end prefetching are usually adopted to improve the performance of service processing. However, this speculative means will generate additional bandwidth overhead and cache overhead. For example, front-end prefetching errors will cause invalid instructions to be fetched along the wrong path and generate invalid memory accesses, and back-end prefetching will generate a large number of memory access requests for speculative addresses, thus polluting the cache and causing bandwidth bottlenecks. The speculative means adopted by the above single-threaded system has poor effects when applied to a multi-core system. To solve this problem, the following embodiments provide corresponding technical means.

[0039] When speculative means are adopted in a single-threaded system and a multi-threaded system, the corresponding performance is usually evaluated by the following three indicators: coverage, accuracy, and timeliness. Coverage: It refers to the proportion of all the data used by a processing core (or called a kernel, also called a processor core) that is prefetched into the cache. The higher the coverage, the better. Accuracy: It refers to the proportion of the prefetched data accessed by the processing core in all the prefetched data. The higher the accuracy, the better. If the accuracy is too low, the useless prefetched data will pollute the cache and occupy the memory bandwidth. Timeliness: It refers to the timing when the data to be accessed is prefetched into the cache. If the data is just prefetched into the cache when it is accessed, the timeliness is good. If the data has not been prefetched into the cache or has been prefetched into the cache too early when it is accessed, the timeliness is poor. Among them, when speculative means are adopted in a multi-threaded system, the above accuracy and coverage often cannot be satisfied at the same time, and generally, only by reducing the performance gain can the accuracy meet the multi-core requirements.

[0040] Based on this, an embodiment of the present application provides a multi-threaded service processing device. In this device, when there is at least partial identity (or called homogeneous services or homogeneous tasks) in the services processed by multiple threads, the running information (or called speculative auxiliary information) obtained by running a service on one of the threads (for example, the thread with higher performance) is determined as the prefetch information for the service running on another thread. In this way, the other thread can run the service according to this prefetch information, thereby greatly improving the performance of the other thread, thereby improving the accuracy of speculation, and further improving the performance of service processing. Exemplarily, as Figure 1 shown Figure 1Multiple arrows in [Figure 0] represent multiple threads, and the length of each arrow represents the running efficiency of that thread. Among them, in the first stage, multiple threads start to run homogeneous services and their corresponding running efficiencies are close. In the second stage, the running efficiency of the first thread is relatively fast. At this time, the running information of the first thread is determined as the prefetch information for the services running on other threads. After other threads run the services according to this prefetch information in the third stage, the finally achieved running efficiency is close to that of the first thread, thereby improving the overall performance of the system. At least part of the services involved above are the same or homogeneous, including at least part of the same data used in different task execution operations. For example, at least part of the same exists in the data address information, instruction address information, branch address information, jump information, or value prediction information involved in the tasks.

[0041] The multi-threaded service processing device provided in the embodiments of the present application can be applied to a multi-threaded system. The structure of this multi-threaded system will be introduced and described below.

[0042] Figure 2 It is a schematic structural diagram of a multi-threaded system provided in the embodiments of the present application. This multi-threaded system can be used to deploy homogeneous services. For example, this multi-threaded system can be applied to a server or a device for performing high performance computing (HPC). Among them, multiple threads are set in this multi-threaded system. These multiple threads can be synchronized multi-thread (SMT), and the number of these multiple threads can be 2, 4, 6, or 8, etc. Exemplarily, as Figure 2 shown in (a) of [Figure 9], this multi-threaded system is an SMT2 system, that is, it has two threads and is represented as thread 0 and thread 1; or, as Figure 2 shown in (b) of [Figure 10], this multi-threaded system can be an SMT4 system, that is, it has four threads and is represented as thread 0 to thread 3; or, as Figure 2 shown in (c) of [Figure 12], this multi-threaded system can be an SMTN system, that is, it has N threads and is represented as thread 0 to thread N, where N is a positive integer.

[0043] Optionally, the multi-threaded system may include a processing unit and at least one buffer (or cache), the processing unit can be used to execute operations through multiple threads, and the at least one buffer can be used to cache data and / or instructions during the operation of the operations. In a possible embodiment, the processing unit includes an instruction fetch unit, an arbiter, an executor, etc., and the at least one buffer includes an instruction cache, a data cache, etc. Among them, the multiple threads can reuse the instruction fetch unit in the processing unit to obtain instructions for operations running on different threads, the arbiter can be used to arbitrate the execution requests of the multiple threads, and the executor can be used to execute the instructions for the operations corresponding to the multiple threads according to the arbitration result of the arbiter. Figure 2 Denote that thread i reuses the instruction fetch unit to obtain the instruction for the operation as thread i fetches instructions, and the value range of i can be from 0 to N.

[0044] In addition, multiple threads in the multi-threaded system can use different physical resources or use the same physical resource time-sharing. For example, the multiple threads can use different cache regions in the same buffer, or use the instruction fetch unit and the executor in the processing unit time-sharing. Optionally, the physical resources used by different threads among the multiple threads can be obtained through allocation, and the sizes of the physical resources used by any two threads can be the same or different.

[0045] Furthermore, the multi-threaded system may further include a storage system access interface, and this interface can be used to couple with the storage system. Optionally, the storage system may include a storage controller and a memory, and the memory can be a volatile memory, a non-volatile memory, or include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. Exemplarily, the RAM can be a static RAM (SRAM), a dynamic random access memory (DRAM), a synchronous DRAM (SDRAM), and a double data rate synchronous DRAM (DDR SDRAM), etc. The memory is not shown in the figure.

[0046] The above multi-threaded system can be an electronic device, or a system on a chip (SoC) applied to an electronic device, or a chipset including multiple chips, or a module including the SoC or the chipset. The electronic device can be a terminal device or a server. Optionally, the electronic device includes but is not limited to: mobile phone, tablet computer, laptop computer, desktop computer, palmtop computer, ultra-mobile personal computer (umPC), mobile internet device (MID), netbook, camera, camera, wearable device (such as smart watch and smart bracelet, etc.), vehicle-mounted device (such as car, bicycle, electric vehicle, airplane, ship, train, high-speed rail, etc.), virtual reality (VR) device, augmented reality (AR) device, wireless terminal in industrial control, smart home device (such as refrigerator, TV, air conditioner, electric meter, etc.), smart robot, workshop device, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, or wireless terminal in smart home, flying device (such as smart robot, hot air balloon, drone, airplane), etc.

[0047] Figure 3 FIG. is a schematic structural diagram of a multi-threaded service processing device provided by an embodiment of the present application. The multi-thread includes a first thread and a second thread. The device includes: a buffer 10 and a controller 20, and the buffer 10 is coupled to the controller 20; further, the device further includes: a processing unit 30 coupled to the buffer 10 and the controller 20. Optionally, the processing unit 30, the buffer 10, and the controller 20 are integrated together and collectively referred to as a processing core. Figure 4 FIG. shows a schematic diagram of service processing corresponding to the device.

[0048] In this service processing device, the buffer 10 is used to: buffer the first running information of the first service running on the first thread and the second running information of the second service running on the second thread; the controller 20 is used to: obtain the first running information and the second running information, and when it is determined according to the first running information and the second running information that there is at least partial identity between the first service and the second service, determine the first running information as the prefetch information of the second service, so that the second thread runs the second service according to the first running information. Optionally, when the first service and the second service are not identical, the controller 20 will not determine the first running information as the prefetch information of the second service.

[0049] Among them, the first thread and the second thread can be threads with different running speeds among the multiple threads. For example, the first thread can be a thread with a relatively fast running speed among the multiple threads, and the second thread can be any other thread among the multiple threads except the first thread, and the running speed of the second thread is less than that of the first thread. The first thread can be called the leading thread or the main thread, and the second thread can be called the following thread or the slave thread.

[0050] In addition, the first thread as the leading thread can be preset or specified. Exemplarily, the leading thread is determined according to the positions of the multiple threads. For example, the first thread among the multiple threads is set as the leading thread; or, the leading thread is determined by the basic input output system (BIOS) after the multi-threaded system is powered on; or, the leading thread is specified by the software running on the multi-threaded system. The embodiments of the present application do not make specific limitations on the manner of specifying the leading thread.

[0051] In addition, the first service and the second service can be completely identical services, or partially identical services (or called similar or analogous services). The partially identical services can mean that there is partial identity in the data and / or instructions that the two tasks need to access during operation. For example, the first service and the second service can be two parallel computing services in the same high-performance computing. The operations of the above first service and the second service have a large degree of repetition and similarity. For example, the instructions, data, or branches during operation have repetition and similarity.

[0052] Furthermore, the controller 20 determines that at least a part of the first service and the second service is the same. It can also be said that the controller 20 determines to enter the cluster learning optimization (CLO) mode or the service learning mode, that is, the second thread can use the prefetch information determined by the controller 20 according to the first operation information to run the second service, or it can be said that the second thread learns the way the first thread runs or processes the first service to process the second service. Optionally, the opening or closing of the above-mentioned homogenization mode can be controlled by the controller 20, or a software routing can be designed to be opened or closed by the upper operating system. The embodiments of the present application do not make specific limitations on this.

[0053] In a possible embodiment, the first operation information and the second operation information cached in the buffer 10 can be written into the buffer 10 by the processing unit 30. Exemplarily, when the processing unit 30 receives an execution instruction for the first service, the processing unit 30 can run the first service through the first thread, obtain the first operation information according to the operation result of the already-run part of the first service, and write the first operation information into the buffer 10; similarly, when the processing unit 30 receives an execution instruction for the second service, the processing unit 30 can run the second service through the second thread, obtain the second operation information according to the operation result of the already-run part of the second service, and write the second operation information into the buffer 10. Among them, the running speed of the first thread is greater than that of the second thread, or it can be said that the running progress of the first thread is greater than that of the second thread. In this way, the controller 20 determines the prefetch information of the second service according to the first operation information, which can improve the speculative performance of the second thread running the second service, and further improve the performance of service processing. In Figure 4 it, the communication between the processing unit 30 and the buffer 10 during the process of the first thread running the first service is denoted as T0 (the first service), and the communication between the processing unit 30 and the buffer 10 during the process of the second thread running the second service is denoted as T1 (the second service). T0 represents the first thread, and T1 represents the second thread.

[0054] Optionally, the processing unit 30 can periodically or aperiodically write the first operation information of the first service into the buffer 10 during the process of running the first service through the first thread; similarly, the processing unit 30 can also periodically or aperiodically write the second operation information of the second service into the buffer 10 during the process of running the second service through the second thread. After that, the controller 20 can be used to determine whether at least a part of the first service and the second service is the same according to the first operation information and the second operation information in the buffer 10, and when it is determined that the two services have at least a part in common, determine the prefetch information required for the subsequent running of the second service according to the first operation information subsequently written by the processing unit 30.

[0055] Optionally, the buffer 10 includes a first buffer queue and a second buffer queue. The first buffer queue is used to buffer the first running information and filter duplicate addresses in the first running information; the second buffer queue is used to buffer the second running information and filter duplicate addresses in the second running information. Wherein, any one of the first buffer queue and the second buffer queue can be a multi-input and single-output buffer queue, which can be used to balance the bandwidth difference between input and output in addition to filtering duplicate addresses, and improve the transmission efficiency of the running information in the buffer 10.

[0056] In a possible implementation, the first running information and the second running information include at least one of the following obtained during the business execution process: data address information, instruction address information, branch address information, jump information, or value prediction information. Wherein, the data address information (also known as memory access information) is the address of the actually accessed data in the memory. The instruction address information is the address of the actually accessed instruction in the memory. The branch address information is the address of the actually executed branch in the memory. The jump information may include a jump direction and a jump destination. The jump direction is used to indicate the actual jump direction of the jump instruction (i.e., jump or not jump), and the jump destination is the destination address after the actual jump. The value prediction information refers to the value obtained by prediction when the memory access is not completed, and the value may be the result of intermediate execution. The memory corresponding to the above addresses may be on the same or different chips as the above business processing device, and may be a volatile or non-volatile memory, which is not limited in this embodiment.

[0057] Wherein, the address information in the above first running information and second running information may include a virtual address, and the virtual address may be a virtual address issued by software corresponding to the first service and the second service. Optionally, the address information in the above first running information and second running information may further include a physical address corresponding to the virtual address.

[0058] Exemplarily, the first running information includes the data address information obtained by the processing unit 30 during the process of running the first service in the first thread, and the second running information includes the data address information obtained by the processing unit 30 during the process of running the second service in the second thread; or, the first running information includes the instruction address information obtained by the processing unit 30 during the process of running the first service in the first thread, and the second running information includes the instruction address information obtained by the processing unit 30 during the process of running the second service in the second thread; or, the first running information includes the branch address information obtained by the processing unit 30 during the process of running the first service in the first thread, and the second running information includes the branch address information obtained during the process of running the second service in the second thread; or, the first running information includes at least two of the data address information, instruction address information or branch address information obtained by the processing unit 30 during the process of running the first service in the first thread, and the second running information includes at least two of the data address information, instruction address information or branch address information obtained by the processing unit 30 during the process of running the second service in the second thread.

[0059] It can be understood that when the information included in the first running information and the second running information is different, the above-mentioned homogeneous mode (or called service learning mode) service learning can include the learning of different information. For example, when the first running information and the second running information include data address information, the service learning mode includes data cache learning; when the first running information and the second running information include instruction address information, the service learning includes instruction cache learning; when the first running information and the second running information include branch address information, the service learning includes branch prediction learning.

[0060] In a possible embodiment, the first running information includes a plurality of first address information, and the second running information includes a plurality of second address information. For example, the plurality of first address information and the plurality of second address information may include at least one of the above-mentioned data address information, instruction address information or branch address information. At this time, the controller 20 is used to determine that there is at least partial identity between the first service and the second service, and specifically can be used to: when the number of identical address information in the plurality of first address information and the plurality of second address information is greater than a preset threshold, or when the ratio of the number of the identical address information to the number of the plurality of second address information is greater than a preset ratio, it is determined that there is at least partial identity between the first service and the second service. Optionally, when the controller 20 determines that there is at least partial identity between the first service and the second service, it can output an enabling learning signal, and the enabling learning signal can be used to indicate entering the service learning mode.

[0061] Wherein, the preset threshold or the preset ratio can be set in advance, and the embodiments of the present application do not limit the specific values.

[0062] In one example, the multiple first address information included in the first operation information may be arranged according to a certain data structure. When the controller 20 determines that a certain second address information among the multiple second address information included in the second operation information is the same as a certain first address information, the controller 20 may record the same second address information in the data structure, count the same address information, and determine that at least part of the first service and the second service are the same when the number of the same address information is greater than the first preset threshold. For example, the data structure may be a homogeneous training table (CLO train table, CTT), and the CTT may be a naturally wound storage structure that is naturally refreshed by continuous writing. Specifically, as Figure 5 shown, the CTT may include n (n is a positive integer) rows and two columns. Each row of the first column includes a first address information, and each row of the second column can be used to store a first address information. When the controller 20 finds a second address information that is the same as a certain first address information among the multiple second address information, the controller 20 may store the second address information in the corresponding row and color the row. Figure 4 In it, the n first address information are represented as VA1 to VAn, the n second address information are represented as TA1 to TAn, and the colored row is represented by filling.

[0063] In another example, the buffer 10 may cache operation information through a structure that separates the address tag and the data. For example, the first operation information and the second operation information are cached in the buffer 10 through the above separation structure. The address tag is responsible for determining whether the corresponding address is hit, and the data part is responsible for providing the data after the hit. When the first thread accesses the data corresponding to a certain address tag in the buffer 10 during the process of running a service, an access trace will be left in the address tag, and the data corresponding to the address tag belongs to the first thread. After that, when the second thread accesses the same address tag during the process of running the second service and the corresponding access threads are inconsistent, the buffer 10 may trigger an event signal. The controller 20 may detect the event signal and count the number of detected event signals through a counter or the like. When the counted number reaches the preset threshold, it is determined that at least part of the first service and the second service are the same.

[0064] In a possible embodiment, when it is determined that at least part of the first service and the second service are the same, the controller 20 determines the first running information as the prefetch information of the second service, which may include: converting a plurality of first address information included in the first running information obtained by the first thread running the first service into a plurality of third address information corresponding to the second thread; writing the storage information in the cache space corresponding to the plurality of first address information into the cache space corresponding to the plurality of third address information. In this way, the second thread can subsequently run the second service according to the storage information in the cache space corresponding to the plurality of third address information, or be said to obtain the corresponding prefetch information according to the plurality of first address information and run the second service according to the prefetch information. Exemplarily, when the first running information includes data address information, the second thread can obtain the corresponding data according to the data address information and run the second service according to the data; when the first running information includes instruction address information, the second thread obtains the corresponding instruction according to the instruction address information and runs the second service according to the instruction; when the first running information includes branch address information, the second thread obtains the corresponding branch according to the branch address information and runs the second service according to the branch.

[0065] The above conversion of the plurality of first address information into the plurality of third address information corresponding to the second thread can also be referred to as converting the address information in the query request of the first thread into the address information in the prefetch request of the second thread.

[0066] Exemplarily, as Figure 6 shown, when it is determined that at least part of the first service and the second service are the same, the plurality of first address information can be converted into a plurality of third address information corresponding to the second thread (i.e., converting the T0 query request into the T1 prefetch request), and after performing address translation and other processing on the plurality of third address information, according to the policy setting, the storage information in the cache space corresponding to the plurality of first address information is written (or referred to as backfilled) into caches such as level 1 (L1), level 2 (L2), level 3 (L3), etc., that is, the data and / or instructions that the second thread will use in the future are filled into the cache in advance. After that, when the query request of the second thread arrives, the data and / or instructions needed are already in the corresponding cache, so that the second thread runs the second service based on the data and / or instructions in the cache to improve the performance of the second thread running the second service. Figure 6 In the figure, T0 represents the first thread, T1 represents the second thread, the backfill process corresponding to the second thread is represented as T1 backfill, and the backfill process corresponding to the first thread is represented as T0 backfill.

[0067] The cache spaces corresponding to the multiple first address information and / or the cache spaces corresponding to the multiple third address information may include the cache space in the buffer 10, and may also include the cache spaces of other memories outside the buffer 10. For example, the other memories include, but are not limited to, dynamic random access memory (DRAM), flash memory, etc.

[0068] In another possible embodiment, when the controller 20 determines that at least a part of the first service and the second service are the same, the controller 20 may also send an enable signal to the processing unit 30. When the processing unit 30 receives the enable signal, it may convert the multiple first address information into multiple third address information corresponding to the second thread, and write the stored information in the cache space corresponding to the multiple first address information into the cache space corresponding to the multiple third address information.

[0069] Furthermore, as Figure 4 shown, the controller 20 is further configured to: obtain the first security information of the first thread and the second security information of the second thread, and determine the security of the communication between the first thread and the second thread according to the first security information and the second security information, that is, the first security information and the second security information can be used to verify the security of the communication between the first thread and the second thread. The first security information and the second security information may be, respectively, identification information and / or permission information assigned to the second thread and the second thread, etc. Exemplarily, the first security information and / or the second security information includes at least one of the following: address space identifier (ASID), virtual machine identifier (VMID), privilege level information.

[0070] Optionally, the first security information and the second security information may be stored in the buffer 10, or may be stored in other memories outside the buffer 10. In addition, when the controller 20 determines that the communication between the first thread and the second thread is secure according to the first security information and the second security information, the controller 20 may perform the step of determining the first running information as the prefetch information of the second service.

[0071] In an embodiment of the present application, the buffer 10 is used to cache the first running information obtained by the first thread running the first service, and to cache the second running information obtained by the second thread running the second service. The controller 20 is used to determine the prefetch information of the second service according to the first running information when it is determined that there is a partial similarity between the first service and the second service based on the first running information and the second running information, so that the second thread can continue to run the second service according to the prefetch information. Since the first running information is the running information actually obtained by the first thread, and there is a partial similarity between the first service and the second service, that is, there is repetition and similarity between the running of the first service and the running of the second service. In this way, when determining the prefetch information of the second service according to the first running information and enabling the second thread to continue running the second service according to the prefetch information, the accuracy and efficiency of speculation can be greatly improved, thereby improving the overall performance of multi-threaded service processing and reducing the power consumption overhead caused by traditional speculation failures. Specifically, when the controller 20 enters data cache sharing learning, the second thread runs the second service according to the corresponding prefetch information, which can greatly reduce the data cache miss rate; when the controller 20 enters instruction cache sharing learning, the second thread runs the second service according to the corresponding prefetch information, which can greatly reduce the instruction cache miss rate; when the controller 20 enters branch prediction sharing learning, the second thread runs the second service according to the corresponding prefetch information, which can greatly reduce the branch prediction failure rate.

[0072] In addition, the controller 20 determines the prefetch information of the second service according to the first running information when it is determined that there is at least partial similarity between the first service and the second service, and does not determine the prefetch information of the second service according to the first running information when it is determined that the first service and the second service are not similar. In this way, the controller 20 can automatically determine whether to determine the prefetch information of the second service according to the first running information (i.e., automatically determine whether it belongs to a profitable scenario), thereby avoiding mis-triggering and negative benefits in non-target scenarios.

[0073] Further, the above embodiments take the running speed of the first thread being higher than that of the second thread as an example for illustration. In fact, this setting is not used to limit the application scenarios of the embodiments. In practical applications, multiple threads may have the same or different capabilities or processing speeds, and it is not limited which one has stronger capabilities. As long as the running information obtained after one of the threads executes the service can be used to determine the prefetch information of another thread, so as to execute the technical solutions mentioned in this embodiment. For example, the functions of the main thread and the slave thread can be swapped, and the controller can determine the prefetch information of the first thread (slave thread) according to the running information of the second thread (main thread). For another example, the second thread may execute at least part of the functions of the service first due to software scheduling or user selection and obtain the running information, and this running information can be used to determine the prefetch information of the first thread, so that the first thread continues to execute similar homogeneous services to achieve a similar effect. It can be understood that the main thread and the slave thread mentioned in this embodiment, as well as the differences between these two types of threads, are only an applicable scenario, but not used to limit the technical solutions.

[0074] Based on this, an embodiment of the present application further provides a multi-thread service processing method, which can be applied to a service processing device including a buffer and a controller. For the description of the service processing device, reference can be made to the description in the above text. The method includes: the buffer caches the first running information of the first service running on the first thread and the second running information of the second service running on the second thread; when the controller determines that there is at least partial identity between the first service and the second service according to the first running information and the second running information, the controller determines the first running information as the prefetch information of the second service.

[0075] Optionally, the first running information includes a plurality of first address information, and the second running information includes a plurality of second address information. The controller determines that there is at least partial identity between the first service and the second service according to the first running information and the second running information, including: when the number of identical address information existing in the plurality of first address information and the plurality of second address information is greater than a preset threshold, it is determined that there is at least partial identity between the first service and the second service.

[0076] Optionally, the first running information includes a plurality of first address information, and the controller determines the first running information as the prefetch information of the second service, including: converting the plurality of first address information into a plurality of third address information corresponding to the second thread; writing the storage information in the cache space corresponding to the plurality of first address information into the cache space corresponding to the plurality of third address information.

[0077] In a possible embodiment, the method further includes: the controller obtains the first security information of the first thread and the second security information of the second thread, and determines the security of the communication between the first thread and the second thread according to the first security information and the second security information.

[0078] Further, the apparatus further includes a processing unit; the method further includes: the processing unit runs a first service through a first thread to obtain first running information, and writes the first running information into the buffer; the processing unit runs a second service through a second thread to obtain second running information, and writes the second running information into the buffer.

[0079] Optionally, the buffer includes a first buffer queue and a second buffer queue; the buffer caches the first running information of the first service and the second running information, including: the first buffer queue caches the first running information and filters duplicate addresses in the first running information; the second buffer queue caches the second running information and filters duplicate addresses in the second running information.

[0080] It can be understood that all the contents in the above apparatus embodiments can be cited in the corresponding embodiments of the service processing method, and the embodiments of the present application will not be described in detail here.

[0081] In the embodiments of the present application, since the first running information is the running information actually obtained by the first thread, and there are some similarities between the first service and the second service, that is, there are repetitiveness and similarity between the running of the first service and the running of the second service. In this way, when determining the prefetch information of the second service according to the first running information, so that the second thread can continue to run the second service according to the prefetch information, the accuracy and efficiency of speculation can be greatly improved, thereby improving the overall performance of multi-threaded service processing and reducing the power consumption overhead caused by traditional speculation failure.

[0082] In another aspect of the present application, an electronic device is further provided, as Figure 7 shown. The electronic device includes a memory and a multi-threaded processor. The memory is used to store computer instructions, and the multi-threaded processor is used to execute the computer instructions so that the electronic device can implement any one of the multi-threaded service processing methods provided above. Optionally, the multi-threaded processor includes the multi-threaded service processing apparatus provided above. For the specific introduction of the memory, refer to the previous embodiments.

[0083] It can be understood that all the relevant contents involved in the above apparatus embodiments can be cited in the embodiments of the service processing method and the embodiments of the electronic device, and the embodiments of the present application will not be described in detail here.

[0084] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0085] The units described as separate components may or may not be physically separated. The components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0086] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. The readable storage medium may include: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product.

[0087] In another embodiment of the present application, a readable storage medium is further provided. Computer-executable instructions are stored in the readable storage medium. When a device (which can be a single-chip microcomputer, a chip, etc.) or a processor executes the steps in the above method embodiment.

[0088] In yet another embodiment of the present application, a computer program product is further provided. The computer program product includes computer instructions, and the computer instructions are stored in a readable storage medium; at least one processor of the device can read the computer instructions from the readable storage medium, and at least one processor executes the computer instructions to enable the device to perform the steps in the above method embodiment.

[0089] Finally, it should be noted that the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multi-threaded service processing device, characterized in that The device includes: a buffer for buffering first running information of a first service running on a first thread and second running information of a second service running on a second thread; a controller for obtaining the first running information and the second running information, and when it is determined that at least part of the first service and the second service are the same according to the first running information and the second running information, determining the first running information as prefetch information of the second service.

2. The device according to claim 1, wherein The first running information and the second running information include at least one of the following: data address information, instruction address information, branch address information, jump information, or value prediction information.

3. The device according to claim 1 or 2, characterized in that, The first running information includes a plurality of first address information, and the second running information includes a plurality of second address information; The controller is further configured to determine that at least part of the first service and the second service are the same when the number of identical address information in the plurality of first address information and the plurality of second address information is greater than a preset threshold.

4. The device according to any one of claims 1-3, wherein the controller is further configured to obtain first security information of the first thread and second security information of the second thread, and determine the security of communication between the first thread and the second thread according to the first security information and the second security information.

5. The device according to claim 4, characterized in that, The first security information and / or the second security information includes at least one of the following: an address space identifier ASID, a virtual machine identifier VMID, or privilege level information.

6. The device according to any one of claims 1-5, characterized in that, The first running information includes a plurality of first address information, and the controller is further configured to: convert the plurality of first address information into a plurality of third address information corresponding to the second thread; write the storage information in the cache space corresponding to the plurality of first address information into the cache space corresponding to the plurality of third address information.

7. The device according to any one of claims 1-6, characterized in that, The device further includes: a processing unit for: running the first service through the first thread to obtain the first running information, and writing the first running information into the buffer; running the second service through the second thread to obtain the second running information, and writing the second running information into the buffer.

8. The device according to any one of claims 1 to 7, characterized in that, The buffer includes a first buffer queue and a second buffer queue; the first buffer queue is used for buffering the first running information and filtering duplicate addresses in the first running information; the second buffer queue is used for buffering the second running information and filtering duplicate addresses in the second running information.

9. A multi-threaded service processing method, characterized in that, Applied to a service processing device including a buffer and a controller, the method includes: the buffer buffers first running information of a first service running on a first thread and second running information of a second service running on a second thread; when the controller determines that at least part of the first service and the second service are the same according to the first running information and the second running information, determining the first running information as prefetch information of the second service.

10. The method according to claim 9, wherein The first running information and the second running information include at least one of the following: data address information, instruction address information, branch address information, jump information, or value prediction information.

11. The method according to claim 9 or 10, characterized in that, The first running information includes a plurality of first address information, and the second running information includes a plurality of second address information; the controller determines that at least part of the first service and the second service are the same according to the first running information and the second running information, including: When the number of identical address information existing in the plurality of first address information and the plurality of second address information is greater than a preset threshold, the controller determines that at least part of the first service and the second service are the same.

12. The method according to any one of claims 9-11, characterized in that, The method further includes: The controller obtains first security information of the first thread and second security information of the second thread, and determines the security of communication between the first thread and the second thread according to the first security information and the second security information.

13. The method according to claim 12, wherein The first security information and / or the second security information includes at least one of the following: address space identifier ASID, virtual machine identifier VMID, or privilege level information.

14. The method according to any one of claims 9 - 13, characterized in that The first running information includes a plurality of first address information, and the controller determines the first running information as prefetch information of the second service, including: Converting the plurality of first address information into a plurality of third address information corresponding to the second thread; Writing the storage information in the cache space corresponding to the plurality of first address information into the cache space corresponding to the plurality of third address information.

15. The method according to any one of claims 9 - 14, characterized in that, The device further includes a processing unit, and the method further includes: The processing unit runs the first service through the first thread to obtain the first running information, and writes the first running information into the buffer; The processing unit runs the second service through the second thread to obtain the second running information, and writes the second running information into the buffer.

16. The method according to any one of claims 9 - 15, characterized in that, The buffer includes a first buffer queue and a second buffer queue, and the buffer caches the first running information of the first service running on the first thread and the second running information of the second service running on the second thread, including: The first buffer queue caches the first running information and filters duplicate addresses in the first running information; The second buffer queue caches the second running information and filters duplicate addresses in the second running information.

17. An electronic device, characterized in that, The electronic device includes a memory and a multi-threaded processor, the memory is used to store computer instructions, and the multi-threaded processor is used to execute the computer instructions so that the electronic device executes the multi-threaded service processing method according to any one of claims 9-16.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions run on the device, the device is enabled to execute the multi-threaded service processing method according to any one of claims 9-16.