Data processing method and apparatus, and device

By establishing a one-to-one correspondence between consumer threads and processing units, and using a ring directed queue mechanism and load balancing strategy, the thread scheduling overhead caused by the change in the number of computing cores in the core group is solved, and data processing performance is improved.

WO2025152399A1PCT designated stage expired Publication Date: 2025-07-24HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/109757
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-15
Filing Date
2024-08-05
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

When the number of cores is calculated in the kernel group changes, the thread scheduling overhead increases, affecting the performance of the kernel group.

Method used

By establishing a one-to-one correspondence between consumer threads and processing units, using a ring directed queue mechanism and load balancing strategy, threads can be avoided switching between processing units, load balancing is achieved, and lock conflicts are reduced.

Benefits of technology

Without kernel-state operations, load balancing between processing units is realized, data processing performance is improved, and thread switching overhead is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024109757_24072025_PF_FP_ABST
    Figure CN2024109757_24072025_PF_FP_ABST
Patent Text Reader

Abstract

A data processing method, comprising: a consumer thread among N consumer threads processes data issued by at least one among M producers, data issued by one producer being processed by one consumer thread, and M being an integer greater than or equal to N; and when a processing unit corresponding to a first consumer thread of data issued by a first producer among the M producers meets a condition, a second consumer thread among the N consumer threads replaces the first consumer thread to process the data issued by the first producer. According to the described method, the overhead of thread scheduling can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing method, device and equipment

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on January 15, 2024, with application number 202410059641.8 and application name “Data Processing Methods, Devices and Equipment”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a data processing method, apparatus, and device. Background Art

[0003] Computing devices such as servers process multiple services simultaneously. To reduce the impact of different services (such as competition for computing resources), different core groups are assigned to different services. Each core group assigned to a service includes one or more computing cores, each dedicated to processing that service.

[0004] In some scenarios, the number of computing cores in a core group may change. For example, if business A1 is under heavy pressure and the core group assigned to it is heavily loaded, one or more cores in the less loaded core group may need to be moved to the core group assigned to business A1. For another example, if business A2 is under less pressure and the core group assigned to it is also lightly loaded, some computing cores in that core group may need to be shut down to reduce power consumption.

[0005] The outflow of computing cores from a core group or the shutdown of some computing cores in a core group will increase the thread scheduling overhead in the core group, thereby affecting the performance of the core group.

[0006] Summary of the Invention

[0007] The present application provides a data processing method, apparatus, and device that can reduce thread scheduling overhead and improve the performance of a processing unit cluster.

[0008] In a first aspect, a data processing method is provided, which is applied to a data processing device, the device comprising: N processing units, where N is an integer greater than or equal to 1; the N processing units correspond one-to-one to N consumer threads, and the consumer threads run on the processing units corresponding to the consumer threads; the method comprises: a consumer thread among the N consumer threads processes data sent by at least one producer among M producers; wherein the data sent by one producer is processed by one consumer thread, where M is an integer greater than or equal to N; when the processing unit corresponding to the first consumer thread that processes the data sent by the first producer among the M producers meets a condition, a second consumer thread among the N consumer threads takes over from the first consumer thread to process the data sent by the first producer.

[0009] In this method, one processing unit runs one consumer thread, one consumer thread can process data sent by one or more producers at the same time, and at the same time, data sent by one producer is processed by one consumer thread.

[0010] In other words, in this method, consumer threads and processing units always maintain a one-to-one correspondence, eliminating thread switching between processing units. If there's a one-to-one correspondence between producers and processing units, or between producer and consumer threads, the producer corresponding to the consumer thread can be switched to the corresponding processing unit. This allows load balancing between processing units without requiring kernel-mode operations, thus avoiding the significant overhead of kernel-mode operations and improving data processing performance.

[0011] In one possible implementation, the processing unit corresponding to the first consumer thread satisfies any of the following conditions: the processing unit corresponding to the first consumer thread is shut down, the processing unit corresponding to the first consumer thread fails, and the processing unit corresponding to the first consumer thread is used to process data sent by producers other than the M producers.

[0012] That is to say, when the number of processing units used to process data sent by M producers decreases, the remaining processing units can take over the reduced processing units to process the data sent by the corresponding producers, ensuring that the data sent by each of the M producers is processed.

[0013] In one possible implementation, consumer threads other than the first consumer thread among N consumer threads form a first circular directed queue; the method also includes: when the load of the processing unit corresponding to the second consumer thread is greater than the load of the processing unit corresponding to the third consumer thread, and the second consumer thread is used to process data sent by at least two of the M producers, the third consumer thread takes over the second consumer to process data sent by some of the at least two producers; wherein, in the first circular queue, the second consumer thread is the next consumer thread of the third consumer thread.

[0014] In this implementation, N consumer threads form a circular directed queue. The current consumer thread in this queue can take over from its next consumer thread to process data from the corresponding producer, but it does not take over from other consumer threads to process data from the corresponding producer. Furthermore, the next consumer thread after the current consumer thread does not take over from the current consumer thread to process data from the corresponding producer. This circular directed queue mechanism ensures a unidirectional flow of producers between consumer threads, significantly reducing lock conflicts when switching producers between consumer threads.

[0015] In one possible implementation, the first consumer thread is used to simultaneously process data sent by the first producer and data sent by the second producer among the M producers; the processing unit corresponding to the first consumer thread satisfies the conditions including: the load of the processing unit corresponding to the first consumer thread is greater than the load of the processing unit corresponding to the second consumer thread.

[0016] When the current consumer thread processes data sent by two or more producers, the consumer thread running on a processing unit with a smaller load among the N processing units can take over the current consumer thread to process the data sent by some of the two or more producers, thereby reducing the load of the processing unit where the current consumer thread is located and achieving load balancing among the N processing units.

[0017] In a possible implementation, N consumer threads form a second circular directed queue; wherein, in the second circular directed queue, the first consumer thread is the next consumer thread of the second consumer thread.

[0018] In this implementation, N consumer threads form a circular directed queue. The current consumer thread in this queue can take over from its next consumer thread to process data from the corresponding producer, but it does not take over from other consumer threads to process data from the corresponding producer. Furthermore, the next consumer thread after the current consumer thread does not take over from the current consumer thread to process data from the corresponding producer. This circular directed queue mechanism ensures a unidirectional flow of producers between consumer threads, significantly reducing lock conflicts when switching producers between consumer threads.

[0019] In one possible implementation, the storage space corresponding to the producer is used to store data sent by the producer; the consumer thread processes the data sent by at least one producer among the M producers, including: the consumer thread is used to obtain data from the storage space corresponding to at least one producer and process the obtained data.

[0020] According to a second aspect, a data processing device is provided, comprising: N processing units, where N is an integer greater than or equal to 1; the N processing units correspond one-to-one to N consumer threads, and the consumer threads run on the processing units corresponding to the consumer threads; wherein the consumer threads among the N consumer threads are used to process data sent by at least one producer among M producers; wherein the data sent by one producer is processed by one consumer thread, where M is an integer greater than or equal to N; when the processing unit corresponding to the first consumer thread processing the data sent by the first producer among the M producers meets a condition, the second consumer thread among the N consumer threads is used to take over the first consumer thread to process the data sent by the first producer.

[0021] In one possible implementation, the processing unit corresponding to the first consumer thread satisfies any of the following conditions: the processing unit corresponding to the first consumer thread is shut down, the processing unit corresponding to the first consumer thread fails, and the processing unit corresponding to the first consumer thread is used to process data sent by producers other than the M producers.

[0022] In one possible implementation, consumer threads other than the first consumer thread among N consumer threads form a first circular directed queue; when the load of the processing unit corresponding to the second consumer thread is greater than the load of the processing unit corresponding to the third consumer thread, and the second consumer thread is used to process data sent by at least two of the M producers, the third consumer thread is used to take over the second consumer to process data sent by some of the at least two producers; wherein, in the first circular queue, the second consumer thread is the next consumer thread of the third consumer thread.

[0023] In one possible implementation, the first consumer thread is used to simultaneously process data sent by the first producer and data sent by the second producer among the M producers; the processing unit corresponding to the first consumer thread satisfies the conditions including: the load of the processing unit corresponding to the first consumer thread is greater than the load of the processing unit corresponding to the second consumer thread.

[0024] In a possible implementation, N consumer threads form a second circular directed queue; wherein, in the second circular directed queue, the first consumer thread is the next consumer thread of the second consumer thread.

[0025] In a possible implementation, the storage space corresponding to the producer is used to store data sent by the producer; the consumer thread is used to obtain data from the storage space corresponding to at least one producer and process the obtained data.

[0026] In a third aspect, a computing device is provided, comprising the data processing apparatus provided in the second aspect and M producers.

[0027] In a fourth aspect, a computer-readable storage medium is provided, comprising computer program instructions. When the computer program instructions are executed by a computing device, the computing device executes the method provided in the first aspect.

[0028] In a fifth aspect, a computer program product comprising instructions is provided, which, when executed by a computing device, causes the computing device to execute the method provided in the first aspect.

[0029] Among them, the beneficial effects of the second to fifth aspects can be found in the above description of the beneficial effects of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] FIG1 is a schematic diagram of a computing core scheduling of a core group;

[0031] FIG2 is a schematic diagram of a system architecture provided in an embodiment of the present application;

[0032] FIG3 is a schematic diagram of a system architecture provided in an embodiment of the present application;

[0033] FIG4 is a schematic diagram of a change in the number of processing units in a processing unit cluster provided by an embodiment of the present application;

[0034] FIG5 is a flow chart of a data processing method provided in an embodiment of the present application;

[0035] FIG6 is a schematic diagram of a data processing method provided in an embodiment of the present application;

[0036] FIG7 is a schematic diagram of a producer-based scheduling strategy provided in an embodiment of the present application;

[0037] FIG8 is a schematic diagram of a circular directed queue provided in an embodiment of the present application;

[0038] FIG9 is a schematic diagram of a data processing method provided in an embodiment of the present application;

[0039] FIG10 is a schematic diagram of a circular directed queue provided in an embodiment of the present application;

[0040] FIG11 is a schematic structural diagram of a data processing device provided in an embodiment of the present application;

[0041] FIG12 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0042] The following describes the solutions provided by the embodiments of the present application in conjunction with the accompanying drawings. In the embodiments of the present application, "plurality" refers to two or more, and "multiple" refers to two or more. Terms such as "first" and "second" are used only to distinguish similar objects and do not necessarily describe a specific order or quantity of objects.

[0043] To facilitate understanding of the solutions provided by the embodiments of the present application, the technical terms that may be involved in the embodiments of the present application are first introduced.

[0044] Processing unit: A computer module, component, or assembly with computing power used to perform one or more computing tasks. A processing unit can be a processor (e.g., a single-core processor) or the computing core of a multi-core processor. A processor can be a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), or a neural-network processing unit (NPU). A processing unit can run one or more threads simultaneously.

[0045] Processing unit cluster: Also known as a processing unit group, it is composed of multiple processing units. A multi-core processor can be divided into one or more processing unit clusters, each of which includes multiple computing cores in the multi-core processor.

[0046] Thread: Sometimes also called lightweight process (LWP), is the smallest unit of computer operation scheduling. A thread is run by one processing unit at a time.

[0047] Consumer thread: This refers to the thread that processes data sent by the producer. A consumer thread can include one or more threads. A consumer thread is run by one processing unit at a time.

[0048] Producer: This refers to the entity that produces and sends the data it produces. The producer can be located on a device other than the device where the consumer thread resides, or the generator can be an application running on the device where the consumer thread resides, and so on.

[0049] In the related art, as shown in FIG1 , a core group includes multiple computing cores, each of which runs a consumer thread. The consumer thread is bound to a producer. The consumer thread is used to process data sent by the producer to which the consumer thread is bound. When at least one computing core in the core group is disconnected (for example, the one or more computing cores are shut down or the one or more computing cores are moved to another core group), the consumer thread running on the at least one computing core is executed by the remaining computing cores in the core group. Since the number of producers has not changed, in order to ensure that the remaining computing cores carry the production business evenly, it is necessary to break the 1:1 correspondence between consumer threads and processing units and schedule consumer threads among the remaining computing cores in the core group. For example, as shown in FIG1 , a core group consisting of 4 computing cores processes data sent by 4 producers. Each computing core runs a consumer thread. The consumer thread and the producer have a 1:1 binding relationship, and one consumer thread processes data sent by one producer. When a computing core is released from the core group, four consumer threads are run by three processing units, resulting in one computing core running two consumer threads. To achieve load balancing, the computing core running the two consumer threads must be frequently adjusted, meaning that the consumer threads must be constantly switched from one processing unit to another. Switching the processing unit running the consumer threads involves kernel state operations, resulting in high switching overhead and impacting overall service performance.

[0050] Figure 2 shows a schematic diagram of a system architecture provided by an embodiment of the present application. As shown in Figure 2, the system architecture includes a processing unit cluster 100, which may include N processing units, and the N processing units correspond one-to-one to N consumer threads, that is, in the processing unit cluster 100, the processing units and the consumer threads are in a 1:1 binding relationship. Among them, the consumer thread runs on the processing unit corresponding to the consumer thread. N is an integer greater than or equal to 1. For the convenience of description, the N processing units can be set to include processing unit 111, processing unit 112, processing unit 113, processing unit 114, etc. Among them, processing unit 111 corresponds to consumer thread 121, processing unit 112 corresponds to consumer thread 122, processing unit 113 corresponds to consumer thread 123, processing unit 114 corresponds to consumer thread 124, and so on.

[0051] As shown in FIG2 , the system architecture includes M producers, where M is an integer greater than or equal to N. The processing unit cluster 100 is configured to process data sent by the M producers. A consumer thread in the processing unit cluster 100 processes data sent by at least one of the M producers. That is, a consumer thread can process data sent by one or more of the M producers at the same time. Data sent by one of the M producers is processed by a consumer thread in the processing unit cluster 100. That is, data sent by one producer is processed by one consumer thread at the same time.

[0052] In some embodiments, as shown in FIG2 , the M producers may include producer 131, producer 132, producer 133, producer 134, etc. Among them, consumer thread 121 processes data sent by producer 131, consumer thread 122 processes data sent by producer 132, consumer thread 123 processes data sent by producer 133, and consumer thread 124 processes data sent by producer 134.

[0053] In some embodiments, as shown in FIG3 , the M producers may include producer 131, producer 132, producer 133, producer 134, producer 135, etc. Among them, consumer thread 121 processes data sent by producer 131, consumer thread 122 processes data sent by producer 132, consumer thread 123 processes data sent by producer 133, and consumer thread 123 processes data sent by producer 134 and data sent by producer 135.

[0054] In some embodiments, as shown in FIG2 or FIG3 , a producer corresponds to a storage space for storing data sent by the producer. A consumer thread retrieves the data sent by the producer from the storage space and processes the retrieved data. Exemplarily, the storage space may be a cache, also known as a buffer space.

[0055] To ensure data processing consistency, only one consumer thread is allowed to retrieve data from the same storage space at the same time. In other words, data sent by a producer is only processed by one consumer thread at a time. For example, when a consumer thread retrieves data from a storage space, it applies a lock to that storage space, preventing other consumer threads from retrieving data from that storage space.

[0056] In some scenarios, one or more processing units in processing unit cluster 100 are no longer used to process data sent by one or more of the M producers. For example, in the scenario shown in FIG4 , the M producers are used to generate and send data based on the business data sent by business A2. The producer corresponding to business A1 is used to generate and send data based on the business data sent by business A1. The data sent by the producer corresponding to business A1 is processed by the consumer thread corresponding to business A1, and the consumer thread of business A1 runs on a processing unit in processing unit cluster 200. When business A2 sends a small amount of business data and business A1 sends a large amount of business data, the load on the processing units in processing unit cluster 100 is relatively small, while the load on the processing units in processing unit cluster 200 is relatively large. In this case, one or more processing units in processing unit cluster 100 can be moved to processing unit cluster 200. These one or more processing units are no longer used to process data sent by the corresponding producers in the M producers, but are instead used to process data sent by the producer corresponding to business A1.

[0057] When one or more processing units in the processing unit cluster 100 are no longer used to process data sent by one or more of the M producers, the consumer threads running on the one or more processing units for processing the data sent by the one or more producers are no longer used to process the data sent by the one or more producers, wherein the consumer threads can be shut down or moved along with the movement of the one or more processing units. The data sent by the one or more producers are processed by the remaining processing units in the processing unit cluster 100. The data sent by the one or more producers are processed by the consumer threads running on the remaining processing units in the processing unit cluster 100.

[0058] Therefore, when one or more processing units in the processing unit cluster 100 are no longer used to process data sent by one or more of the above-mentioned M producers, the data sent by the one or more producers are processed by consumer threads running by the remaining processing units in the processing unit cluster 100. This does not involve switching of consumer threads between different processing units, thereby avoiding the overhead caused by switching of consumer threads between different processing units.

[0059] The above examples introduce the system architecture provided by the embodiment of the present application. Next, in conjunction with the system architecture, the data processing method provided by the embodiment of the present application is introduced.

[0060] Through this method, the processing units in processing unit cluster 100 process the data sent by the M producers. As described above, processing unit cluster 100 includes N processing units, each of which corresponds to N consumer threads. A consumer thread runs on the corresponding processing unit, where N is an integer greater than or equal to 1. As shown in Figure 5, this method includes the following steps.

[0061] In step 501 , a consumer thread among N consumer threads processes data sent by at least one producer among M producers; wherein, data sent by one producer is processed by one consumer thread, and M is an integer greater than or equal to N.

[0062] Specifically, a consumer thread can process data sent by a producer, for example, consumer thread 121 can process data sent by producer 131. A consumer thread can also process data sent by multiple producers at the same time, for example, consumer thread 124 can process data sent by producer 134 and data sent by producer 135 at the same time.

[0063] To ensure data processing consistency, data sent by a producer is processed by a consumer thread at the same time. When a consumer thread processes data sent by a producer, it can apply a lock to the data sent by the producer to prevent other consumer threads from processing the data sent by the producer.

[0064] In some embodiments, the storage space corresponding to the producer is used to store data sent by the producer; the consumer thread is used to obtain data from the storage space and process the obtained data. In one example, the storage space is a cache, which can also be called a cache space.

[0065] After the producer generates data, the generated data can be stored in the storage space corresponding to the producer. The consumer thread used to process the data sent by the producer obtains the data from the storage space corresponding to the producer and processes the obtained data.

[0066] Step 502: When a processing unit corresponding to a first consumer thread processing data sent by a first producer among the M producers meets a condition, a second consumer thread among the N consumer threads takes over from the first consumer thread to process the data sent by the first producer.

[0067] The first consumer thread may be one or more threads among the N consumer threads mentioned above. Accordingly, the processing unit corresponding to the first consumer thread is one or more processing units among the N processing units mentioned above.

[0068] When the processing unit corresponding to the first consumer thread meets the conditions, the processing unit corresponding to the first consumer thread no longer processes the data sent by the first producer, and the first consumer thread no longer processes the data sent by the first producer. The second consumer thread takes over the processing of the data sent by the first producer.

[0069] In some embodiments, the condition satisfied by the processing unit corresponding to the first consumer thread means that the processing unit corresponding to the first consumer thread no longer or is unable to process data sent by the first producer. In this embodiment, the number of processing units in processing unit cluster 100 used to process data sent by the M producers has been reduced. In one example, the condition satisfied by the processing unit corresponding to the first consumer thread specifically means that the processing unit corresponding to the first thread has been shut down. For example, when the load on the processing units in processing unit cluster 100 is low, some processing units in processing unit cluster 100 may be shut down, including the processing unit corresponding to the first consumer thread. In another example, the condition satisfied by the processing unit corresponding to the first consumer thread specifically means that the processing unit corresponding to the first thread has failed. The failure of the processing unit prevents the processing unit from operating the consumer thread normally and, therefore, from processing data sent by the first producer. In yet another example, the condition satisfied by the processing unit corresponding to the first consumer thread specifically means that the processing unit corresponding to the first consumer thread is used to process data sent by producers other than the M producers. For example, in the scenario shown in Figure 5, the processing unit corresponding to the first consumer thread is moved to processing unit cluster 200 to process data sent by the producer corresponding to business A1.

[0070] In an example of this embodiment, referring to FIG6 , the first consumer thread can be set to be consumer thread 124, and accordingly, the processing unit corresponding to the first consumer thread is processing unit 114. The second consumer thread can also be set to be consumer thread 123. The data sent by producer 134 is originally processed by consumer thread 124. When the processing unit corresponding to consumer thread 124 meets the conditions, consumer thread 123 takes over from consumer thread 124 to process the data sent by producer 134. That is, when the processing unit corresponding to consumer thread 124 meets the conditions, consumer thread 124, which originally processed the data sent by producer 134, no longer processes the data sent by producer 134, but consumer thread 123 processes the data sent by producer 134. Consumer thread 123 runs on processing unit 113, and consumer thread 123 processing the data sent by producer 134 can also be referred to as processing unit 113 processing the data sent by producer 134.

[0071] In this way, the data sent by the producer can be processed by different processing units without involving thread switching between processing units, thus avoiding the overhead caused by thread switching between different processing units.

[0072] The second consumer thread takes over from the first consumer thread to process the data sent by the first producer. Furthermore, the second consumer thread also processes data sent by other producers. This results in the second consumer thread processing data sent by two or more producers, resulting in a heavy load on the processing unit where the second consumer thread resides. To this end, load balancing can be performed among the processing units of the N consumer threads, excluding the processing unit corresponding to the first consumer thread.

[0073] Among them, in the embodiment of the present application, the load balancing between processing units is based on the producer scheduling strategy, that is, scheduling the consumer thread to process the data sent by the producer. As shown in Figure 7, the producer is constantly switched between consumer threads to achieve load balancing between consumer threads. Consumer threads and processing units correspond one to one, and the load balancing between consumer threads is also the load balancing between processing units. Compared with the related art of scheduling consumer threads between processing units, this is a user-mode scheduling strategy based on the producer scheduling strategy, which does not involve kernel-mode operations, thereby avoiding the large overhead of kernel-mode operations.

[0074] In one example, as described above, to ensure data processing consistency, data sent by a producer is processed by a single consumer thread at a time. While a consumer thread processes data sent by a producer, it can apply a lock to that data to prevent other consumer threads from processing the same data. When multiple consumer threads simultaneously process data sent by the same producer, a lock conflict occurs. To avoid lock conflicts, or to prevent multiple consumer threads from processing data from the same producer simultaneously, a circular directed queue mechanism is designed. This mechanism effectively avoids lock conflicts, as detailed below.

[0075] Referring to FIG. 8 , the N consumer threads, excluding the first consumer thread, can be configured as consumer thread 121, consumer thread 122, and consumer thread 123. Consumer threads 121, 122, and 123 form a circular directed queue 800 that is connected at the end. In circular directed queue 800, the next consumer thread after consumer thread 121 is consumer thread 122, the next consumer thread after consumer thread 122 is consumer thread 123, and the next consumer thread after consumer thread 123 is consumer thread 121.

[0076] The load of the processing unit corresponding to the current consumer thread can be compared with the load of the processing unit corresponding to the next consumer thread in the circular directed queue 800. If the load of the processing unit corresponding to the current consumer thread is less than the load of the processing unit corresponding to the next consumer thread, and the next consumer thread processes data sent by at least two producers, the current consumer thread will take over from the next consumer thread to process data sent by some of the at least two producers. For example, consumer thread 123 can be set to process data sent by producer 133 and data sent by producer 134. When the load of processing unit 112 corresponding to consumer thread 122 is less than the load of processing unit 113 corresponding to consumer thread 123, consumer thread 122 will take over from consumer thread 123 to process data sent by one of producer 133 and producer 134, and consumer thread 123 will continue to process data sent by the other of producer 133 and producer 134.

[0077] In this way, the producer can be guaranteed to flow in one direction between consumer threads, which greatly reduces the lock conflict when switching the consumer thread that processes the data sent by the producer.

[0078] The loads of different processing units can be compared in the following two ways.

[0079] In one approach, the loads of different processing units can be compared using the scheduling delays of different processing units. The scheduling delay, which can also be referred to as the execution wait time, refers to the time between the moment a producer sends data and the moment the data is processed by a consumer thread on the processing unit. Still taking processing unit 112 and processing unit 113 as an example, if the most recent one or more scheduling delays of processing unit 112 are less than the most recent one or more scheduling delays of processing unit 112, it can be confirmed that the load of processing unit 112 is less than the load of processing unit 113. In one example, the load of processing unit 112 and the load of processing unit 113 can be compared using formula (1).

[0080] in, is the most recent k scheduling delays of the j processing unit (eg, processing unit 112). is the most recent k scheduling delays of the j+1 processing unit (e.g., processing unit 113). diff_max is a constant, which is the maximum difference in the most recent k scheduling delays between different processing units. Where k is an integer greater than or equal to 1, and k is a preset value.

[0081] In another embodiment, as described above, the producer stores the generated data in the storage space corresponding to the producer, and the consumer thread retrieves the data from the storage space and processes the retrieved data. If all data in the storage space corresponding to the current consumer thread has been processed and no new data has arrived, that is, the current consumer thread is in an idle state, then it is determined that the load of the processing unit corresponding to the current consumer thread is less than the load of the processing unit corresponding to the next consumer thread in the circular directed queue 800.

[0082] In one example, the storage space corresponding to the producer can also be called a queue. A queue steal mechanism can be enabled based on a partition tree structure to ensure unidirectional flow between different queue consumer threads, thereby ensuring unidirectional flow of producers between consumer threads.

[0083] In some embodiments, when the first consumer thread is used to process data sent by the first producer and data sent by the second producer among M producers at the same time, the processing unit corresponding to the first consumer thread satisfies the following conditions: the load of the processing unit corresponding to the first consumer thread is greater than the load of the processing unit corresponding to the second consumer thread. That is, in this embodiment, the first consumer thread is used to process data sent by two or more producers at the same time. If the load of the processing unit corresponding to the first consumer thread is greater than the load of the processing unit corresponding to the second consumer thread, the second consumer thread can take over the first consumer thread to process data sent by some of the two or more producers. The first producer is one of the two or more producers. In addition, the first consumer thread continues to process data sent by producers other than the first producer among the two or more producers.

[0084] For example, referring to FIG9 , consumer thread 124 can be configured to simultaneously process data sent by producer 134 and data sent by producer 135. If the load of processing unit 114 corresponding to consumer thread 124 is greater than the load of processing unit 113 corresponding to consumer thread 123, consumer thread 123 takes over from consumer thread 124 to process data sent by one of producer 134 and producer 135, while consumer thread 124 continues to process data sent by the other of producer 134 and producer 135.

[0085] In one example of this embodiment, the N consumer threads described above form a circular directed queue 900 that is connected at the end. In this example, the first consumer thread is the next consumer thread of the second consumer thread in the circular directed queue 900. In this example, the load of the processing unit corresponding to the current consumer thread can be compared with the load of the processing unit corresponding to the next consumer thread in the circular directed queue 900. If the load of the processing unit corresponding to the current consumer thread is less than the load of the processing unit corresponding to the next consumer thread, and the next consumer thread processes data sent by at least two producers, the current consumer thread takes over from the next consumer thread to process data sent by some of the at least two producers.

[0086] For example, the N consumer threads can be configured as consumer thread 121, consumer thread 122, consumer thread 123, and consumer thread 124. As shown in FIG10 , consumer threads 121, 122, 123, and 124 form a circular directed queue 900 that is connected at the end. In circular directed queue 900, the next consumer thread after consumer thread 121 is consumer thread 122, the next consumer thread after consumer thread 122 is consumer thread 123, the next consumer thread after consumer thread 123 is consumer thread 124, and the next consumer thread after consumer thread 124 is consumer thread 121. Consumer thread 124 can be configured to execute data sent by producer 134 and producer 135 simultaneously. If the load of the processing unit 114 corresponding to the consumer thread 124 is greater than the load of the processing unit 113 corresponding to the consumer thread 123, the consumer thread 123 takes over the processing of the data sent by one of the producers 134 and 135 from the consumer thread 124, and the consumer thread 124 continues to process the data sent by the other one of the producers 134 and 135.

[0087] The specific method of comparing the loads can be referred to the above introduction and will not be repeated here.

[0088] In this way, the producer can be guaranteed to flow in one direction between consumer threads, which greatly reduces the lock conflict when switching the consumer thread that processes the data sent by the producer.

[0089] In summary, in the data processing method provided in the embodiments of the present application, a one-to-one correspondence is maintained between consumer threads and processing units, and threads are not switched between processing units. In situations where there is not a one-to-one correspondence between producers and processing units, or where there is not a one-to-one correspondence between producer and consumer threads, the producer corresponding to a processing unit can be switched by switching the producer between consumer threads or switching the consumer thread between producers. This allows load balancing between processing units without requiring kernel-mode operations, thereby avoiding the significant overhead of kernel-mode operations and improving data processing performance.

[0090] Furthermore, in the data processing method provided in the embodiment of the present application, a ring-shaped directed queue mechanism is used to ensure unidirectional flow of producers between consumer threads, thereby greatly reducing lock conflicts when switching producers between consumer threads.

[0091] Referring to FIG11 , an embodiment of the present application further provides a data processing device 1100. The device 1100 includes: N processing units 1110, where N is an integer greater than or equal to 1; the N processing units 1110 correspond to N consumer threads one by one, and the consumer threads run on the processing units 1110 corresponding to the consumer threads; wherein,

[0092] The consumer thread of the N consumer threads is used to process data sent by at least one producer of the M producers; wherein the data sent by one producer is processed by one consumer thread, and M is an integer greater than or equal to N;

[0093] When the processing unit 1110 corresponding to the first consumer thread processing the data sent by the first producer among the M producers meets the conditions, the second consumer thread among the N consumer threads is used to take over the first consumer thread to process the data sent by the first producer.

[0094] In some embodiments, the processing unit 1110 corresponding to the first consumer thread satisfies any of the following conditions:

[0095] The processing unit 1110 corresponding to the first consumer thread is shut down, the processing unit 1110 corresponding to the first consumer thread fails, or the processing unit 1110 corresponding to the first consumer thread is used to process data sent by producers other than the M producers.

[0096] In an example of this embodiment, the consumer threads other than the first consumer thread among the N consumer threads form a first circular directed queue; when the load of the processing unit 1110 corresponding to the second consumer thread is greater than the load of the processing unit 1110 corresponding to the third consumer thread, and the second consumer thread is used to process data sent by at least two of the M producers, the third consumer thread is used to take over from the second consumer to process data sent by some of the at least two producers; wherein, in the first circular queue, the second consumer thread is the next consumer thread of the third consumer thread.

[0097] In some embodiments, the first consumer thread is used to simultaneously process data sent by the first producer and data sent by the second producer among the M producers; the processing unit 1110 corresponding to the first consumer thread satisfies the conditions including: the load of the processing unit 1110 corresponding to the first consumer thread is greater than the load of the processing unit 1110 corresponding to the second consumer thread.

[0098] In an example of this embodiment, the N consumer threads form a second circular directed queue; wherein, in the second circular directed queue, the first consumer thread is the next consumer thread of the second consumer thread.

[0099] In some embodiments, the storage space corresponding to the producer is used to store data sent by the producer; the consumer thread is used to obtain data from the storage space corresponding to the at least one producer and process the obtained data.

[0100] Referring to FIG. 12 , an embodiment of the present application provides a computing device 1200 . As shown in FIG. 12 , computing device 120 includes the data processing apparatus 1100 shown in FIG. 11 , M producers 1210 , and memory 1220 . M is an integer greater than or equal to 1. Memory 1220 is used to store executable programs. The processing unit in data processing apparatus 1100 is used to execute the executable program stored in memory 1220 , enabling computing device 1200 to perform the method shown in FIG. 5 .

[0101] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on a computing device, the computing device executes the method shown in FIG5 .

[0102] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the method shown in Figure 5.

[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data processing method, characterized in that, Applied to a data processing device, the device includes: N processing units, where N is an integer greater than or equal to 1; the N processing units correspond to N consumer threads one by one, and the consumer threads run on the processing units corresponding to the consumer threads; the method includes: The consumer threads among the N consumer threads process data sent by at least one of the M producers; wherein, the data sent by one producer is processed by one consumer thread, and M is an integer greater than or equal to N; When the processing unit corresponding to the first consumer thread that processes the data sent by the first producer among the M producers meets the condition, the second consumer thread among the N consumer threads takes over the first consumer thread to process the data sent by the first producer.

2. The method according to claim 1, wherein The processing unit corresponding to the first consumer thread meeting the condition includes any one of the following: The processing unit corresponding to the first consumer thread is shut down, the processing unit corresponding to the first consumer thread fails, the processing unit corresponding to the first consumer thread is used to process data sent by producers other than the M producers.

3. The method according to claim 2, wherein The consumer threads other than the first consumer thread among the N consumer threads form a first circular directed queue; the method further includes: When the load of the processing unit corresponding to the second consumer thread is greater than the load of the processing unit corresponding to the third consumer thread, and the second consumer thread is used to process data sent by at least two of the M producers, the third consumer thread takes over the second consumer to process the data sent by some of the at least two producers; Wherein, in the first circular queue, the second consumer thread is the next consumer thread of the third consumer thread.

4. The method according to claim 1, wherein The first consumer thread is simultaneously used to process the data sent by the first producer and the data sent by the second producer among the M producers; the processing unit corresponding to the first consumer thread meeting the condition includes: the load of the processing unit corresponding to the first consumer thread is greater than the load of the processing unit corresponding to the second consumer thread.

5. The method according to claim 4, wherein The N consumer threads form a second circular directed queue; wherein, in the second circular directed queue, the first consumer thread is the next consumer thread of the second consumer thread.

6. The method according to any one of claims 1-5, characterized in that The storage space corresponding to the producer is used to store the data sent by the producer; The consumer thread processes data sent by at least one of the M producers, including: the consumer thread is used to obtain data from the storage space corresponding to the at least one producer and process the obtained data.

7. A data processing device, characterized in that, The device includes: N processing units, where N is an integer greater than or equal to 1; the N processing units correspond to N consumer threads one by one, and the consumer threads run on the processing units corresponding to the consumer threads; wherein, The consumer threads among the N consumer threads are used to process data sent by at least one of the M producers; wherein, the data sent by one producer is processed by one consumer thread, and M is an integer greater than or equal to N; When the processing unit corresponding to the first consumer thread that processes the data sent by the first producer among the M producers meets the condition, the second consumer thread among the N consumer threads is used to take over the first consumer thread to process the data sent by the first producer.

8. The device according to claim 7, characterized in that, The processing unit corresponding to the first consumer thread meeting the condition includes any one of the following: The processing unit corresponding to the first consumer thread is shut down, the processing unit corresponding to the first consumer thread fails, the processing unit corresponding to the first consumer thread is used to process the data sent by producers other than the M producers.

9. The device according to claim 8, characterized in that The consumer threads among the N consumer threads other than the first consumer thread form a first circular directed queue; When the load of the processing unit corresponding to the second consumer thread is greater than the load of the processing unit corresponding to the third consumer thread, and the second consumer thread is used to process the data sent by at least two producers among the M producers, the third consumer thread is used to take over the second consumer to process the data sent by some of the at least two producers; Wherein, in the first circular queue, the second consumer thread is the next consumer thread of the third consumer thread.

10. The device according to claim 7, wherein The first consumer thread is simultaneously used to process the data sent by the first producer and the data sent by the second producer among the M producers; the processing unit corresponding to the first consumer thread meeting the condition includes: the load of the processing unit corresponding to the first consumer thread is greater than the load of the processing unit corresponding to the second consumer thread.

11. The device according to claim 10, wherein, The N consumer threads form a second circular directed queue; wherein, in the second circular directed queue, the first consumer thread is the next consumer thread of the second consumer thread.

12. The device according to any one of claims 7-11, characterized in that, The storage space corresponding to the producer is used to store the data sent by the producer; the consumer thread is used to obtain data from the storage space corresponding to at least one producer and process the obtained data.

13. A computing device, characterized in that, It includes the data processing device according to any one of claims 7-12 and M producers.

14. A computer-readable storage medium, characterized in that, It includes computer program instructions, when the computer program instructions are executed by a computing device, the computing device executes the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Data processing method, device and equipment

    CN120315851A

  • Data processing method and device

    CN107239343A

  • Message processing method and device, and computer readable storage medium

    CN108509299A

  • A task processing method and device

    CN109684091A

  • Batch processing of messages

    US20180217882A1