Event processing device and event processing method
The event processing apparatus addresses event congestion in shared resource systems by regulating event reception based on worker performance, enhancing system efficiency and reducing retransmissions.
Patent Information
- Application Number
- PCT/JP2024/001167
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2025-07-24
AI Technical Summary
Existing load distribution systems experience event congestion and retransmission due to unpredictable resource allocation in shared resource environments, particularly in systems using UDP protocols, where traffic volume exceeds worker performance limits, leading to inefficiencies and signal accumulation.
An event processing apparatus with an input control unit that regulates event reception based on performance values from event processing units, determining a regulation value to prevent congestion by rejecting or discarding events exceeding the units' processing capacity, using shared resources to manage traffic volume effectively.
The system effectively controls event congestion and reduces retransmissions by dynamically adjusting traffic regulation based on real-time performance measurements, ensuring efficient resource utilization even in shared resource systems.
Smart Images

Figure JP2024001167_24072025_PF_FP_ABST
Abstract
Description
Event processing device and event processing method
[0001] The present invention relates to an event processing device and an event processing method.
[0002] A load balancing system works in conjunction with a load balancer, which acts as a gateway for receiving signals requesting event processing, and multiple workers, which process signals allocated from the load balancer. Below, we will explain examples 1 to 3 of a load balancing system that handles signals of L (Layer) 5 to L7 protocols using UDP (User Datagram Protocol).
[0003] FIG. 8 is an explanatory diagram showing a first example of an event processing system. The event processing system has a pipeline configuration in which the left side of the drawing is the upstream side (first pipeline stage) and the right side of the drawing is the downstream side (fourth pipeline stage). Each pipeline stage is made up of one or more threads T. Each thread T has a queue that temporarily stores events to be processed (black circles in the drawing). An event received from the previous stage is stored at the end of the queue, and the next event to be processed is read into thread T from the top of the queue. In this event processing system, parallelism can be increased and system processing performance can be improved by increasing the number of pipeline stages or the number of threads per pipeline stage.
[0004] Non-Patent Document 1 describes a second example, SEDA (Staged Event Driven Architecture). SEDA is a type of event-driven architecture with a multi-stage configuration, and employs an architecture in which the following components (stages) are connected in a pipeline: - An event handler that receives events and distributes them to processing threads. - A thread pool that holds multiple threads that process the distributed events.
[0005] 9 is an explanatory diagram showing a cloud-native system of a third example. The load balancer LB temporarily stores received events in an upstream queue and assigns the events to each worker W. When the load balancer LB receives the event processing results from each worker W, it temporarily stores the results in a downstream queue and then returns them to the sender of the event.
[0006] Each worker W temporarily stores events received from the load balancer LB in an upstream queue and references the database DB to process the assigned events. When each worker W receives an event reference result from the database DB, it temporarily stores the event in a downstream queue, then creates an event processing result and returns it to the load balancer LB. The database DB temporarily stores the events received from each worker W in a queue, creates an event reference result and returns it to each worker W. In this way, in the cloud-native system of the third example, each processing unit both sends and receives events, but the architecture is structurally the same as the event processing system of the first example shown in FIG. 8.
[0007] M. Welsh et al., "SEDA: An Architecture for Scalable, Well-Conditioned Internet Services," [online], SOSP'01, Lake Louise, Canada, [Retrieved January 4, 2024], Internet <URL: https: / / www.eecg.toronto.edu / ~amza / ece1747h / papers / seda-sosp01-talk.pdf>
[0008] Each worker is assigned resources, such as a CPU and memory, for signal processing, and has a limited performance limit—the amount of processing that can be performed with those resources at any given time. Therefore, unprocessed signals exceeding the worker's performance are retained in the worker's buffer, resulting in signal congestion. For example, if the signal being processed is UDP, the load balancing system will experience event congestion within the system if it receives traffic volumes exceeding its performance limits. On the other hand, if the signal being processed is an L5-L7 protocol using UDP, a long retention time for an event will result in retransmission, even for signals that the worker has already processed. For example, if too many events are received in the first stage of the pipeline in Figure 8, the events will be retained in the queue of each thread T. Since a long retention time will result in retransmissions, it is necessary to keep the retention time below the retransmission timer.
[0009] Therefore, the load balancer in the first stage of the pipeline performs congestion control by regulating the amount of signals it accepts (incoming traffic volume) to prevent worker congestion. To determine the amount and timing of this congestion control, it is effective to measure the performance of the workers in advance using a traffic generator or other device in a test environment. The traffic generator applies traffic up to the upper limit that the load balancing system can handle and measures that upper limit as a performance value. The load balancer then uses the measured performance value to regulate the amount of incoming traffic, preventing signal congestion and protecting the system from burst traffic and suppressing signal retransmission.
[0010] Here, we consider shared resource systems such as virtual machines and public clouds as load balancing systems in which the same hardware is shared among multiple users. In these shared resource systems, the amount of shared resources allocated to each user fluctuates depending on the usage status of other users sharing the same hardware. For example, the longer the CPU time used by other users, the shorter the CPU time allocated to each user. In this situation where the amount of available resources is uncertain, it is not possible to use values measured in advance using a traffic generator in a test environment, and therefore the content of congestion control cannot be determined.
[0011] In addition, while conventional technologies such as Non-Patent Document 1 present a rough architecture for load balancing events as a load balancing system, they do not provide a means for determining the content of event congestion control in a shared resource system.
[0012] Therefore, a main object of the present invention is to control congestion of events to be processed in a load balancing system in which resources are shared.
[0013] In order to solve the above problems, the event processing device of the present invention comprises the following means: The event processing device comprises an ingress control unit that accepts an event processing request, and a plurality of event processing units that use shared resources to process events assigned by the ingress control unit, each of the event processing units determining whether it is in a congested state based on the amount of unprocessed events assigned to it, and if it is in a congested state, notifying the ingress control unit of the amount of events that it has processed up to a predetermined time ago as its own performance value, and the ingress control unit determining a limit value for the amount of events that it will accept in the future based on the performance values notified from each of the event processing units.
[0014] According to the present invention, it is possible to control congestion of events to be processed in a load balancing system in which resources are shared.
[0015] FIG. 1 is a configuration diagram of an event processing system according to the present embodiment. FIG. 2 is a configuration diagram of an event processing unit according to the present embodiment. FIG. 3 is a configuration diagram of a first example in which the event processing system according to the present embodiment is applied to a pipeline configuration. FIG. 4 is a configuration diagram of a second example in which the event processing system according to the present embodiment is applied to a pipeline configuration. FIG. 5 is a hardware configuration diagram of each device in the event processing system according to the present embodiment. FIG. 6 is a sequence diagram showing the processing of the event processing system according to the present embodiment. FIG. 7 is an example of a time series graph showing the relationship between counter values and performance values according to the present embodiment. FIG. 8 is an explanatory diagram showing a first example of an event processing system. FIG. 9 is an explanatory diagram showing a third example of a cloud native system.
[0016] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings.
[0017] 1 is a configuration diagram of an event processing system 100. The event processing system 100 has an ingress control unit 10 that accepts events (black circles in the figure), and event processing units 20A to 20D that process events assigned from the ingress control unit 10 using resources (the same hardware) shared among multiple users. In other words, the event processing system 100 is a load balancing system in which the ingress control unit 10 functions as a load balancer, and the event processing units 20A to 20D function as workers.
[0018] The ingress control unit 10 is connected to an ingress queue 11 that reads received events from a socket queue (an event queue from the OS to an application), not shown, and temporarily stores them at the end of the queue. Each of the event processing units 20A to 20D is connected to a processing queue 29A to 29D that temporarily stores events assigned by the ingress control unit 10 at the end of the queue. The ingress control unit 10 distributes the load so that the number of events residing in the processing queues 29A to 29D is as uniform as possible, and assigns events retrieved from the head of the ingress queue 11 to each of the event processing units 20A to 20D (the tail of the processing queues 29A to 29D).
[0019] The ingress control unit 10 determines a traffic regulation value that regulates the amount of events to be processed based on the performance values of each event processing unit 20A-20D notified by each event processing unit 20A-20D. The performance value of each event processing unit 20A-20D is, for example, the maximum amount of processing that can be performed with the resources at that time, such as the CPU allocation time. If the ingress control unit 10 receives events that exceed the regulation value per unit time, it rejects the events by returning a message such as "rejected due to congestion" instead of accepting them into the ingress queue 11. Alternatively, the ingress control unit 10 omits the message reply and simply discards the events.
[0020] The ingress control unit 10 calculates the regulation value as, for example, the sum of the performance values notified from each of the event processing units 20A to 20D. This notified performance value is, for example, one of the following values: - The performance value notified currently (or within a predetermined period from now) when each of the event processing units 20A to 20D is congested (at its performance limit). - If each of the event processing units 20A to 20D is not congested, the maximum value of the performance values notified in the past. Note that the event processing unit 20A notifies the ingress control unit 10 of its own current performance value (the calculation method is explained in Figure 2) when it is congested.
[0021] FIG. 2 is a configuration diagram of the event processing unit 20A. Note that while the configuration of the event processing unit 20A will be described in FIG. 2, the other event processing units 20B to 20D also have the same configuration. The event processing unit 20A includes a reading unit 21, a processing unit 22, an output unit 23, a reading counter 21C, an output counter 23C, a limit determination unit 24, and a performance measurement unit 25. The reading unit 21 reads the next event to be processed from the top of the processing side queue 29A. The processing unit 22 processes the event read by the reading unit 21. The output unit 23 outputs the result of event processing by the processing unit 22. The reading counter 21C counts the amount of events read by the reading unit 21 (the number of events to be processed) per unit time. The output counter 23C counts the amount of events output by the output unit 23 (the number of events resulting from processing) per unit time.
[0022] The limit determination unit 24 determines whether the current state is congested (i.e., performance limit state) based on the amount of events accumulated in the processing queue 29A. For example, the limit determination unit 24 determines that its own event processing unit 20A is in a congested state when, for example, an event in which the amount of events in the processing queue 29A exceeds a threshold occurs a predetermined number of times or more consecutively within a predetermined period T1. In other words, the limit determination unit 24 detects, as a congested state, signs that the events accumulated in the processing queue 29A are beginning to overflow from the queue. Furthermore, after determining that the event processing unit 20A is in a congested state, the limit determination unit 24 determines that its own event processing unit 20A has recovered from the congested state to a normal state (a state that is not congested) when an event in which the amount of events in the processing queue 29A does not exceed a threshold occurs a predetermined number of times or more consecutively within a predetermined period T1.
[0023] The performance measurement unit 25 measures the count increment value of the read counter 21C per unit time or the count increment value of the output counter 23C per unit time as the performance value per unit time for its own event processing unit 20A. That is, each event processing unit 20A-20D measures the number of times it reads unprocessed events for further processing or the number of times it outputs processed events as the amount of events processed up to a predetermined time ago. For example, let us express the counter value of the read counter 21C at each time as [time (minutes:seconds), counter value], and assume that the following values are measured: [9:00, 180], [9:01, 183], [9:02, 186], [9:03, 191], [9:04, 197], and [9:05, 203]. In this case, the performance measurement unit 25 calculates the performance value per predetermined period T2 (e.g., 3 seconds) as follows:・The count increment value for the 3 seconds from 9:00 to 9:02 = 6 (186 - 180), so the performance value at 9:02 = 2 (6 / 3). ・The count increment value for the 3 seconds from 9:03 to 9:05 = 12 (203 - 191), so the performance value at 9:05 = 4 (12 / 3).
[0024] The counter value of the read counter 21C and the counter value of the output counter 23C each increase by one for processing one event. Here, the size of the variable that stores the counter value is finite (e.g., 2 bytes). Therefore, it is desirable for the performance measurement unit 25 to reset (e.g., to 0) a counter value that is close to the upper limit value (e.g., exceeding 65,000) after completing the calculation of the performance value so as not to exceed the upper limit value (e.g., 65,535) of the variable.
[0025] If the limit determination unit 24 determines that the current state is congested, the performance measurement unit 25 notifies the ingress control unit 10 of the measured current performance value. This allows the performance measurement unit 25 to restrict the acceptance of events so as to prevent further overflow of the processing queue 29A. Here, "present" does not necessarily refer strictly to the current time, but may also refer to a period up to three seconds before the current time. The following parameters indicating the period are set values, for example, entered in advance by an administrator. The predetermined period T1 is the period during which the limit determination unit 24 determines whether congestion (performance limit) exists. The predetermined period T2 is the period during which the performance measurement unit 25 calculates the performance value during congestion. The values of T1 and T2 may be the same or different. However, because both the values of T1 and T2 are related to the sensitivity of congestion detection, it is desirable that there is not too much difference between them (for example, 0.1 seconds and 100 seconds).
[0026] Thus, the event processing system 100, which includes an ingress control unit 10 that accepts event processing requests and multiple event processing units 20A that process events assigned by the ingress control unit 10 using shared resources, executes the following procedure. (Step 1) Each event processing unit 20A-20D determines whether it is in a congested state based on the amount of unprocessed events assigned to it. To this end, each event processing unit 20A-20D checks the amount of events in its processing queue 29A-29D every predetermined period T1 (seconds). When it finds that the amount of events exceeds a threshold, it determines that it is in a congested state. (Step 2) Each event processing unit 20A-20D in a congested state notifies the ingress control unit 10 of the amount of events processed up to the previous predetermined time as its own performance value. To this end, each event processing unit 20A-20D reads the counter value of the read counter 21C or output counter 23C and calculates its current performance value by dividing the counter value by the increment per predetermined period T2 (seconds) by the predetermined period T2. (Step 3) Based on the performance values notified by each of the event processing units 20A to 20D, the ingress control unit 10 sets the sum of these performance values as the limit value for the amount of events to be accepted. In other words, the ingress control unit 10 sets the performance value of each of the event processing units 20A to 20D as the limit value, and rejects any event amount that exceeds the limit value.
[0027] An example in which the event processing system 100 described above in Figures 1 and 2 is applied to a pipeline configuration such as that shown in Figure 8 will be described below. Figure 3 is a configuration diagram of a first example in which the event processing system 100 is applied to a pipeline configuration. In the four-stage pipeline configuration shown in Figure 3, the first stage of the pipeline is the ingress control unit 10, and the second to fourth stages of the pipeline are the event processing unit 20A. This configuration is effective when congestion frequently occurs in the second stage of the pipeline.
[0028] 4 is a configuration diagram of a second example in which the event processing system 100 is applied to a pipeline configuration. In the four-stage pipeline configuration shown in FIG. 4, the first and second stages of the pipeline are the ingress control unit 10, and the third and fourth stages of the pipeline are the event processing unit 20A. This configuration is effective when congestion frequently occurs in the third stage of the pipeline. In this way, it is desirable to have a configuration in which the ingress control unit 10 and the event processing unit 20A are separated by the number of pipeline stages at which congestion frequently occurs.
[0029] FIG. 5 is a hardware configuration diagram of each device in the event processing system 100. The event processing system 100 may be configured as a single device including the ingress control unit 10 and the event processing unit 20A, or may be configured by connecting the ingress control unit 10 and the event processing unit 20A via a network. In either case, each device in the event processing system 100 is configured as a computer 900 including a CPU 901, RAM 902, ROM 903, HDD 904, communication I / F 905, input / output I / F 906, and media I / F 907. The communication I / F 905 is connected to an external communication device 915. The input / output I / F 906 is connected to an input / output device 916. The media I / F 907 reads and writes data from a recording medium 917. Furthermore, the CPU 901 controls each unit by executing a program (event processing program) loaded into the RAM 902. This program (also called an application, or simply "app") can be distributed via a communication line or recorded on a recording medium 917 such as a USB memory and distributed.
[0030] 6 is a sequence diagram showing the processing of the event processing system 100. In S101, the ingress control unit 10 receives an event request from incoming traffic and stores it at the end of the ingress queue 11. In S102, the ingress control unit 10 distributes the event request read from the top of the ingress queue 11 to each of the event processing units 20A to 20D. The following illustrates an example in which an event request is distributed to the event processing unit 20A. In S111, the reading unit 21 of the event processing unit 20A stores the received event request at the end of the processing queue 29A. In S112, the processing unit 22 of the event processing unit 20A processes the event request read from the processing queue 29A by the reading unit 21, and returns the processing result to the ingress control unit 10 from the output unit 23.
[0031] In S113, the limit determination unit 24 determines whether congestion (performance limit) exists based on the event volume (amount of backlogged event requests) in the processing queue 29A. If the event volume in the processing queue 29A falls below the first threshold (or falls below the first threshold multiple times in succession), the limit determination unit 24 determines No (not at performance limit) in S113 and returns the process to S111. On the other hand, if the event volume in the processing queue 29A exceeds the second threshold (or exceeds the second threshold multiple times in succession), the limit determination unit 24 determines Yes (performance limit) in S113 and proceeds to S114. In S114, the performance measurement unit 25 calculates the current performance value of the event processing unit 20A based on the counter value read from the read counter 21C or the output counter 23C, as described in FIG. 2, and notifies the ingress control unit 10 of the performance value. Whether the performance measurement unit 25 uses the read counter 21C, which is not affected by an event processing error, or the output counter 23C, which is affected by an event processing error, can be determined according to the system policy.
[0032] In S103, the ingress control unit 10 receives the performance value calculated in S114 from the event processing unit 20A and reflects it in the regulation details (regulation value) for future incoming traffic, as described in Fig. 1. As a result, when the ingress control unit 10 receives incoming traffic (S104), similar to S101, if the incoming traffic is subject to regulation (exceeds the regulation value) (S105, Yes), it discards the incoming traffic (S106). On the other hand, if the result in S105 is No, the ingress control unit 10 returns to S101 and processes the incoming traffic received in S104 in the same way as the incoming traffic received in S101.
[0033] Here, to facilitate understanding of the calculations made by the performance measurement unit 25, an explanation will be provided with reference to FIG. 7. FIG. 7 is an example of a time-series graph showing the relationship between counter values and performance values. The vertical axis of this time-series data represents traffic per unit time (t / s) as a performance value. The dashed curve L1 indicates changes in the amount of resources available to the output unit 23 of the event processing unit 20A. Even if the amount of resources provided is fixed, the amount of available resources changes over time because the resources are shared with other users. The thick solid curve L2 indicates changes in the amount of events allocated to the event processing unit 20A. During the period when (value of dashed curve L1) > (value of thick solid curve L2), a sufficient amount of resources is allocated for event processing, the amount of events in the processing queue 29A is gradually consumed, and the congestion state is resolved.
[0034] On the other hand, during periods C1, C2, and C3 where the value of the dashed curve L1 is less than the value of the thick solid curve L2, the amount of resources allocated to event processing is insufficient, and the amount of events in the processing queue 29A gradually accumulates, causing congestion. In other words, periods C1, C2, and C3 are periods during which the limit determination unit 24 determines that a congestion state exists. The thin solid curve L3 indicates the counter value read from the read counter 21C or the output counter 23C during each of these periods C1, C2, and C3. In other words, because only the amount of events processed corresponds to the amount of available resources, the thin solid curve L3 and the dashed curve L1 have the same value.
[0035] Therefore, the performance measurement unit 25 notifies the ingress control unit 10 of the counter value read from the read counter 21C or the output counter 23C shown in the thin solid curve L3 during each period C1, C2, and C3 as the value (performance value) of the dashed curve L1, thereby allowing the ingress control unit 10 to grasp the degree of resource shortage.
[0036] On the other hand, during periods where the value of the dashed curve L1 is greater than the value of the thick solid curve L2, resources are not fully utilized and remain in reserve, resulting in a discrepancy between the counter value and the performance value. Therefore, during periods when there is no congestion, the performance measurement unit 25 does not need to notify the ingress control unit 10 of the counter value as a performance value. Instead, when the performance measurement unit 25 of each event processing unit 20A-20D is not in a congested state, it notifies the ingress control unit 10 of a substitute value based on the performance value measured so far as its own performance value. The substitute value may be 1.5 times the maximum performance value observed so far or an initial setting value entered by the administrator. This allows the ingress control unit 10 to constantly determine the traffic regulation value, regardless of whether there is congestion, such as when the event processing unit 20A is not in a congested state but the other event processing units 20B-20D are in a congested state.
[0037] [Effect] The present invention is an event processing system 100 having an ingress control unit 10 that accepts event processing requests, and multiple event processing units 20A-20D that use shared resources to process events assigned by the ingress control unit 10, wherein each event processing unit 20A-20D determines whether it is in a congested state based on the amount of unprocessed events assigned to it, and if it is in a congested state, notifies the ingress control unit 10 of the amount of events it has processed up to a predetermined time ago as its own performance value, and the ingress control unit 10 determines a limit value for the amount of events it will accept in the future based on the performance values notified from each of the event processing units 20A-20D.
[0038] As a result, the event processing system 100 is able to control congestion by appropriately setting the regulation value of the ingress control unit 10, even in cases where resources are shared, such as in virtual machines or systems on the cloud.
[0039] The present invention is characterized in that, when each event processing unit 20A to 20D is not in a congested state, it notifies the ingress control unit 10 of an alternative value based on the performance value measured up to that point as its own performance value.
[0040] This allows the event processing system 100 to set an appropriate regulation value for the ingress control unit 10 even in a situation where a congested event processing unit 20A and event processing units 20B to 20D that are not congested are mixed.
[0041] The present invention is characterized in that the number of times each event processing unit 20A to 20D reads out an unprocessed event for future processing or outputs an event that has already been processed is measured as the amount of events processed up to a predetermined time ago.
[0042] This allows the event processing system 100 to properly grasp the volume of events to be processed simply by counting the number of inputs and outputs of events, regardless of the content of the events being processed.
[0043] 10 Ingress control unit 11 Ingress queue 20A to 20D Event processing unit 21 Reading unit 21C Reading counter 22 Processing unit 23 Output unit 23C Output counter 24 Limit determination unit 25 Performance measurement unit 29A to 29D Processing side queue 100 Event processing system (event processing device)
Claims
1. An event processing apparatus, comprising: an input control unit that receives a processing request for an event; and a plurality of event processing units that process events assigned from the input control unit using shared resources, wherein each event processing unit determines whether it is in a congestion state based on the amount of unprocessed events assigned to itself, and if it is in the congestion state, notifies the input control unit of the amount of events processed from the present time to a predetermined time before as its performance value, and the input control unit determines a regulation value for the amount of events to be received hereafter based on the performance values notified from the respective event processing units.
2. The event processing apparatus according to claim 1, wherein each event processing unit notifies the input control unit of an alternative value based on the performance value measured so far as its performance value if it is not in the congestion state.
3. The event processing apparatus according to claim 1, wherein each event processing unit measures the number of times of reading unprocessed events to be processed hereafter or the number of times of outputting processed events as the amount of events processed from the present time to a predetermined time before.
4. An event processing method by an event processing apparatus, comprising: an input control unit that receives a processing request for an event; and a plurality of event processing units that process events assigned from the input control unit using shared resources, wherein each event processing unit determines whether it is in a congestion state based on the amount of unprocessed events assigned to itself, and if it is in the congestion state, notifies the input control unit of the amount of events processed from the present time to a predetermined time before as its performance value, and the input control unit determines a regulation value for the amount of events to be received hereafter based on the performance values notified from the respective event processing units.
Citation Information
Patent Citations
Intelligent load balancing method based on c / s (Client / Server) architecture
CN103368864A
Centralized Load Balancer Using Weighted Hash Functions
JP2019522936A