A high-throughput stream processing method and device

By building a bolt model and thread pool, using a combination of memory cache and hard disk cache to process real-time streaming data, the overload problem of real-time stream computing system during burst data stream processing is solved, and high throughput and high resource utilization is achieved.

CN113934531BActive Publication Date: 2025-05-23ZTE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010610083.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-29
Publication Date
2025-05-23
Estimated Expiration
2040-06-29

AI Technical Summary

Technical Problem

In the prior art, real-time stream computing systems are prone to overload when processing burst data streams, resulting in low data efficiency and serious waste of hardware resources.

Method used

By building a bolt model and thread pool, the data receiving thread determines whether the memory cache space is sufficient. If it is sufficient, the data will be cached, otherwise it will be saved to the hard disk; the data processing thread obtains data from the memory cache for processing.

Benefits of technology

It improves the throughput and resource utilization of the stream processing system, reduces operation and maintenance costs, and improves the stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113934531B_ABST
    Figure CN113934531B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides a high-throughput stream processing method and device. The stream processing method in the present invention is: the data receiving thread determines whether the remaining cache space of the current data memory cache is greater than a preset threshold. If so, the received data is cached in the memory cache, otherwise, the received data is saved to the local hard disk; the data processing thread obtains data from the memory cache and performs calculation processing on the obtained data. In the scheme of the present invention, the data receiving Bolt adopts a two-level cache of memory cache and hard disk cache, while the data processing Bolt does not have an independent memory cache, but directly obtains data from the memory cache in the data receiving Bolt for calculation processing. The data receiving thread and the data processing thread are independent of each other, which improves the throughput, the utilization rate of computing resources and storage resources, and improves the stability of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to, but are not limited to, the field of large data stream technology, and specifically, relate to, but are not limited to, a high-throughput stream processing method and device. Background Art

[0002] As a technology that can process data quickly and provide quick feedback, real-time stream computing has been widely used in various fields. There are a large number of excellent stream computing engines, such as Apache storm (distributed real-time computing system), Apache Filnk (general data processing platform), Spark Streaming, etc. In these systems, data generation is completely determined by the data source. The dynamic changes and inconsistent states of the data source cause the rate of the data stream to present a bursty feature, and the bursty feature of the data stream often leads to overload. There are also several reasons for overload: network congestion, sudden peaks of user requests, etc. In real-time stream computing, overload is common and difficult to avoid.

[0003] The stream computing engine based on Apache Storm has its own back pressure mechanism, but its efficiency is low and hardware resources are seriously wasted. Therefore, it is necessary to improve throughput, increase computing resource utilization, and reduce operation and maintenance costs. Summary of the invention

[0004] The embodiments of the present invention provide a high-throughput stream processing method and device, which mainly solve the technical problem of low data efficiency and serious waste of hardware resources.

[0005] To solve the above technical problems, an embodiment of the present invention provides a high-throughput stream processing method, including:

[0006] The data receiving thread determines whether the remaining cache space of the current data memory cache is greater than a preset threshold, and if so, caches the received data into the memory cache, otherwise, saves the received data to the local hard disk;

[0007] The data processing thread obtains data from the memory cache and performs computational processing on the obtained data.

[0008] The high-throughput stream processing method further comprises:

[0009] Constructing a bolt model, the bolt model includes: a data receiving bolt model, a data processing bolt model, and a data sending bolt model; the data receiving bolt model is used to receive a data message sent by the data sending bolt model, and the data processing bolt model is used to obtain data from the data receiving bolt model and perform calculation processing;

[0010] A thread pool is established according to the constructed bolt model, and the thread pool includes: a data receiving thread, a data processing thread, and a data sending thread; the data receiving thread is used to receive data messages sent by the data sending thread, and the data processing thread is used to obtain data from the data receiving thread for calculation and processing;

[0011] The ratio of the data receiving thread to the data processing thread is configured according to a preset ratio.

[0012] The embodiment of the present invention also provides a high-throughput stream processing device, including a data receiving module and a data processing module;

[0013] The data receiving module is used to receive data and determine whether the remaining cache space of the current data memory cache is greater than a preset threshold, and if so, cache the received data into the memory cache, otherwise, save the received data into the local hard disk;

[0014] The data processing module is used to obtain data from the memory cache and perform calculations on the obtained data.

[0015] The high-throughput stream processing apparatus further comprises:

[0016] A creation module, the creation module is used to build a bolt model; the bolt model includes: a data receiving bolt model, a data processing bolt model, and a data sending bolt model; the data receiving bolt model is used to receive a data message sent by the data sending bolt model, and the data processing bolt model is used to obtain data from the data receiving bolt model and perform calculation processing;

[0017] The creation module is also used to establish a thread pool according to the constructed bolt model, and the thread pool includes: a data receiving thread, a data processing thread, and a data sending thread; the data receiving thread is used to receive data messages sent by the data sending thread, and the data processing thread is used to obtain data from the data receiving thread for calculation and processing;

[0018] A configuration module is used to configure the ratio of the data receiving thread and the data processing thread according to a preset ratio.

[0019] An embodiment of the present invention further provides a computer storage medium, wherein the computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the high-throughput stream processing method as described above.

[0020] According to the high-throughput stream processing method, device and computer storage medium provided by the embodiments of the present invention, the data receiving thread determines whether the remaining cache space of the current data memory cache is greater than a preset threshold. If so, the received data is cached in the memory cache, otherwise, the received data is saved to the local hard disk; the data processing thread obtains data from the memory cache and performs computational processing on the obtained data. The data receiving Bolt uses a two-level cache of memory cache and hard disk cache, while the data processing Bolt does not have an independent memory cache, but directly obtains data from the memory cache in the data receiving Bolt for computational processing. The data receiving thread and the data processing thread are independent of each other. In some embodiments, the throughput, computing resource and storage resource utilization are improved, and the system stability is improved.

[0021] Other features and corresponding beneficial effects of the present invention are described in the latter part of the specification, and it should be understood that at least part of the beneficial effects become obvious from the description in the specification of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 A flow chart of a high-throughput stream processing method provided in Embodiment 1 of the present invention;

[0023] Figure 2 A unified data cache state diagram provided in the first embodiment of the present invention;

[0024] Figure 3 A unified task scheduling state diagram provided in the first embodiment of the present invention;

[0025] Figure 4 A flow chart of a high-throughput stream processing method provided in Embodiment 2 of the present invention;

[0026] Figure 5 A diagram of a high-throughput stream processing device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the following is a further detailed description of the embodiments of the present invention through specific implementation methods combined with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0028] Embodiment 1:

[0029] In order to improve throughput, increase computing resource utilization, and reduce operation and maintenance costs, this embodiment provides a high-throughput stream processing method. Figure 1 , the method comprises the following steps:

[0030] S101, the data receiving thread determines whether the remaining cache space of the current data memory cache is greater than a preset threshold, and if so, caches the received data into the memory cache, otherwise, saves the received data to the local hard disk;

[0031] S102: The data processing thread obtains data from the memory cache and performs computational processing on the obtained data.

[0032] In this embodiment, the high-throughput stream processing method further includes: constructing a bolt model, the bolt model including: a data receiving bolt model, a data processing bolt model, and a data sending bolt model; the data receiving bolt model is used to receive a data message sent by the data sending bolt model, and the data processing bolt model is used to obtain data from the data receiving bolt model and perform calculation processing;

[0033] A thread pool is established according to the constructed bolt model, and the thread pool includes: a data receiving thread, a data processing thread, and a data sending thread; the data receiving thread is used to receive data messages sent by the data sending thread, and the data processing thread is used to obtain data from the data receiving thread for calculation and processing;

[0034] The ratio of the data receiving thread to the data processing thread is configured according to a preset ratio.

[0035] In the Apache storm distributed real-time computing system, the conventional bolt model is a level-by-level bolt followed by a level-by-level bolt to form a complete data receiving and processing link. In the embodiment of the present invention, a bolt model with multiple dimensions is constructed, including a data receiving bolt model, a data processing bolt model, and a data sending bolt model. These three groups of bolt models with independent dimensions, i.e., three independent data links, each link only assumes a single responsibility. The data processing bolt model is used to process data, the data receiving bolt model is used to receive data, and the data sending bolt model is used to send data.

[0036] The data receiving bolt in this embodiment has two levels of storage space, namely, data memory cache space and local hard disk storage space, which are used to store received data. When the data receiving thread receives data, it will determine the remaining cache space size of the current data memory cache. If the remaining space of the current memory cache is greater than the preset threshold, the received data will be cached in the memory cache for caching, otherwise, the received data will be saved in the local hard disk. The data processing thread directly obtains the data to be processed from the memory cache for calculation and processing. In this embodiment, when the remaining memory cache space is less than 5%, the received data can be saved to the local hard disk, or when the memory cache space is full, that is, the remaining cache space is 0, the received data can be saved to the local hard disk. The specific threshold can be set flexibly, and this embodiment does not limit it. The processing thread directly obtains data from the memory cache, and the data processing thread uses a process-level public queue to actively obtain the data to be processed from the memory cache, and continuously calculates and processes the data cached in the memory cache.

[0037] In an example of this embodiment, the ratio of data receiving threads and data processing threads is configured according to a preset ratio. Specifically, 10% of the processing threads can be configured to process data, and the others are used as data receiving threads and data sending threads. The specific ratio can be flexibly configured according to the data processing requirements of the system, and this embodiment does not limit it.

[0038] In this embodiment, the received data is uniformly cached, such as Figure 2 The figure shows the unified data cache state diagram provided in this embodiment. The data receiving thread uses spout to obtain the data message sent by the data sending thread from the real-time queue, and the spout sends the data to the data receiving thread according to the rules. After receiving the data, the data receiving thread will group the acquired data according to the preset rules, and determine whether the remaining space of the current memory cache is greater than the preset threshold. If it is determined that the remaining space of the current memory cache is greater than the preset threshold, the received data will be inserted into the data list of the memory cache, otherwise, the received data will be saved to the local hard disk for caching. The space size of the memory cache described in this embodiment is automatically initialized according to a certain ratio based on the memory of the local machine when the system starts, and cache space is allocated to the memory cache.

[0039] The data receiving thread of the data receiving bolt will dynamically calculate the time required to process the received data, and dynamically adjust the granularity of the data cache based on the calculated data processing time, thereby ensuring that the real-time performance of the data will not be affected in any way.

[0040] The data receiving thread will scan the data cache status in the memory cache and the local hard disk in real time. If there is cached data in the local hard disk and the current remaining space of the memory cache is greater than a preset threshold, the data receiving thread will cache the data cached in the local hard disk into the memory cache space so that the data processing thread can process the data in time, because the data processing thread will only obtain the data to be processed from the memory cache.

[0041] The data processing thread obtains data from the memory cache for calculation and processing according to a preset time interval. Specifically, the data processing thread actively obtains a set of tuple data from the memory cache for calculation and processing at preset time intervals. The data processing thread will process the data obtained from the memory cache in a timely manner. After processing the currently obtained data, it will immediately actively obtain the next set of data to continue processing, thereby continuously obtaining data from the memory cache and processing the data, ensuring that the data in the memory cache can be processed in a timely manner. If the data processing thread does not obtain data from the memory cache, the processing thread will enter a sleep state, and after sleeping for a preset time, it will obtain data from the memory cache again for processing. When the data processing thread does not obtain the data that needs to be processed, entering a sleep state can save the CPU and avoid idle time. If the data processing bolt frequently and continuously accesses and obtains the memory cache, it will cause the CPU to surge, which seriously wastes resources. Therefore, in this embodiment, if the data processing thread does not obtain data, it will enter a sleep state for a preset time and wait for data to be processed again. In this embodiment, the data processing bolt continuously obtains data from the memory cache for processing, and enters a preset sleep state when no data to be processed is obtained. There will be no situation where server resources are idle due to a certain imbalance in the data, but the data cannot be processed in time.

[0042] refer to Figure 3 , Figure 3 The figure shows the state diagram of the unified task scheduling provided by this embodiment. In the worker node, the task data to be processed cached in the memory cache is uniformly scheduled, and the task to be processed is obtained from the memory cache by the task distributor and assigned to the data processing thread for calculation and processing. The data processing thread actively obtains the data task to be processed from the memory cache for processing through the task allocation. After processing a batch of data, new data to be processed is requested again to continue processing, so that data is continuously obtained from the memory cache for processing, ensuring that the data in the memory cache can be processed in a timely manner.

[0043] Actual test results show that the high-throughput stream processing method provided by the present invention has a throughput 10 times higher than that of the native system. Through the overall strategy of a single server, the bolt model is divided into three dimensions: data receiving bolt, data processing bolt, and data sending bolt, which improves the overall real-time performance of the system, saves memory resources, and copes with strong bursts. The memory resource configuration is sufficient, and the data receiving bolt adopts a two-level cache of memory cache and local hard disk, which greatly improves the stability of the system. The receiving thread and the processing thread are simply proportionally configured, and no manual intervention is required, and no additional hardware is required, thereby reducing the operation and maintenance cost.

[0044] The high-throughput stream processing method provided by the embodiment of the present invention is as follows: the data receiving thread determines whether the remaining cache space of the current data memory cache is greater than a preset threshold value. If so, the received data is cached in the memory cache, otherwise, the received data is saved to the local hard disk; the data processing thread obtains data from the memory cache and performs calculation processing on the obtained data. The data receiving Bolt adopts two-level caches, namely, memory cache and hard disk cache, while the data processing Bolt does not have an independent memory cache, but directly obtains data from the memory cache in the data receiving Bolt for calculation processing. The data receiving thread and the data processing thread are independent of each other, the data reception adopts a unified cache, and the data processing adopts a unified task allocation, which improves the throughput of the stream processing system, the utilization rate of computing resources and storage resources, improves the system stability, and reduces the operation and maintenance costs.

[0045] Embodiment 2:

[0046] In order to improve throughput, enhance computing resource utilization, and reduce operation and maintenance costs, this embodiment provides a high-throughput stream processing method.

[0047] This embodiment describes the high-throughput flow processing method with reference to a specific example. There is a current MR measurement report of a certain telecommunications operator in City A, which requires that the data can be processed in real time.

[0048] like Figure 4 As shown, a flow chart of a high-throughput stream processing method is provided for this embodiment, including the following steps:

[0049] S401, allocate memory space inside the Worker as a unified memory cache space;

[0050] Memory space is allocated inside the Worker as a unified memory cache space to cache data received by the receiving bolt.

[0051] S402, initializing the unified memory cache space size;

[0052] The size of the unified memory cache space is automatically initialized according to a certain ratio based on the local memory size when the system starts.

[0053] S403, using spout to obtain data from the real-time queue;

[0054] Use spout to get data from the real-time queue and send it to the data receiving bolt according to the rules.

[0055] S404, the data receiving bolt groups the received data according to preset rules;

[0056] The data receiving thread uses spout to obtain the data message sent by the data sending thread from the real-time queue, and the spout sends the data to the data receiving thread according to the rules. After receiving the data, the data receiving thread will group the obtained data according to the preset rules.

[0057] S405, determining whether the current memory cache space usage has reached 100%;

[0058] After receiving the data, the data receiving thread determines whether the current memory cache usage has reached 100%. If the usage has not reached 100%, it enters step S406 to cache the received data into the unified memory cache space for unified caching; otherwise, it enters step 407 to save the received data in the local high-speed hard disk for unified caching.

[0059] S408, real-time scanning of the current local high-speed hard disk cache status and unified memory cache usage. If there is data in the local high-speed hard disk and the unified memory cache space usage has not reached 100%, the data in the local high-speed hard disk is cached in the unified memory.

[0060] The data receiving thread will scan the data cache status in the memory cache and the local hard disk in real time. If there is cached data in the local hard disk and the current remaining space of the memory cache is greater than a preset threshold, the data receiving thread will cache the data cached in the local hard disk into the memory cache space so that the data processing thread can process the data in time, because the data processing thread will only obtain the data to be processed from the memory cache.

[0061] S409, the data processing bolt actively obtains a set of data from the unified memory cache for processing every 100ms;

[0062] S410, data processing bolt did not obtain data, sleep 100ms;

[0063] The data processing thread obtains data from the memory cache for calculation and processing according to a preset time interval. Specifically, the data processing thread actively obtains a set of tuple data from the memory cache for calculation and processing every 100ms time interval. The data processing thread will process the data obtained from the memory cache in a timely manner. After processing the currently obtained data, it will immediately actively obtain the next set of data to continue processing, thereby continuously obtaining data from the memory cache and processing the data, ensuring that the data in the memory cache can be processed in a timely manner. If the data processing thread does not obtain data from the memory cache, the processing thread will enter a sleep state, and after sleeping for 100ms, it will obtain data from the memory cache again for processing.

[0064] In the high-throughput stream processing method provided by the embodiment of the present invention, the data receiving Bolt adopts two-level cache of memory cache and hard disk cache, while the data processing Bolt does not have an independent memory cache, but directly obtains data from the memory cache in the data receiving Bolt for calculation and processing. The data receiving thread and the data processing thread are independent of each other, data reception adopts a unified cache, and data processing adopts a unified task allocation, which improves the throughput of the stream processing system, the utilization rate of computing resources and storage resources, improves the system stability, and reduces the operation and maintenance costs.

[0065] Embodiment three:

[0066] In order to improve throughput, enhance computing resource utilization, and reduce operation and maintenance costs, this embodiment provides a high-throughput stream processing method.

[0067] This embodiment provides a high-throughput stream processing device, such as Figure 5 The above is a block diagram of a high-throughput stream processing device provided by this embodiment, including a data receiving module and a data processing module;

[0068] The data receiving module is used to receive data and determine whether the remaining cache space of the current data memory cache is greater than a preset threshold, and if so, cache the received data into the memory cache, otherwise, save the received data into the local hard disk;

[0069] The data processing module is used to obtain data from the memory cache and perform calculations on the obtained data.

[0070] The high-throughput stream processing device provided in this embodiment further includes: a creation module, the creation module is used to construct a bolt model; the bolt model includes: a data receiving bolt model, a data processing bolt model, and a data sending bolt model; the data receiving bolt model is used to receive a data message sent by the data sending bolt model, and the data processing bolt model is used to obtain data from the data receiving bolt model and perform calculation processing;

[0071] The creation module is also used to establish a thread pool according to the constructed bolt model, and the thread pool includes: a data receiving thread, a data processing thread, and a data sending thread; the data receiving thread is used to receive data messages sent by the data sending thread, and the data processing thread is used to obtain data from the data receiving thread for calculation and processing;

[0072] A configuration module is used to configure the ratio of the data receiving thread and the data processing thread according to a preset ratio.

[0073] In the Apache storm distributed real-time computing system, the conventional bolt model is a level-by-level bolt followed by a level-by-level bolt to form a complete data receiving and processing link. In the embodiment of the present invention, a bolt model with multiple dimensions is constructed, including a data receiving bolt model, a data processing bolt model, and a data sending bolt model. These three groups of bolt models with independent dimensions, i.e., three independent data links, each link only assumes a single responsibility. The data processing bolt model is used to process data, the data receiving bolt model is used to receive data, and the data sending bolt model is used to send data.

[0074] The data receiving bolt in this embodiment has two levels of storage space, namely, data memory cache space and local hard disk storage space, which are used to store received data. When the data receiving thread receives data, it will determine the remaining cache space size of the current data memory cache. If the remaining space of the current memory cache is greater than the preset threshold, the received data will be cached in the memory cache for caching, otherwise, the received data will be saved in the local hard disk. The data processing thread directly obtains the data to be processed from the memory cache for calculation and processing. In this embodiment, when the remaining memory cache space is less than 5%, the received data can be saved to the local hard disk, or when the memory cache space is full, that is, the remaining cache space is 0, the received data can be saved to the local hard disk. The specific threshold can be set flexibly, and this embodiment does not limit it. The processing thread directly obtains data from the memory cache, and the data processing thread uses a process-level public queue to actively obtain the data to be processed from the memory cache, and continuously calculates and processes the data cached in the memory cache.

[0075] In an example of this embodiment, the ratio of data receiving threads and data processing threads is configured according to a preset ratio. Specifically, 10% of the processing threads can be configured to process data, and the others are used as data receiving threads and data sending threads. The specific ratio can be flexibly configured according to the data processing requirements of the system, and this embodiment does not limit it.

[0076] After the data receiving module groups the received data according to the preset rules, it determines whether the remaining cache space of the current data memory cache is greater than the preset threshold. If so, the received data is inserted into the data list of the memory cache, otherwise, the received data is saved to the local hard disk. The data receiving module dynamically calculates the time required to process the received data, and dynamically adjusts the granularity of the data cache according to the calculated data processing time. The data receiving module scans the data cache status in the local hard disk and the memory cache in real time. If there is data in the local hard disk and the remaining cache space of the memory cache is greater than the preset threshold, the data in the local hard disk is cached in the memory cache.

[0077] In this embodiment, the data processing module obtains data from the memory cache, and performs calculation processing on the obtained data, including: the data processing module obtains a group of tuple data from the memory cache for calculation processing according to a preset time interval;

[0078] If the data processing module fails to obtain data from the memory cache, it enters a sleep state for a preset time.

[0079] This embodiment provides a high-throughput stream processing device, including: a data receiving module, a data processing module;

[0080] The data receiving module is used to receive data and determine whether the remaining cache space of the current data memory cache is greater than a preset threshold. If so, the received data is cached in the memory cache, otherwise, the received data is saved in the local hard disk; the data processing module is used to obtain data from the memory cache and perform calculations on the obtained data; the data receiving Bolt adopts a two-level cache of memory cache and hard disk cache, while the data processing Bolt does not have an independent memory cache, but directly obtains data from the memory cache in the data receiving Bolt for calculations. The data receiving thread and the data processing thread are independent of each other, the data reception adopts a unified cache, and the data processing adopts a unified task allocation, which improves the throughput of the stream processing system, the utilization rate of computing resources and storage resources, improves the system stability, and reduces the operation and maintenance costs.

[0081] It can be seen that those skilled in the art should understand that all or some of the steps, systems, and functional modules / units in the above disclosed methods can be implemented as software (which can be implemented with computer program code executable by a computing device), firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be performed by several physical components in cooperation. Some or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit.

[0082] In addition, it is well known to those skilled in the art that communication media generally contain computer readable instructions, data structures, computer program modules or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery media. Therefore, the present invention is not limited to any specific hardware and software combination.

[0083] The above contents are further detailed descriptions of the embodiments of the present invention in combination with specific implementation methods, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A high throughput stream processing method, include: The data receiving thread determines whether the remaining cache space of the current data memory cache is greater than a preset threshold, and if so, caches the received data into the memory cache, otherwise, saves the received data to the local hard disk; The data processing thread obtains data from the memory cache and performs calculation processing on the obtained data, wherein the data processing thread does not set a memory cache; If the data processing thread fails to obtain data from the memory cache, it enters a sleep state for a preset time.

2. The high-throughput stream processing method according to claim 1, It is characterized in that Also includes: Constructing a bolt model, the bolt model includes: a data receiving bolt model, a data processing bolt model, and a data sending bolt model; the data receiving bolt model is used to receive a data message sent by the data sending bolt model, and the data processing bolt model is used to obtain data from the data receiving bolt model and perform calculation processing; A thread pool is established according to the constructed bolt model, and the thread pool includes: a data receiving thread, a data processing thread, and a data sending thread; the data receiving thread is used to receive data messages sent by the data sending thread, and the data processing thread is used to obtain data from the data receiving thread for calculation and processing; The ratio of the data receiving thread to the data processing thread is configured according to a preset ratio.

3. The high throughput stream processing method according to claim 2, It is characterized in that The data sending thread obtains data from the real-time queue using spout, and sends the obtained data to the data receiving thread according to the rules.

4. The high throughput stream processing method according to claim 3, It is characterized in that After the data receiving thread groups the received data according to the preset rules, it determines whether the remaining cache space of the current data memory cache is greater than the preset threshold. If so, the received data is inserted into the data list of the memory cache; otherwise, the received data is saved to the local hard disk.

5. The high throughput stream processing method according to claim 4, It is characterized in that The data receiving thread dynamically calculates the time required to process the received data, and dynamically adjusts the granularity of the data cache according to the calculated data processing time.

6. The high throughput stream processing method according to claim 5, It is characterized in that The data receiving thread scans the data cache status of the local hard disk and the memory cache in real time. If there is data in the local hard disk and the remaining cache space of the memory cache is greater than a preset threshold, the data in the local hard disk is cached in the memory cache.

7. The high-throughput stream processing method according to any one of claims 1 to 6, It is characterized in that The data processing thread obtains data from the memory cache, and performs computational processing on the obtained data, including: The data processing thread obtains a set of tuple data from the memory cache for calculation and processing according to a preset time interval.

8. A high throughput stream processing device, include: Data receiving module, data processing module; The data receiving module is used to receive data and determine whether the remaining cache space of the current data memory cache is greater than a preset threshold, and if so, cache the received data into the memory cache, otherwise, save the received data into the local hard disk; The data processing module is used to obtain data from the memory cache and perform calculations on the obtained data, wherein the data processing thread does not set a memory cache; If the data processing module fails to obtain data from the memory cache, it enters a sleep state for a preset time.

9. The high-throughput stream processing apparatus according to claim 8, It is characterized in that Also includes: A creation module, the creation module is used to build a bolt model; the bolt model includes: a data receiving bolt model, a data processing bolt model, and a data sending bolt model; the data receiving bolt model is used to receive a data message sent by the data sending bolt model, and the data processing bolt model is used to obtain data from the data receiving bolt model and perform calculation processing; The creation module is also used to establish a thread pool according to the constructed bolt model, and the thread pool includes: a data receiving thread, a data processing thread, and a data sending thread; the data receiving thread is used to receive data messages sent by the data sending thread, and the data processing thread is used to obtain data from the data receiving thread for calculation and processing; A configuration module is used to configure the ratio of the data receiving thread and the data processing thread according to a preset ratio.

10. The high-throughput stream processing apparatus according to claim 9, It is characterized in that The data sending thread obtains data from the real-time queue using spout, and sends the obtained data to the data receiving thread according to the rules.

11. The high-throughput stream processing device according to claim 10, It is characterized in that After the data receiving module groups the received data according to preset rules, it determines whether the remaining cache space of the current data memory cache is greater than a preset threshold. If so, the received data is inserted into the data list of the memory cache; otherwise, the received data is saved to the local hard disk.

12. The high-throughput stream processing apparatus according to claim 11, It is characterized in that The data receiving module dynamically calculates the time required to process the received data, and dynamically adjusts the granularity of the data cache according to the calculated data processing time.

13. The high-throughput stream processing apparatus according to claim 12, It is characterized in that The data receiving module scans the data cache status of the local hard disk and the memory cache in real time. If there is data in the local hard disk and the remaining cache space of the memory cache is greater than a preset threshold, the data in the local hard disk is cached in the memory cache.

14. The high-throughput stream processing device according to any one of claims 8 to 13, It is characterized in that The data processing module obtains data from the memory cache, and performs calculation processing on the obtained data, including: The data processing module obtains a set of tuple data from the memory cache for calculation and processing according to a preset time interval.

Citation Information

Patent Citations

  • Method for processing high throughput data stream

    CN106850740A

  • Parallel processing framework supporting large-scale dynamic data query and design method

    CN107807983A