Streaming-based real-time distributed big data acquisition method and system

Through the combination of distributed cloud clustering and memory model, efficient and scalable data acquisition and processing are achieved, solving the problems of insufficient processing performance and data loss in traditional methods, ensuring real-time and security of data.

CN120386772APending Publication Date: 2025-07-29SUZHOU COLLABORATIVE INNOVATION INTELLIGENT MFG EQUIP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510461984.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Traditional data acquisition methods have insufficient processing performance and poor scalability, resulting in data accumulation and loss.

Method used

The distributed cloud cluster is used to process data acquisition tasks, detect data changes in real time, store incremental data using memory models, process data flows in parallel, and use the parallel processing capabilities of the cloud cluster to encrypt and compress data.

Benefits of technology

Improves the processing performance and scalability of data acquisition, reduces data accumulation and loss, and ensures real-time and security of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386772A_ABST
    Figure CN120386772A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a streaming-based real-time distributed big data acquisition method, which comprises the following steps: S1, constructing a partition association task queue, processing a data acquisition task by adopting a distributed cloud cluster mode, and when the data volume is increased or the processing demand is improved, executing the step S2; the processing capacity of the system is expanded by increasing the number of nodes; the invention also provides a streaming-based real-time distributed big data acquisition system, which comprises a system module assembly, and the system module assembly comprises a distributed cloud cluster processing module, a real-time detection module, a memory storage module, a streaming processing module, a parallel computing module, a dynamic node module and a data fragmentation module. The data acquisition task is processed in a distributed cloud cluster mode, the processing performance and expandability of data acquisition are improved, and when the data volume is increased or the processing requirement is improved, the processing capacity of the system can be expanded by increasing the number of nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a method and system for streaming real-time distributed big data collection. Background Art

[0002] With the rapid development of big data technology, real-time data collection and processing have become important requirements in many industries. Big data collection refers to the process of obtaining data from sensors and intelligent devices, enterprise online systems, enterprise offline systems, social networks, and Internet platforms. The data includes various types of structured, semi-structured, and unstructured massive data such as RFID data, sensor data, user behavior data, social network interaction data, and mobile Internet data.

[0003] Traditional data collection methods often directly perform online collection, but they cannot be well processed after collection, suffering from problems such as insufficient processing performance, poor scalability, data accumulation, and data loss. Therefore, a method and system for streaming real-time distributed big data collection are proposed. Summary of the Invention

[0004] (1) Technical Problems to be Solved

[0005] In view of the deficiencies of the prior art, the present invention provides a method and system for streaming real-time distributed big data collection, which solves the problems that traditional data collection methods often directly perform online collection but cannot be well processed after collection, suffering from insufficient processing performance, poor scalability, data accumulation, and data loss.

[0006] (2) Technical Solutions

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A method for streaming real-time distributed big data collection includes the following steps:

[0009] S1: Construct a partition-associated task queue and use a distributed cloud cluster to process data collection tasks. When the data volume increases or the processing requirement improves, expand the processing capacity of the system by increasing the number of nodes;

[0010] S2: Real-time detect changes in business data, detect data changes, and adopt an efficient data capture mechanism without waiting for the data to accumulate to a certain amount before processing, ensuring the real-time and accuracy of the data;

[0011] S3: Store incremental data in a memory model, adopt a memory model to efficiently store incrementally collected data, reduce the space occupied by local temporary file storage, avoid data accumulation and loss, and the incrementally collected data is first stored in the memory to reduce the space occupied by local temporary file storage;

[0012] S4: Whether to perform fluidization processing. If fluidization processing is not required, the data is directly processed. If fluidization processing is required, the data block is fluidized in memory, and the data stream is divided into multiple small blocks for parallel processing;

[0013] S5: Update the analysis data set. Utilize the parallel processing ability of the cloud cluster to perform parallel processing on the data stream, and the processed data is updated to the analysis data set in real time, providing a data basis for subsequent real-time data analysis;

[0014] S6: Data preparation is completed. After data preparation is completed, data encryption and compression are performed. During data transmission, the data is encrypted to ensure data security.

[0015] As a further solution of the present invention, the S1 includes a data collection module. The data collection module collects data before constructing partitions and performs distributed cloud cluster processing after collection.

[0016] Furthermore, the S1 includes a data subscription and push module. Data subscribers register subscription relationships with the data acquisition system. The system assigns corresponding nodes to subscribers according to the subscription relationships. The nodes actively establish connections with data subscribers. After the connection is successful, the data is sent to the subscribers in real time in an active push manner, supporting resume from breakpoint.

[0017] On the basis of the foregoing solution, the data capture mechanism in S2 is specifically as follows:

[0018] Step 1: Change data capture, adopt log scanning, and support the incremental capture mode;

[0019] Step 2: Adaptive polling strategy, dynamically adjust the polling frequency, and automatically switch between high-frequency and low-frequency modes according to the data source activity;

[0020] High-frequency mode: Triggered when it is detected that the data update rate is greater than 100 records per second;

[0021] Low-frequency mode: The frequency is reduced when the data is stationary for more than 5 seconds;

[0022] Step 3: Event timestamp alignment, attach a physical clock mark accurate to microseconds to each piece of data.

[0023] Furthermore, the S3 includes a storage module, the storage module is connected to the data collection module, the S4 includes a data processing module, and the data processing module is connected to a data analysis module.

[0024] The present invention also proposes a streaming real-time distributed big data acquisition system, including a system module assembly, wherein the system module assembly includes a distributed cloud cluster processing module, a real-time detection module, a memory storage module, a streaming processing module, a parallel computing module, a dynamic node module, and a data sharding module. When a new node in the dynamic node module is started, the metadata is registered with the Coordinator, and the Gossip protocol is used to spread the node status to ensure that the cluster topology synchronization is completed within 10 seconds. The data sharding module calculates the sharding key based on the data characteristics, and the routing strategy automatically migrates according to the hot data.

[0025] On the basis of the above solution, the system module assembly includes a fault recovery module, which detects node failure and determines that the node is offline after three heartbeat timeouts.

[0026] Furthermore, after the fault recovery module determines that the node is offline, it performs data reconstruction and gives priority to restoring from other copies.

[0027] (III) Beneficial effects

[0028] Compared with the existing technology, the present invention provides a method and system for real-time distributed big data collection based on streaming, which has the following beneficial effects:

[0029] 1. In the present invention, a distributed cloud cluster approach is adopted to process data acquisition tasks, thereby improving the processing performance and scalability of data acquisition. When the amount of data increases or the processing demand increases, the processing capacity of the system can be expanded by increasing the number of nodes.

[0030] 2. In the present invention, a memory model is used to efficiently store incrementally collected data, reduce the space occupied by local temporary files when they are saved, and avoid data accumulation and loss. The incrementally collected data is first stored in the memory to reduce the space occupied by local temporary files when they are saved. The high efficiency of the memory model enables data to be read and written quickly, thereby improving the speed of data processing.

[0031] 3. In the present invention, based on the memory model, the data blocks are streamed, the data streams are directly processed in parallel in the memory and updated to the analysis data set in real time, thereby improving data processing efficiency.

[0032] 4. In the present invention, the parallel processing capability of the cloud cluster is utilized to process data streams in parallel. Multiple nodes can process different data blocks simultaneously, which speeds up data processing. The processed data is updated to the analysis data set in real time, providing a data basis for subsequent real-time data analysis. The real-time feedback mechanism enables the system to adjust the collection strategy and processing method in a timely manner to adapt to the needs of data changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 Schematic diagram of the step process of a streaming real-time distributed big data collection method proposed by the present invention. Specific implementation mode

[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0035] Embodiment 1

[0036] Refer to Figure 1 , a streaming real-time distributed big data collection method, including the following steps:

[0037] S1: Construct a partition association task queue, and use a distributed cloud cluster method to process data collection tasks, improving the processing performance and scalability of data collection. When the data volume increases or the processing requirements increase, the processing capacity of the system can be expanded by increasing the number of nodes. S1 includes a data collection module, which collects data before constructing partitions and performs distributed cloud cluster processing after collection;

[0038] S2: Real-time detect changes in business data, detect data changes, without waiting for the data to accumulate to a certain amount before processing, and adopt an efficient data capture mechanism to ensure the real-time and accuracy of data. The data capture mechanism in S2 is specifically Step 1: Change Data Capture (CDC).

[0039] Adopt log scanning (such as database binlog, Kafka log offset) instead of full table scanning.

[0040] Support incremental capture mode (such as the Debezium framework realizes real-time change monitoring of MySQL / Oracle).

[0041] Step 2: Adaptive polling strategy

[0042] Dynamically adjust the polling frequency: Automatically switch between high-frequency / low-frequency modes according to the data source activity.

[0043] High-frequency mode (such as 100ms interval): Triggered when the detected data update rate > 100 records / second.

[0044] Low-frequency mode (such as 1s interval): Reduce the frequency when the data is stationary for more than 5 seconds.

[0045] Step 3: Event timestamp alignment

[0046] Attach a physical clock mark accurate to microseconds to each piece of data (synchronize the clocks of each node using the PTPv2 protocol);

[0047] S3: The memory model stores incremental data. An efficient memory model is used to store the incrementally collected data, reducing the space occupied by local temporary files during saving, avoiding data accumulation and loss. The incrementally collected data is first stored in memory to reduce the space occupied by local temporary files during saving. The efficiency of the memory model enables fast reading and writing of data, improving the speed of data processing;

[0048] S4: Whether to perform streaming processing. If streaming processing is not required, the data is directly processed. If streaming processing is required, the data blocks are streamed in memory, and the data stream is divided into multiple small blocks for parallel processing. Based on the memory model, the data blocks are streamed, and the data stream is directly processed in parallel in memory and updated in real time to the analysis dataset, improving data processing efficiency;

[0049] S5: Update the analysis dataset. Utilize the parallel processing capabilities of the cloud cluster to perform parallel processing on the data stream. Multiple nodes can simultaneously process different data blocks, accelerating the speed of data processing. The processed data is updated in real time to the analysis dataset, providing a data basis for subsequent real-time data analysis. The real-time feedback mechanism enables the system to promptly adjust the acquisition strategy and processing method to adapt to the changing data requirements;

[0050] S6: Data preparation is completed. After data preparation is completed, data encryption and compression are performed. During data transmission, the data is encrypted to ensure data security.

[0051] The present invention also proposes a streaming real-time distributed big data acquisition system, including a system module assembly. The system module assembly includes a distributed cloud cluster processing module, a real-time detection module, a memory storage module, a streaming processing module, a parallel computing module, a dynamic node module, and a data sharding module. When a new node starts in the dynamic node module, it registers metadata (IP / CPU / memory / disk) with the Coordinator and spreads the node status using the Gossip protocol to ensure that the cluster topology synchronization is completed within 10 seconds. The data sharding module calculates the sharding key based on data characteristics, and the routing policy automatically migrates according to hot data: when the load of a certain node > 80% lasts for 5 minutes, data rebalancing is triggered, and cross-data center disaster recovery: 3 replicas are maintained through the Raft algorithm, allowing one data center to go down.

[0052] Embodiment 2

[0053] Refer to Figure 1 , a streaming real-time distributed big data acquisition method, including the following steps:

[0054] S1: Build a partition association task queue and use a distributed cloud cluster to process data collection tasks, improving the processing performance and scalability of data collection. When the data volume increases or the processing requirements rise, the system's processing capacity can be expanded by adding the number of nodes. S1 includes a data collection module that collects data before building partitions and performs distributed cloud cluster processing after collection. S1 also includes a data subscription and push module. Data subscribers register subscription relationships with the data collection system. The system assigns corresponding nodes to subscribers based on the subscription relationships. The nodes actively establish connections with the data subscribers. After successful connection, data is sent to the subscribers in real time using the active push method, supporting resume from breakpoint to ensure the integrity and reliability of data during transmission.

[0055] S2: Detect changes in business data in real time. Detect data changes without waiting for the data to accumulate to a certain amount before processing. Adopt an efficient data capture mechanism to ensure the real-time and accuracy of data. The data capture mechanism in S2 is specifically as follows: Step 1: Change Data Capture (CDC).

[0056] Adopt log scanning (such as database binlog, Kafka log offset) instead of full table scanning.

[0057] Support the incremental scraping mode (such as the Debezium framework to achieve real-time change monitoring of MySQL / Oracle).

[0058] Step 2: Adaptive polling strategy

[0059] Dynamically adjust the polling frequency: Automatically switch between high-frequency / low-frequency modes according to the data source activity.

[0060] High-frequency mode (such as at 100ms intervals): Triggered when the detected data update rate > 100 records / second.

[0061] Low-frequency mode (such as at 1s intervals): Reduce the frequency when the data is stationary for more than 5 seconds.

[0062] Step 3: Event timestamp alignment

[0063] Attach a physical clock mark accurate to microseconds to each piece of data (synchronize the clocks of each node using the PTPv2 protocol);

[0064] S3: The memory model stores incremental data. An in-memory model is adopted to efficiently store the incrementally collected data, reducing the space occupied by local temporary files during saving, avoiding data accumulation and loss. The incrementally collected data is first stored in memory to reduce the space occupied by local temporary files during saving. The high efficiency of the in-memory model enables fast reading and writing of data, improving the speed of data processing. S3 includes a storage module, which is connected to the data collection module;

[0065] S4: Whether to perform streaming processing. If streaming processing is not required, the data is directly processed. If streaming processing is required, the data blocks are streamed in memory, and the data stream is divided into multiple small blocks for parallel processing. Based on the in-memory model, the data blocks are streamed, and the data stream is directly processed in parallel in memory and updated in real time to the analysis dataset, improving the data processing efficiency. S4 includes a data processing module, which is connected to a data analysis module;

[0066] S5: Update the analysis dataset. Utilize the parallel processing ability of the cloud cluster to perform parallel processing on the data stream. Multiple nodes can process different data blocks simultaneously, accelerating the data processing speed. The processed data is updated in real time to the analysis dataset, providing a data basis for subsequent real-time data analysis. The real-time feedback mechanism enables the system to timely adjust the acquisition strategy and processing method to adapt to the requirements of data changes;

[0067] S6: Data preparation is completed. After data preparation is completed, data encryption and compression are performed. During data transmission, the data is encrypted to ensure data security.

[0068] The present invention also proposes a streaming real-time distributed big data acquisition system, including a system module assembly. The system module assembly includes a distributed cloud cluster processing module, a real-time detection module, a memory storage module, a streaming processing module, a parallel computing module, a dynamic node module, and a data sharding module. When a new node starts in the dynamic node module, it registers metadata (IP / CPU / memory / disk) with the Coordinator, and uses the Gossip protocol to spread the node status to ensure that the cluster topology synchronization is completed within 10 seconds. The data sharding module calculates the sharding key based on data characteristics, and the routing policy automatically migrates according to hot data: when the load of a certain node > 80% lasts for 5 minutes, data rebalancing is triggered, and cross-data center disaster recovery: 3 replicas are maintained through the Raft algorithm, allowing one data center to go down.

[0069] In particular, the system module assembly includes a fault recovery module. The fault recovery module detects node failures. Three consecutive heartbeat timeouts (with an interval of 2 seconds) determine that the node is offline. After the fault recovery module determines that the node is offline, data reconstruction is performed, and recovery is preferentially from other replicas.

[0070] In the description herein, it should be noted that relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises", "comprising", or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0071] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the invention, and the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for streaming real-time distributed big data collection, characterized in that, It includes the following steps: S1: Construct a partition association task queue, and use a distributed cloud cluster to process data collection tasks. When the data volume increases or the processing requirements improve, expand the processing capacity of the system by increasing the number of nodes; S2: Detect changes in business data in real time. Detect data changes and adopt an efficient data capture mechanism. There is no need to wait for the data to accumulate to a certain amount before processing, ensuring the real-time and accuracy of the data; S3: Store incremental data in a memory model. Use a memory model to efficiently store the incrementally collected data, reduce the space occupied by local temporary files during saving, and avoid data accumulation and loss. The incrementally collected data is first stored in memory to reduce the space occupied by local temporary files during saving; S4: Whether to perform streaming processing. If streaming processing is not required, directly process the data. If streaming processing is required, perform streaming processing on the data blocks in memory, and divide the data stream into multiple small blocks for parallel processing; S5: Update the analysis data set. Utilize the parallel processing ability of the cloud cluster to perform parallel processing on the data stream, and the processed data is updated to the analysis data set in real time, providing a data basis for subsequent real-time data analysis; S6: Data preparation is completed. After data preparation is completed, data encryption and compression are performed. During data transmission, the data is encrypted to ensure data security.

2. The method for streaming real-time distributed big data collection according to claim 1, wherein, The S1 includes a data collection module. The data collection module collects data before constructing partitions and performs distributed cloud cluster processing after collection.

3. A real-time distributed big data acquisition method based on streaming according to claim 1, characterized in that The S1 includes a data subscription and push module. Data subscribers register subscription relationships with the data collection system. The system assigns corresponding nodes to subscribers according to the subscription relationships. The nodes actively establish connections with data subscribers. After the connection is successful, the data is sent to the subscribers in real time in an active push manner, supporting resume from breakpoint.

4. A method for streaming real-time distributed big data collection according to claim 1, characterized in that, The data capture mechanism in the S2 is specifically as follows: Step 1: Change data capture. Adopt log scanning and support the incremental capture mode; Step 2: Adaptive polling strategy. Dynamically adjust the polling frequency and automatically switch between high-frequency and low-frequency modes according to the data source activity; High-frequency mode: Triggered when the detected data update rate is greater than 100 records per second; Low-frequency mode: Reduce the frequency when the data is stationary for more than 5 seconds; Step 3: Event timestamp alignment. Attach a physical clock mark accurate to microseconds to each piece of data.

5. A real-time distributed big data acquisition method based on streaming according to claim 1, characterized in that The S3 includes a storage module. The storage module is connected to the data collection module. The S4 includes a data processing module. The data processing module is connected to a data analysis module.

6. A streaming real-time distributed big data acquisition system, characterized in that, It includes a system module assembly. The system module assembly includes a distributed cloud cluster processing module, a real-time detection module, a memory storage module, a streaming processing module, a parallel computing module, a dynamic node module, and a data sharding module. When a new node starts in the dynamic node module, it registers metadata with the Coordinator and spreads the node status using the Gossip protocol to ensure that the cluster topology synchronization is completed within 10 seconds. The data sharding module calculates the sharding key based on data characteristics, and the routing policy automatically migrates according to hot data.

7. A real-time distributed big data acquisition system based on streaming, according to claim 6, characterized in that The system module assembly includes a fault recovery module. The fault recovery module detects node failures and determines that a node is offline after 3 heartbeat timeouts.

8. A method for streaming real-time distributed big data collection according to claim 6, characterized in that, After the fault recovery module determines that a node is offline, it performs data reconstruction and preferentially recovers from other replicas.