A data stream connection method, system, and storage medium

By creating a data queue for each data stream in the cache structure and connecting based on timestamp information, the problems of delay and fixed window limitations in data stream processing are solved, and efficient, flexible and accurate data processing is achieved.

CN116881286BActive Publication Date: 2025-07-25DOLPHINDB INC (CN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310907059.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-22
Publication Date
2025-07-25
Estimated Expiration
2043-07-22

AI Technical Summary

Technical Problem

In scenarios where real-time requirements are high and complex connection logic are required, existing data stream processing methods may lead to computational delay and performance problems, especially when data stream speeds are inconsistent or delays, new data records cannot be responded to in a timely manner, and the fixed window size cannot adapt to data changes.

Method used

A corresponding data queue is created for each original data stream using a cache structure, data is stored according to the characteristic information of the data, and data connection is performed based on the timestamp information, and two connection modes are set to improve flexibility and efficiency, including the first mode and the second mode. The first mode determines the connection through the time range, and the second mode triggers the connection through the local system time stamp.

Benefits of technology

It improves the efficiency and accuracy of data processing, reduces data conflicts, ensures the timing and real-time nature of data processing, and optimizes system performance and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116881286B_ABST
    Figure CN116881286B_ABST
Patent Text Reader

Abstract

This application relates to the field of data processing, and discloses a data stream connection method, system, and storage medium. A data stream connection method includes: accepting a plurality of original data streams; creating a plurality of cache structures, and establishing a plurality of data queues in each cache structure; storing data in corresponding cache structures; for each cache structure, connecting the data in the main queue with the data in at least one slave queue based on the timestamp information of the data in the plurality of data queues to generate and output a target data stream. This application uses cache structures to temporarily store data, making data processing smoother, avoiding connection delays caused by too fast inflow of instantaneous data. At the same time, corresponding data queues are created for each original data stream in the cache structures, and the data of the same original data stream can be independently processed in multiple cache structures, realizing parallel processing of multiple data and improving the overall data processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular, to a data stream connection method, system, and storage medium. Background Art

[0002] Batch data refers to data processed in a batch processing manner. Batch processing is a traditional data processing method that divides data into batches of jobs and processes the data of each batch sequentially. Streaming data refers to continuously generated and dynamically flowing data. Different from traditional batch data, streaming data is continuously generated, has no clear boundaries and fixed sizes. It usually flows through the system in a time series manner, and the data is processed or transmitted soon after it is generated. Stream computing is a real-time data processing technology used for real-time processing and analysis of continuously generated streaming data.

[0003] Apache Flink is a distributed open-source computing framework that supports both streaming data processing and batch data processing. In Flink, there are two common join operations: window join and interval join.

[0004] Window join is an operation that joins data within a fixed window. It requires the following parameters: two streams (stream and otherStream) to be joined, a selector for specifying the join condition (such as which attributes in the stream are used for joining), a window, and a join function.

[0005] Specifically, selectors for specifying the join condition such as the where selector and the equalTo selector. The where selector is used to specify the join condition between the left table (stream) and the right table (otherStream) in the join operation, defining the filtering condition in the join operation. Only data records that meet the conditions will be joined. The equalTo selector is a special case of the where selector, used to specify the equality condition in the join operation. It specifies which attributes in the two data streams should be compared for equality. Only when the specified attributes have the same values in the left table and the right table will they be considered matching, and then the join operation is performed.

[0006] A window refers to a mechanism for segmenting or chunking data streams. It divides an infinite data stream into finite data sets for bounded operations and analysis on this data. According to the way of window division, windows can be mainly divided into two types: time window (Time Window) and count window (Count Window). The time window (Time Window) divides the data stream according to the time range. It defines a fixed time interval or window length, such as every 5 seconds, every minute, or every hour. Events in the data stream are assigned to the corresponding time windows according to the timestamps of the events. The time window (Time Window) can be rolling (fixed length and non-overlapping) or sliding (overlapping). The count window (Count Window) divides the data stream according to the quantity of data. It defines a fixed data volume threshold, such as every 10, every 100, or every 1000. When the specified data volume is reached, the window is triggered and the data in the window is processed.

[0007] The join function is used to specify the logic of the join operation and defines how to join and merge two data streams.

[0008] Interval Join: It is an operation that performs join processing on two stream data with the same key within a specified time interval. For each element in one stream, it looks for matching elements in the other stream within the specified time interval and associates them.

[0009] However, in some cases, the above join operations may have a relatively high computational latency, mainly because the above join operations need to wait for matching records in both streams to arrive before performing the join operation. If the data in the data stream arrives at inconsistent speeds or with delays, it will lead to an increase in the waiting time of the join operation. And when performing a join within an interval, only two stream data with the same key can be supported for joining. If multiple-key stream data needs to be supported, it will greatly increase the complexity of the join process. Furthermore, the static window join operation in Flink cannot trigger the join in real time but performs the join at the boundaries of the window. This may result in a certain delay between the join operation and the actual data arrival, being unable to respond immediately to new data records. And since the size of the window is fixed and cannot be adjusted dynamically, it may not be able to adapt to changes in data and actual requirements. If the characteristics of the data change or window adjustment is required according to dynamic situations, this limitation of the fixed window size may lead to performance and accuracy problems.

[0010] Therefore, the above data stream processing methods may affect the performance, real-time nature, and flexibility of the application program in scenarios with high requirements for real-time nature and complex join logic to be processed. Summary of the Invention

[0011] To address the above problems, the present application provides a data stream connection method, system, and storage medium.

[0012] In a first aspect, a data stream connection method provided by the present application adopts the following technical solution:

[0013] A data stream connection method, the method comprising:

[0014] Receiving a plurality of original data streams, and continuously inputting data in the original data streams according to a time series;

[0015] Creating a plurality of cache structures, and respectively establishing a plurality of data queues corresponding to the plurality of original data streams in each of the cache structures, the plurality of data queues including a main queue and at least one slave queue;

[0016] Storing the data into the corresponding cache structure according to the characteristic information of each data, and for each cache structure, storing the data from different original data streams into the corresponding data queues in sequence;

[0017] For each cache structure, connecting the data in the main queue with the data in the at least one slave queue based on the timestamp information of the data in the plurality of data queues to generate and output a target data stream.

[0018] By adopting the above technical solution, the present application can achieve the following functions: parallel processing. By creating corresponding data queues for each original data stream in the cache structure, multiple data streams can be processed in parallel, thereby improving the overall data processing speed. Further, by storing the data into the corresponding cache structure according to the characteristic information of the data, different data of the same original data stream can be independently processed in multiple cache structures, which can further improve the data processing efficiency and reduce the conflict of data with different characteristic information, and improve the accuracy of data processing.

[0019] In summary, the efficiency of stream data processing is improved.

[0020] Exemplarily, the connecting the data in the main queue with the data in the at least one slave queue based on the timestamp information of the data in the plurality of data queues includes:

[0021] Determining a data connection mode according to preset configuration information, and connecting the data in the main queue with the data in the at least one slave queue according to the determined data connection mode;

[0022] Among them, the data connection mode includes a first mode and a second mode. In the first mode, the data to be connected in the slave queue of the cache structure is determined by the timestamp t of the earliest unconnected data in the master queue and a preset time range [a, b]. In the second mode, the data to be connected in the slave queue of the cache structure is determined by the timestamp t of the earliest unconnected data in the master queue and the timestamp t0 of the previously stored data.

[0023] By adopting the above technical solution, two data connection modes are set, so that the most suitable connection mode can be selected according to the actual situation, increasing the flexibility of data processing.

[0024] In the first mode, the data to be connected is determined by the timestamp of the earliest unconnected data in the master queue and the preset time range, which can ensure that only the data within the specified time range is considered, thus avoiding unnecessary incorrect matching. And in the first mode, the amount of data to be processed can be controlled by adjusting the preset time range, so as to optimize the performance of data processing while ensuring the accuracy of data processing.

[0025] In the second mode, the data to be connected is determined by the timestamp of the earliest unconnected data in the master queue and the timestamp of the previously stored data, which can dynamically respond to the changes in the data stream and adjust the connected data in real time, thus providing more timely and accurate results. Since two timestamps are used to determine the data to be connected, the relevant data can be more accurately matched, reducing the inaccurate connection caused by timestamp errors. And because only the earliest unconnected data and the previous data in the master queue need to be considered, the complexity and quantity of data processing can be effectively reduced, thus optimizing the system performance.

[0026] Optionally, the connecting the data in the master queue with the data in the at least one slave queue based on the timestamp information of the data in the multiple data queues includes:

[0027] Performing a preset aggregation operation on the data in some or all of the slave queues based on the timestamp information of the data in the multiple slave queues to obtain aggregated data;

[0028] Connecting the data in the master queue with the aggregated data based on the timestamp information of the data in the multiple data queues, or connecting the data in the master queue with the aggregated data and the data in the at least one slave queue based on the timestamp information of the data in the multiple data queues.

[0029] By adopting the above technical solution, the aggregation operation can calculate the data in the slave queue in advance when adding data to the slave queue. For example, when the preset aggregation operation is a summation operation, the sum can be directly calculated when new data in the slave queue triggers the summation operation, and the aggregation result of the aggregation operation can be directly retrieved during subsequent connection, thereby reducing the complexity of data processing, reducing the pressure of data processing, and improving the efficiency of data processing. By connecting the data in the main queue with the aggregated data, or the data in the main queue with the aggregated data and the data in at least one slave queue, the most suitable processing method can be selected according to specific requirements, improving the flexibility of data processing. Further, connecting based on the timestamp information can ensure the timeliness of the data, which is very useful for some scenarios that need to process real-time data or ensure the order of data.

[0030] In summary, this method of aggregating and connecting based on timestamp information can not only simplify data processing, reduce data redundancy, and improve the efficiency of data processing, but also ensure the timeliness of the data.

[0031] Exemplarily, the multiple original data streams only include a first data stream and a second data stream, and each of the cache structures includes a main queue and a slave queue. The main queue corresponds to the first data stream, and the slave queue corresponds to the second data stream.

[0032] Optionally, when the determined data connection mode is the first mode, the connecting of the data in the main queue with the data in the at least one slave queue according to the determined data connection mode includes:

[0033] Obtain the timestamp t of the earliest unconnected data existing in the main queue;

[0034] Query the timestamp t1 of the latest deposited data in the slave queue;

[0035] When the timestamp t1 of the latest deposited data in the slave queue is greater than or equal to the sum of the timestamp t of the earliest unconnected data existing in the main queue and the left boundary a of the preset time range [a, b], trigger the connection of the data in the slave queue within the time range [t + a, t + b] with the data with the timestamp t in the main queue, and output the result when the timestamp t1 of the latest deposited data in the slave queue is greater than or equal to the sum of the timestamp t of the earliest unconnected data existing in the main queue and the right boundary b of the preset time range [a, b].

[0036] By adopting the above technical solution, by comparing the timestamps of the main queue and the slave queue, it is possible to accurately determine when to perform data connection. This method can not only ensure the timeliness of the data but also avoid invalid connection operations; only when the timestamp t1 of the latest data stored in the slave queue is greater than or equal to the sum of the timestamp t of the earliest unconnected data existing in the main queue and the left boundary a of the preset time range [a, b], will the connection operation be triggered. This can avoid invalid connection operations caused by unready data. By triggering the connection operation within the time range [t + a, t + b], the amount of data to be processed can be limited, and all relevant data within this time range can be effectively processed.

[0037] The above triggering method not only ensures the timeliness of data processing but also effectively avoids invalid operations.

[0038] Optionally, when the determined data connection mode is the second mode, the connecting the data in the main queue with the data in the at least one slave queue according to the determined data connection mode includes:

[0039] Determine whether to trigger the data connection operation according to the local system timestamp according to the configuration information;

[0040] When it is determined to trigger the data connection operation according to the local system timestamp:

[0041] As long as new data is stored in the main queue, trigger the connection of the data in the slave queue within the time range [t0, t] with the data with the timestamp t in the main queue;

[0042] Wherein, t is the timestamp of the latest data stored in the main queue, and t0 is the timestamp of the previous data stored in the main queue;

[0043] When it is determined not to trigger the data connection operation according to the local system timestamp:

[0044] Obtain the timestamp t of the earliest unconnected data existing in the main queue;

[0045] Query the timestamp t1 of the latest data stored in the slave queue;

[0046] When the timestamp t1 of the latest data stored in the slave queue is greater than or equal to the timestamp t of the earliest unconnected data existing in the main queue, trigger the connection of the data in the slave queue within the time range [t0, t] with the data with the timestamp t in the main queue;

[0047] Wherein, t0 is the timestamp of the previous data stored in the main queue.

[0048] By adopting the above technical solution, when triggering the data connection operation based on the local system timestamp, once new data is stored in the main queue, it will trigger the connection of the data within the time range [t0, t] in the slave queue and the corresponding aggregated data with the data at the timestamp t of the main queue. This mechanism is more proactive and real-time. As long as new data enters the main queue, it will attempt to perform the data connection operation, and can immediately respond when new data arrives, providing good real-time performance.

[0049] When not triggering the data connection operation based on the local system timestamp, obtain the timestamp t of the earliest unconnected data existing in the main queue, and then query the timestamp t1 of the latest stored data in the slave queue. When the timestamp t1 of the latest stored data in the slave queue is greater than or equal to the timestamp t of the earliest unconnected data existing in the main queue, it will trigger the connection of the data within the time range [t0, t] in the slave queue and the corresponding aggregated data with the data at the timestamp t of the main queue. This mechanism is more conservative. It needs to ensure that the data in the slave queue is ready before performing the connection operation, which can avoid premature connection operations and thus reduce the possible computational load.

[0050] Both of these two methods ensure the real-time performance and accuracy of data processing, making the connection of the data stream more flexible and efficient.

[0051] Optionally, when the determined data connection mode is the first mode, the connecting the data in the main queue with the data in the at least one slave queue according to the determined data connection mode includes:

[0052] Obtain the timestamp t of the earliest unconnected data existing in the main queue;

[0053] Query the timestamp t1 of the latest stored data in the slave queue;

[0054] When the timestamp t1 of the latest stored data in the slave queue is greater than or equal to the sum of the timestamp t of the earliest unconnected data existing in the main queue and the left boundary a of the preset time range [a, b], trigger the connection of the aggregated data within the time range [t + a, t + b] in the slave queue with the data at the timestamp t of the main queue, or trigger the connection of the data within the time range [t + a, t + b] in the slave queue and the corresponding aggregated data with the data at the timestamp t of the main queue, and output the result when the timestamp t1 of the latest stored data in the slave queue is greater than or equal to the sum of the timestamp t of the earliest unconnected data existing in the main queue and the right boundary b of the preset time range [a, b].

[0055] Optionally, when the determined data connection mode is the second mode, the connecting of the data in the main queue and the data in the at least one slave queue according to the determined data connection mode includes:

[0056] Determine whether to trigger a data connection operation according to the local system timestamp according to the configuration information;

[0057] When it is determined to trigger a data connection operation according to the local system timestamp:

[0058] Whenever new data is stored in the main queue, trigger the connection of the aggregated data in the slave queue within the time range [t0, t] and the data with the timestamp t in the main queue, or trigger the connection of the data within the time range [t0, t] in the slave queue and the corresponding aggregated data and the data with the timestamp t in the main queue;

[0059] Wherein, t is the timestamp of the latest stored data in the main queue, and t0 is the timestamp of the previous stored data in the main queue;

[0060] When it is determined not to trigger a data connection operation according to the local system timestamp:

[0061] Obtain the timestamp t of the earliest unconnected data existing in the main queue;

[0062] Query the timestamp t1 of the latest stored data in the slave queue;

[0063] When the timestamp t1 of the latest stored data in the slave queue is greater than or equal to the timestamp t of the earliest unconnected data existing in the main queue, trigger the connection of the aggregated data in the slave queue within the time range [t0, t] and the data with the timestamp t in the main queue, or trigger the connection of the data within the time range [t0, t] in the slave queue and the corresponding aggregated data and the data with the timestamp t in the main queue;

[0064] Wherein, t0 is the timestamp of the previous stored data in the main queue.

[0065] By adopting the above technical solution, when connecting the main queue and the slave queue, the slave queue has completed the aggregation operation first. In this way, when connecting, the amount of data to be calculated will be greatly reduced, the complexity of the connection can be reduced, and thus the computing resources and time can be saved, and the data processing efficiency can be improved.

[0066] In a second aspect, a processing system for streaming data provided by the present application adopts the following technical solution:

[0067] A processing system for streaming data, comprising:

[0068] A data acquisition module, configured to receive a plurality of original data streams, and the data in the original data streams is continuously input according to a time series; a data caching module, configured to create a plurality of caching structures, and respectively establish a plurality of data queues in one-to-one correspondence with the plurality of original data streams in each of the caching structures, where the plurality of data queues include a main queue and at least one slave queue;

[0069] A data distribution module, configured to distribute the data to the corresponding caching structures according to the feature information of each piece of data, and for each of the caching structures, store the data from different original data streams into the corresponding data queues in sequence; a data connection module, configured to, for each of the caching structures, connect the data in the main queue with the data in the at least one slave queue based on the timestamp information of the data in the plurality of data queues, so as to generate and output a target data stream.

[0070] By adopting the above technical solution, corresponding data queues can be created for each original data stream in the caching structure, so that multiple data streams can be processed in parallel, improving the overall data processing speed. The data is distributed to the corresponding caching structures according to the feature information of the data, and the data of the same original data stream can be independently processed in multiple caching structures, further improving the data processing efficiency, reducing the conflict of data with different feature information, and improving the data processing accuracy. Since the original data streams are continuously input, directly processing may cause a problem that the processing capacity does not match the data input rate, especially when the data input rate suddenly increases. Using the caching structure can temporarily store the data and wait for processing, making the data processing smoother, preventing the system from being pressured by the instantaneous peak of data processing, and avoiding connection delays caused by too fast inflow of instantaneous data.

[0071] In a third aspect, a computer-readable storage medium provided by the present application adopts the following technical solution:

[0072] A computer-readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement: the above data stream connection method.

[0073] In summary, the present application includes at least one of the following beneficial technical effects:

[0074] 1. By creating corresponding data queues for each original data stream in the cache structure, multiple data streams can be processed in parallel, thereby improving the overall data processing speed. According to the characteristic information of the data, the data is stored in the corresponding cache structure, so that different data of the same original data stream can be independently processed in multiple cache structures, which can further improve the data processing efficiency, reduce the data conflict of the characteristic information, and improve the accuracy of data processing.

[0075] 2. The data is temporarily stored through the cache structure, making the data processing smoother, preventing the pressure on the system caused by the instantaneous peak of data processing, and avoiding connection delays caused by the too-fast inflow of instantaneous data. Description of the Drawings

[0076] Figure 1 It is a step diagram of a data stream connection method.

[0077] Figure 2 It is a structural diagram of a data stream connection system. Detailed Implementation Modes

[0078] The following further describes the present application in detail with reference to the drawings. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0079] An embodiment of the present application discloses a data stream connection method. Referring to Figure 1 , a data stream connection method includes:

[0080] S1. Accept multiple original data streams;

[0081] Specifically, the data in the original data stream is continuously input according to the time series;

[0082] S2. Create multiple cache structures, and respectively establish multiple data queues corresponding to the multiple original data streams in each cache structure;

[0083] Specifically, the multiple data queues include a main queue and at least one slave queue;

[0084] S3. According to the characteristic information of each data, the data is stored in the corresponding cache structure, and for each cache structure, the data from different original data streams is respectively stored in the corresponding data queue in sequence;

[0085] S4. For each cache structure, based on the timestamp information of the data in the multiple data queues, connect the data in the main queue with the data in at least one slave queue to generate and output the target data stream.

[0086] Specifically, the above step S4 of connecting the data in the main queue with the data in at least one slave queue based on the timestamp information of the data in multiple data queues includes:

[0087] Determine the data connection mode according to the preset configuration information, and connect the data in the main queue with the data in at least one slave queue according to the determined data connection mode;

[0088] Among them, the preset configuration information is set by the user before data processing. The data connection mode includes a first mode and a second mode. In the first mode, the data to be connected in the slave queue in the cache structure is determined by the timestamp t of the earliest unconnected data existing in the main queue and the preset time range [a, b]. In the second mode, the data to be connected in the slave queue in the cache structure is determined by the timestamp t of the earliest unconnected data existing in the main queue and the timestamp t0 of the previously stored data.

[0089] Furthermore, the connection methods of the main queue and the slave queue are different in the first mode and the second mode. By way of example, the original data stream includes a first data stream and a second data stream. Then each cache structure includes a main queue and a slave queue. The main queue corresponds to the first data stream, and the slave queue corresponds to the second data stream.

[0090] In the first embodiment of this application, when the determined data connection mode is the first mode, connecting the data in the main queue with the data in at least one slave queue according to the determined data connection mode includes:

[0091] Obtain the timestamp t of the earliest unconnected data existing in the main queue;

[0092] Query the timestamp t1 of the latest stored data in the slave queue;

[0093] When the timestamp t1 of the latest stored data in the slave queue is greater than or equal to;

[0094] Trigger the connection of the data in the slave queue within the time range [t + a, t + b] with the data with the timestamp t in the main queue when the timestamp t1 of the latest stored data in the slave queue is greater than or equal to the sum of the timestamp t of the earliest unconnected data existing in the main queue and the left boundary a of the preset time range [a, b], and output the result when the timestamp t1 of the latest stored data in the slave queue is greater than or equal to the sum of the timestamp t of the earliest unconnected data existing in the main queue and the right boundary b of the preset time range [a, b].

[0095] Specifically, for example, if the preset time range is [-2, +2], and there is no unconnected data in a certain cache structure, when adding the data at the 10th second to the main queue of a certain cache structure, if it is queried that the data in the slave queue of this cache structure has been added up to the data at the 12th second or the data after the 12th second at this time, it means that the connection has been triggered and the result is ready to be output, and the data in the slave queue within the time range of [8, 12] is connected to the data at the 10th second of the main queue. Further, when adding data to the slave queue of this cache structure, the timestamp t of the earliest unconnected data existing in the main queue will also be obtained each time. When the earliest and unconnected data in the main queue is the data at the 10th second, and the timestamp of the latest data in the slave queue of this cache structure is the 11th second at this time, the connection will be triggered, but the output will not be triggered. However, when adding the data at the 12th second to the slave queue, the result will be output immediately.

[0096] Further, in the first mode, generating and outputting the target data stream includes:

[0097] Determining the cache structure where the connection is triggered;

[0098] Retrieving the data with the timestamp t in the main queue of the cache structure and the data in the slave queue within the time range of [t + a, t + b];

[0099] Connecting the data with the timestamp t in the main queue to each data in the slave queue within the time range of [t + a, t + b] to generate the target data stream.

[0100] In the first embodiment of the present application, in the case where the determined data connection mode is the second mode, connecting the data in the main queue with the data in at least one slave queue according to the determined data connection mode includes:

[0101] Determining whether to trigger the data connection operation according to the local system timestamp according to the configuration information;

[0102] In the case of determining to trigger the data connection operation according to the local system timestamp:

[0103] As long as new data is stored in the main queue, trigger the connection of the data in the time range of [t0, t] in the slave queue with the data with the timestamp t in the main queue;

[0104] Wherein, t is the timestamp of the latest stored data in the main queue, and t0 is the timestamp of the previous stored data in the main queue;

[0105] Specifically, when triggering the data connection operation based on the local system timestamp, there is data input to the main queue in the cache structure at the 8th second and the 10th second. When adding the data at the 8th second to the main queue, a connection is triggered once. When adding the data at the 10th second to the main queue, the connection between the data in the slave queue within the time range of [8, 10] and the data at the 10th second in the main queue is triggered.

[0106] In the case of determining not to trigger the data connection operation based on the local system timestamp:

[0107] Obtain the timestamp t of the earliest unconnected data existing in the main queue;

[0108] Query the timestamp t1 of the latest stored data in the slave queue;

[0109] When the timestamp t1 of the latest stored data in the slave queue is greater than or equal to the timestamp t of the earliest unconnected data existing in the main queue, trigger the connection between the data in the slave queue within the time range of [t0, t] and the data with the timestamp t in the main queue;

[0110] Wherein, t0 is the timestamp of the previously stored data in the main queue.

[0111] Specifically, when not triggering the data connection operation based on the local system timestamp, there is data input to the main queue in the cache structure at the 8th second and the 10th second. When adding the data at the 8th second to the main queue, a connection is triggered once. After adding the data at the 10th second to the main queue and adding the data at the 10th second to the slave queue, the connection between the data in the slave queue within the time range of [8, 10] and the data at the 10th second in the main queue will be triggered.

[0112] Triggering the data connection operation based on the local system timestamp is applicable to scenarios with high requirements for real-time performance. For example, in financial transactions, online advertising, real-time monitoring systems, etc., once new data is stored in the main queue, it is necessary to immediately trigger the data connection and make corresponding decisions or feedback. In these scenarios, processing the newly entered data as soon as possible is usually more important than optimizing the computing efficiency. Not triggering the data connection operation based on the local system timestamp is more applicable to scenarios with high requirements for computing efficiency and resource optimization. For example, in big data analysis, machine learning model training, etc., the access and processing of data often involve a large amount of computing and storage resources. If the data connection is triggered too early or too frequently, it may lead to waste of computing resources and even affect the stability of the system. Therefore, in these scenarios, a more conservative connection triggering mechanism (that is, only when the data in the slave queue is ready, the connection is made) can effectively optimize the computing efficiency and resource usage.

[0113] Furthermore, in the second mode, generating and outputting the target data stream includes:

[0114] Determine the cache structure that triggers the connection;

[0115] Retrieve the data of the main queue at the timestamp t and the data of the slave queue within the time range [t0, t];

[0116] Connect the data of the main queue at the timestamp t with each data of the slave queue within the time range [t0, t] to generate the target data stream.

[0117] In the second mode, the data to be connected is determined by the timestamp of the earliest unconnected data existing in the main queue and the timestamp of the previously stored data, which can dynamically respond to the changes in the data stream, adjust the connected data in real time, and thus provide more timely and accurate results. Since two timestamps are used to determine the data to be connected, the relevant data can be more precisely matched, reducing the inaccurate connection caused by timestamp errors.

[0118] It is worth mentioning that when processing data, sometimes it is necessary to calculate the data in the corresponding time range of the slave queue. For example, when it is necessary to calculate the sum, product, average, standard deviation, variance, covariance, correlation coefficient, covariance, least squares estimation, weighted average, weighted sum, minimum value, maximum value, start value, end value, median, percentile... etc. of the data in the corresponding time range of the slave queue, in order to further improve the efficiency of data processing, specifically but not limitedly, the present application proposes a method for connecting the data in the main queue with the data in at least one slave queue based on the timestamp information of the data in multiple data queues:

[0119] Perform a preset aggregation operation on the data in some or all of the slave queues based on the timestamp information of the data in multiple slave queues to obtain the aggregated data;

[0120] Connect the data in the main queue with the aggregated data based on the timestamp information of the data in multiple data queues, or connect the data in the main queue with the aggregated data and the data in at least one slave queue based on the timestamp information of the data in multiple data queues.

[0121] Specifically, for example, if one of the multiple slave queues is the price and the sum of the price needs to be obtained during connection, then when creating the cache structure, an aggregator for the sum operation can be preset in the slave queue using metaprogramming code, so that each time new data is added to the slave queue, the aggregation operation will be triggered, and the sum operation will be performed on the data in the slave queue in advance. When outputting the target data stream finally, the result of the sum operation can be directly retrieved. With such settings, the complexity of data processing is reduced, the pressure of data processing is lowered, and the efficiency of data processing is improved.

[0122] Further, in the case where multiple original data streams include more than two data streams, for example, including three original data streams, each created cache structure will include a main queue and two slave queues. Both of the two slave queues can be preset with aggregation operations and can be preset with multiple aggregation operations. When outputting the target data stream, the data in the main queue is connected to the data in the two slave queues, and the corresponding aggregated data of the slave queues is retrieved. In the case where the original data stream only includes two data streams, for example, only includes the first data stream and the second data stream, each cache structure includes a main queue and a slave queue. The main queue corresponds to the first data stream, and the slave queue corresponds to the second data stream. The slave queue can also be preset with aggregation operations and can also be preset with multiple aggregation operations. When outputting the target data stream, the data in the main queue is connected to the data in the slave queue and the corresponding aggregated data, or can also be only connected to the aggregated data of the slave queue.

[0123] Based on the above method, still taking the example where the original data stream only contains the first data stream and the second data stream, in the second embodiment of the present application, when aggregating data is incorporated into the connection, in the case where the determined data connection mode is the first mode, connecting the data in the main queue to the data in at least one slave queue according to the determined data connection mode includes:

[0124] Obtain the timestamp t of the earliest unconnected data existing in the main queue;

[0125] Query the timestamp t1 of the latest deposited data in the slave queue;

[0126] When the timestamp t1 of the latest deposited data in the slave queue is greater than or equal to the sum of the timestamp t of the earliest unconnected data existing in the main queue and the left boundary a of the preset time range [a, b], trigger the connection of the aggregated data in the slave queue within the time range [t + a, t + b] to the data with the timestamp t in the main queue, or trigger the connection of the data within the time range [t + a, t + b] in the slave queue and the corresponding aggregated data to the data with the timestamp t in the main queue, and output the result when the timestamp t1 of the latest deposited data in the slave queue is greater than or equal to the sum of the timestamp t of the earliest unconnected data existing in the main queue and the right boundary b of the preset time range [a, b].

[0127] Further, in the first mode, generating and outputting the target data stream includes:

[0128] Determine the cache structure where the connection is triggered;

[0129] Retrieve the data with the timestamp t in the main queue of the cache structure, the data within the time range [t + a, t + b] in the slave queue, and the corresponding aggregated data;

[0130] Connect the data with the timestamp of the main queue at time t to the aggregated data corresponding to the slave queue within the time range [t + a, t + b], or, connect the data with the timestamp of the main queue at time t to the aggregated data corresponding to the slave queue within the time range [t + a, t + b] and each piece of data to generate a target data stream.

[0131] Specifically, if it is determined that the cache structure triggering the connection is the first cache structure, retrieve the data of the main queue at the tenth second as a, and the data of the slave queue within the time range [8, 12] as X1, X2, X3. The aggregated data obtained by performing a summation aggregation operation on the data of the slave queue within the time range [8, 12] is W1. The target data stream generated after connecting the data in the main queue, the data in the slave queue, and the corresponding aggregated data is:

[0132]

[0133]

[0134] Furthermore, after multiple cache structures all trigger connections, the multiple generated target data streams can be summarized and output. When summarizing and outputting, the cache structure and the timestamp of the main queue data when the connection is triggered can also be output. As an example:

[0135] Cache structure Timestamp Main queue Slave queue Aggregate data First cache structure 000 a X1 W1 First cache structure 000 a X2 W1 First cache structure 000 a X3 W1 First cache structure 001 b Y1 W2 First cache structure 001 b Y2 W2 First cache structure 001 b Y3 W2 Second cache structure 000 c Z1 W3 Second cache structure 002 d U1 W4

[0136] In the first mode, by setting a preset time range [a, b], it is possible to precisely determine which data needs to be processed. Taking the data with the timestamp of the main queue at time t as an example, all the data in the slave queue within the time range [t + a, t + b] will be connected and processed. If this preset time range is adjusted, the amount of data to be processed can be changed. If the set time range is larger, then more data will be processed, and since more data is covered, the accuracy of data processing will be improved to some extent. On the contrary, if the set time range is smaller, then the amount of data processed will be reduced, which will reduce the burden of data processing and improve the performance of data processing. Therefore, a balance point can be found between ensuring the accuracy of data processing and optimizing the performance of data processing by adjusting the preset time range.

[0137] In the second embodiment of the present application, the aggregated data is incorporated into the connection. In the case where the determined data connection mode is the second mode, connecting the data in the main queue to the data in at least one slave queue according to the determined data connection mode includes: determining whether to trigger the data connection operation based on the local system timestamp according to the configuration information;

[0138] In the case where it is determined to trigger the data connection operation based on the local system timestamp:

[0139] Whenever new data is stored in the main queue, it triggers the connection of the aggregated data within the time range [t0, t] of the slave queue with the data at the timestamp t of the main queue, or triggers the connection of the data within the time range [t0, t] of the slave queue and the corresponding aggregated data with the data at the timestamp t of the main queue;

[0140] where t is the timestamp of the latest stored data in the main queue, and t0 is the timestamp of the previous stored data in the main queue;

[0141] In the case of determining not to trigger the data connection operation based on the local system timestamp:

[0142] Obtain the timestamp t of the earliest unconnected data existing in the main queue;

[0143] Query the timestamp t1 of the latest stored data in the slave queue;

[0144] When the timestamp t1 of the latest stored data in the slave queue is greater than or equal to the timestamp t of the earliest unconnected data existing in the main queue, trigger the connection of the aggregated data within the time range [t0, t] of the slave queue with the data at the timestamp t of the main queue, or trigger the connection of the data within the time range [t0, t] of the slave queue and the corresponding aggregated data with the data at the timestamp t of the main queue.

[0145] Furthermore, in the second mode, generating and outputting the target data stream includes:

[0146] Determine the cache structure that triggers the connection;

[0147] Retrieve the data at the timestamp t of the main queue, the data within the time range [t0, t] of the slave queue, and the corresponding aggregated data from the cache structure;

[0148] Connect the data at the timestamp t of the main queue with the corresponding aggregated data within the time range [t0, t] of the slave queue, or with the corresponding aggregated data within the time range [t0, t] of the slave queue and each data to generate the target data stream.

[0149] Refer to Figure 2, Embodiment 2 of the present application discloses a processing system for streaming data. A processing system for streaming data includes: a data acquisition module, a data caching module, a data distribution module, and a data connection module. The data acquisition module is used to receive multiple original data streams, and the data in the original data streams is continuously input according to the time sequence. The data caching module is used to create multiple caching structures, and respectively establish multiple data queues corresponding to the multiple original data streams in each caching structure. The multiple data queues include a main queue and at least one slave queue. The data distribution module is used to distribute data to the corresponding caching structures according to the characteristic information of each data, and for each caching structure, store the data from different original data streams into the corresponding data queues in sequence. The data connection module is used to, for each caching structure, connect the data in the main queue with the data in at least one slave queue based on the timestamp information of the data in the multiple data queues to generate and output a target data stream.

[0150] The original data streams acquired by the data acquisition module are stored in the corresponding data queues in the caching structures. When the original data streams enter the caching structures, the data distribution module will distribute the data of the original data streams to different caching structures according to the characteristic information of the data. With this setting, corresponding data queues are created for each original data stream in the caching structures, and each caching structure can process multiple data streams in parallel, thereby improving the overall data processing speed. Distributing the data to the corresponding caching structures according to the characteristic information of the data enables the data of the same original data stream to be independently processed in multiple caching structures, further improving the data processing efficiency, reducing the conflict of data with different characteristic information, and improving the accuracy of data processing.

[0151] Embodiment 3 of the present application discloses a computer-readable storage medium. A computer-readable storage medium stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the above-mentioned data stream connection method.

[0152] The above are all the preferred embodiments of the present application. The protection scope of the present application is not limited by this. Therefore, all equivalent changes made according to the structure, shape, and principle of the present application should be covered within the protection scope of the present application.

Claims

1. A data stream connection method, characterized in that, The method includes: Receiving a plurality of original data streams, and continuously inputting the data in the original data streams in a time series; Creating a plurality of cache structures, and respectively establishing a plurality of data queues corresponding one-to-one to the plurality of original data streams in each of the cache structures, where the plurality of data queues include a main queue and at least one slave queue; Storing the data into the corresponding cache structure according to the feature information of each data, and for each cache structure, storing the data from different original data streams into the corresponding data queues in sequence; For each cache structure, connecting the data in the main queue with the data in the at least one slave queue based on the timestamp information of the data in the plurality of data queues, so as to generate and output a target data stream; Wherein, the connecting the data in the main queue with the data in the at least one slave queue based on the timestamp information of the data in the plurality of data queues includes: Determining a data connection mode according to preset configuration information, and connecting the data in the main queue with the data in the at least one slave queue according to the determined data connection mode; Wherein, the data connection mode includes a first mode and a second mode. In the first mode, the data to be connected in the slave queue in the cache structure is determined by the timestamp t of the earliest unconnected data existing in the main queue and a preset time range [a, b]. In the second mode, the data to be connected in the slave queue in the cache structure is determined by the timestamp t of the earliest unconnected data existing in the main queue and the timestamp t0 of the previously stored data.

2. The data stream connection method according to claim 1, wherein The connecting the data in the main queue with the data in the at least one slave queue based on the timestamp information of the data in the plurality of data queues includes: Performing a preset aggregation operation on the data in some or all of the slave queues based on the timestamp information of the data in the plurality of slave queues to obtain aggregated data; Connecting the data in the main queue with the aggregated data based on the timestamp information of the data in the plurality of data queues, or connecting the data in the main queue with the data in the at least one slave queue and the aggregated data based on the timestamp information of the data in the plurality of data queues.

3. The data stream connection method according to claim 1, characterized in that The plurality of original data streams only include a first data stream and a second data stream, and each cache structure includes a main queue and a slave queue. The main queue corresponds to the first data stream, and the slave queue corresponds to the second data stream.

4. The data stream connection method according to claim 3, wherein When the determined data connection mode is the first mode, the connecting the data in the main queue with the data in the at least one slave queue according to the determined data connection mode includes: Obtaining the timestamp t of the earliest unconnected data existing in the main queue; Querying the timestamp t1 of the latest stored data in the slave queue; When the timestamp t1 of the data newly stored in the slave queue is greater than or equal to the sum of the timestamp t of the earliest unconnected data existing in the master queue and the left boundary a of the preset time range [a, b], trigger the connection of the data in the slave queue within the time range [t + a, t + b] to the data with the timestamp t in the master queue, and output the result when the timestamp t1 of the data newly stored in the slave queue is greater than or equal to the sum of the timestamp t of the earliest unconnected data existing in the master queue and the right boundary b of the preset time range [a, b].

5. The data stream connection method according to claim 3, wherein In the case where the determined data connection mode is the second mode, the connection of the data in the master queue to the data in the at least one slave queue according to the determined data connection mode includes: Determine whether to trigger the data connection operation based on the local system timestamp according to the configuration information; In the case of determining to trigger the data connection operation based on the local system timestamp: As long as new data is stored in the master queue, trigger the connection of the data in the slave queue within the time range [t0, t] to the data with the timestamp t in the master queue; where t is the timestamp of the data newly stored in the master queue, and t0 is the timestamp of the previous data stored in the master queue; In the case of determining not to trigger the data connection operation based on the local system timestamp: Obtain the timestamp t of the earliest unconnected data existing in the master queue; Query the timestamp t1 of the data newly stored in the slave queue; When the timestamp t1 of the data newly stored in the slave queue is greater than or equal to the timestamp t of the earliest unconnected data existing in the master queue, trigger the connection of the data in the slave queue within the time range [t0, t] to the data with the timestamp t in the master queue; where t0 is the timestamp of the previous data stored in the master queue.

6. The data stream connection method according to claim 3, wherein In the case where the determined data connection mode is the first mode, the connection of the data in the master queue to the data in the at least one slave queue according to the determined data connection mode includes: Obtain the timestamp t of the earliest unconnected data existing in the master queue; Query the timestamp t1 of the data newly stored in the slave queue; When the timestamp t1 of the data newly stored in the slave queue is greater than or equal to the sum of the timestamp t of the earliest unconnected data existing in the master queue and the left boundary a of the preset time range [a, b], trigger the connection of the aggregated data in the slave queue within the time range [t + a, t + b] to the data with the timestamp t in the master queue, or trigger the connection of the data and the corresponding aggregated data in the slave queue within the time range [t + a, t + b] to the data with the timestamp t in the master queue, and output the result when the timestamp t1 of the data newly stored in the slave queue is greater than or equal to the sum of the timestamp t of the earliest unconnected data existing in the master queue and the right boundary b of the preset time range [a, b].

7. The data stream connection method according to claim 3, wherein In the case where the determined data connection mode is the second mode, the connection of the data in the master queue to the data in the at least one slave queue according to the determined data connection mode includes: Determine whether to trigger a data connection operation based on the local system timestamp according to the configuration information; In the case of determining to trigger a data connection operation based on the local system timestamp: Whenever new data is stored in the main queue, trigger the connection of the aggregated data within the time range [t0, t] of the slave queue to the data with the timestamp t in the main queue, or trigger the connection of the data within the time range [t0, t] of the slave queue and the corresponding aggregated data to the data with the timestamp t in the main queue; Wherein, t is the timestamp of the latest stored data in the main queue, and t0 is the timestamp of the previous stored data in the main queue; In the case of determining not to trigger a data connection operation based on the local system timestamp: Obtain the timestamp t of the earliest unconnected data existing in the main queue; Query the timestamp t1 of the latest stored data in the slave queue; When the timestamp t1 of the latest stored data in the slave queue is greater than or equal to the timestamp t of the earliest unconnected data existing in the main queue, trigger the connection of the aggregated data within the time range [t0, t] of the slave queue to the data with the timestamp t in the main queue, or trigger the connection of the data within the time range [t0, t] of the slave queue and the corresponding aggregated data to the data with the timestamp t in the main queue; Wherein, t0 is the timestamp of the previous stored data in the main queue.

8. A processing system for streaming data, characterized in that, It includes: A data acquisition module for receiving multiple original data streams, and the data in the original data streams is continuously input in time series; A data caching module for creating multiple caching structures, and respectively establishing multiple data queues corresponding to the multiple original data streams in each of the caching structures, and the multiple data queues include a main queue and at least one slave queue; A data distribution module for storing the data into the corresponding caching structure according to the characteristic information of each data, and for each caching structure, storing the data from different original data streams into the corresponding data queues in sequence; A data connection module for, for each caching structure, connecting the data in the main queue to the data in the at least one slave queue based on the timestamp information of the data in the multiple data queues to generate and output a target data stream; Wherein, the connecting the data in the main queue to the data in the at least one slave queue based on the timestamp information of the data in the multiple data queues includes: Determine a data connection mode according to preset configuration information, and connect the data in the main queue to the data in the at least one slave queue according to the determined data connection mode; Wherein, the data connection mode includes a first mode and a second mode. In the first mode, the data to be connected in the slave queue in the caching structure is determined by the timestamp t of the earliest unconnected data existing in the main queue and a preset time range [a, b]. In the second mode, the data to be connected in the slave queue in the caching structure is determined by the timestamp t of the earliest unconnected data existing in the main queue and the timestamp t0 of the previous stored data.

9. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement: the data stream connection method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Flow aggregation method and device and electronic equipment

    CN111488222A

  • Data processing method and device of Flink computing framework, equipment and storage medium

    CN114116802A