Method and System for Merging and Processing Multi-Source Data Streams in a Cloud Environment

By using a combination of ring cache queues and slots in a cloud environment, the time series challenges during multi-data stream merging are solved, and the time sequence reorganization and efficient processing of data packets are realized, and the performance and reliability of network monitoring and data analysis are improved.

CN119892753BActive Publication Date: 2025-07-01SHANGHAI NETIS TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510387046.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-01
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

In a cloud environment, there are time series challenges when merging multiple data streams, resulting in out-of-order packet time, affecting the accuracy and efficiency of network monitoring and analysis.

Method used

The ring cache queue is combined with slot, and the ring cache queue is initialized according to the maximum cache time and time granularity to ensure that the data packets are reorganized in the original chronological order. By calculating the cache slot of the data packet and sorting it when cached or output, the problem of out-of-order packet time is solved.

Benefits of technology

It realizes efficient and accurate processing of multi-source data flows in a cloud environment, ensures that data packets are saved and analyzed in chronological order, and improves the performance and reliability of network monitoring and data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119892753B_ABST
    Figure CN119892753B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for merging and processing multi-source data streams in a cloud environment. The method includes: Step S1: Initialize a circular buffer queue according to configuration parameters; the configuration parameters include a maximum caching time and a time granularity; the circular buffer queue is used to store data packets from multiple collection points; Step S2: Calculate the corresponding starting slot position according to the starting data packet, cache the starting data packet into the circular buffer queue according to the slot, and record the cache start time at the same time; Step S3: Calculate the cache slot of the data packet according to the time of the newly arrived data packet and the cache start time, and cache the data packet into the corresponding slot. The present invention designs a circular buffer queue of a certain size according to the maximum caching time and granularity, and through the combination of the circular buffer queue and the slot, ensures that data packets from multiple collection points can be correctly reorganized in the order of their original occurrence time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer network technology and data processing, and particularly to a method and system for merging and processing multi-source data streams in a cloud environment. Background Art

[0002] In traditional network monitoring, data packets carrying services are copied through network devices (switches) or network mirroring devices (TAPs) and then transmitted to the network ports of network monitoring and analysis devices for capture, and then analyzed and processed by the analysis devices. The time of data packets is uniformly added by the network monitoring and analysis devices when capturing the data packets. In this case, the time series of data on one capture port is consistent, which is called a data stream, as Figure 1 shown.

[0003] In a cloud environment (or virtualized environment), the capture and analysis processing of data packets are often separated. The capture of data packets occurs in cloud containers (or virtual machines) or on hosts, which can be called capture at the capture point. Then, the captured data packets are transmitted to a network monitoring and analysis device at a remote end (even under the cloud) through network protocols for unified analysis and processing, which can be called processing at the processing point, as Figure 2 shown. Obviously, in this case, the time of data packets at each capture point is in order, which are time-ordered data streams one by one; at the processing point, it is the merger of multiple data streams, and the data packet time is out of order.

[0004] However, for network monitoring and analysis processing, it is necessary to ensure that the data is in time order. The data needs to be saved by time, statistically analyzed by time, alarm processed by time, and visually presented, etc.

[0005] Patent document CN101369869A provides a network monitoring method, a network monitoring device, a line fault prevention system, and a computer program for a network monitoring device, aiming to solve the problems of degradation of line quality and line interruption caused by rainfall, which is essentially different from the technical method adopted by the present invention.

[0006] The present invention aims to solve the time-order challenge in the merger of multi-data streams in a cloud environment, and provides an efficient and accurate data packet processing solution, which is applicable to network monitoring and data analysis applications in a modern cloud computing environment. Summary of the Invention

[0007] Aiming at the defects in the prior art, the purpose of the present invention is to provide a method and system for merging and processing multi-source data streams in a cloud environment.

[0008] According to a method for merging and processing multi-source data streams in a cloud environment provided by the present invention, it includes:

[0009] Step S1: Initialize a circular buffer queue according to configuration parameters;

[0010] The configuration parameters include the maximum cache time and the time granularity;

[0011] The circular buffer queue is used to store data packets from multiple collection points;

[0012] Step S2: Calculate the corresponding starting slot position according to the starting data packet, cache the starting data packet into the circular buffer queue according to the slot, and record the cache start time at the same time;

[0013] Step S3: Calculate the cache slot of the data packet according to the time of the newly arrived data packet and the cache start time, and cache the data packet into the corresponding slot; when the slot position is greater than the size of the cache queue, drive the output of the cached data packet, and update the starting slot position and the starting data packet at the same time.

[0014] Preferably, it further includes:

[0015] Step S4: When caching or outputting data packets, select whether to sort the data packets according to requirements.

[0016] Preferably, the time granularity is used to round the data packet time and file the data packet into a certain time slot.

[0017] Preferably, the calculation formula of the circular buffer queue is:

[0018] slot_count = max_timespan / time_granularity;

[0019] where max_timespan is the maximum cache time span and time_granularity is the time granularity.

[0020] Preferably, the calculation process of the cache slot of the data packet includes:

[0021] Read in the time pkt_time of the data packet, the cache start time cache_start_time and the time granularity time_granularity;

[0022] Judge whether cache_start_time == 0. If so, it means that no data packet has been received;

[0023] Update cache_start_time using pkt_time and time_granularity:

[0024] The value obtained by dividing cache_start_time by time_granularity, after rounding, is then multiplied by time_granularity;

[0025] Calculate the time difference pkt_delta_time using pkt_time and cache_start_time. The value of pkt_delta_time is: pkt_time minus cache_start_time;

[0026] Calculate the slot position pkt_slot using pkt_delta_time and time_granularity. The value of pkt_slot is:

[0027] The value obtained by dividing pkt_delta_time by time_granularity, and then rounding.

[0028] Preferably, according to the value of pkt_slot, the data packets are divided into cached data packets, output data packets, and discarded data packets and then processed separately:

[0029] Read pkt_slot, slot_count. If pkt_slot < 0, then the data packet is divided into discarded data packets for processing;

[0030] If pkt_slot ≥ 0 and pkt_slot ≥ slot_count, then the data packet is divided into output data packets for processing, otherwise the data packet is divided into cached data packets for processing.

[0031] Preferably, the processing process of the cached data packet includes:

[0032] Read in pkt_slot, cache_start_slot, slot_count, and cache_ring;

[0033] Among them, cache_start_slot is the starting slot of the cache;

[0034] Calculate index using pkt_slot, cache_start_slot, and slot_count. index is actually the physical storage position of the slot in the circular cache cache_ring; The value of index is:

[0035] (pkt_slot + cache_start_slot) % slot_count;

[0036] Among them, % is the modulo operation;

[0037] Add the data packet to the array of cache_ring[index]; use index to obtain the storage unit of cache_ring[index], and add the data packet to the array of this storage unit.

[0038] Preferably, the processing process of outputting the data packet includes:

[0039] Read in pkt_slot, slot_count, cache_ring, cache_start_slot, pkt_time, and time_granularity;

[0040] Calculate export_count according to pkt_slot, slot_count, and export_count takes the value of:

[0041] (pkt_slot - slot_count + 1) % slot_count;

[0042] where, % is the modulo operation;

[0043] Output export_count units of cache_ring starting from cache_slot_start;

[0044] Update cache_start_slot according to cache_slot_star, export_count, and slot_count, and cache_start_slot takes the value of:

[0045] (cache_slot_star + export_count) % slot_count;

[0046] Update cache_start_time using pkt_time and time_granularity.

[0047] A system for merging and processing multi-source data streams in a cloud environment according to the present invention includes:

[0048] Module M1: Initialize a circular buffer queue according to configuration parameters;

[0049] The configuration parameters include the maximum cache time and time granularity;

[0050] The circular buffer queue is used to store data packets from multiple collection points;

[0051] Module M2: Calculate the corresponding starting slot position according to the starting data packet, cache the starting data packet into the circular cache queue according to the slot, and record the cache start time at the same time;

[0052] Module M3: Calculate the cache slot of the data packet according to the time of the newly arrived data packet and the cache start time, and cache the data packet into the corresponding slot; when the slot position is greater than the cache queue size, drive the output of the cached data packet, and update the starting slot position and the starting data packet at the same time.

[0053] Preferably, it further includes:

[0054] Module M4: When caching or outputting data packets, select whether to sort the data packets according to requirements.

[0055] Preferably, the time granularity is used to round the data packet time and file the data packet into a certain time slot.

[0056] Preferably, the calculation formula of the circular cache queue is:

[0057] slot_count = max_timespan / time_granularity;

[0058] Among them, max_timespan is the maximum cache time span, and time_granularity is the time granularity.

[0059] Preferably, the calculation process of the cache slot of the data packet includes:

[0060] Read in the time pkt_time of the data packet, the cache start time cache_start_time, and the time granularity time_granularity;

[0061] Judge cache_start_time == 0, if so, it means that no data packet has been received;

[0062] Update cache_start_time using pkt_time and time_granularity:

[0063] The value of cache_start_time divided by time_granularity, after rounding, is then multiplied by time_granularity;

[0064] Calculate the time difference pkt_delta_time using pkt_time and cache_start_time. The value of pkt_delta_time is: pkt_time minus cache_start_time;

[0065] Calculate the slot position pkt_slot using pkt_delta_time and time_granularity. The value of pkt_slot is:

[0066] The value obtained by dividing pkt_delta_time by time_granularity, and then taking the integer.

[0067] Preferably, divide the data packet into cached data packets, output data packets, and discarded data packets according to the value of pkt_slot, and then process them separately:

[0068] Read pkt_slot and slot_count. If pkt_slot < 0, then process the data packet as a discarded data packet;

[0069] If pkt_slot ≥ 0 and pkt_slot ≥ slot_count, then process the data packet as an output data packet, otherwise process the data packet as a cached data packet.

[0070] Preferably, the processing process of the cached data packet includes:

[0071] Read pkt_slot, cache_start_slot, slot_count, and cache_ring;

[0072] Among them, cache_start_slot is the starting slot of the cache;

[0073] Calculate index using pkt_slot, cache_start_slot, and slot_count. index is actually the physical storage position of the slot in the circular cache cache_ring; The value of index is:

[0074] (pkt_slot + cache_start_slot) % slot_count;

[0075] Among them, % is the modulo operation;

[0076] Add the data packet to the array of cache_ring[index]; Use index to obtain the storage unit of cache_ring[index], and add the data packet to the array of this storage unit.

[0077] Preferably, the process of processing the output data packet includes:

[0078] Read in pkt_slot, slot_count, cache_ring, cache_start_slot, pkt_time, and time_granularity;

[0079] Calculate export_count according to pkt_slot, slot_count, and export_count takes the value of:

[0080] (pkt_slot - slot_count + 1) % slot_count;

[0081] where % is the modulo operation;

[0082] Output export_count units of cache_ring starting from cache_slot_start;

[0083] Update cache_start_slot according to cache_slot_star, export_count, and slot_count, and cache_start_slot takes the value of:

[0084] (cache_slot_star + export_count) % slot_count;

[0085] Update cache_start_time using pkt_time and time_granularity.

[0086] Compared with the prior art, the present invention has the following beneficial effects:

[0087] 1. The present invention designs a circular buffer queue of a certain size according to the maximum cache time and granularity, and combines the circular buffer queue with slots to ensure that data packets from multiple collection points can be correctly reorganized in the order of their original occurrence time.

[0088] 2. The present invention calculates the corresponding cache slot according to the first packet time, caches the data into the circular queue by slot, and records the cache start time at the same time. Then, continuously calculates the cache slot of the newly arrived data packet according to the time of the newly arrived data packet and the cache start time, and caches the data packet into the corresponding slot; when the slot position is greater than the size of the cache queue, drives the output of the cached data packet, and updates the start slot position and the start data packet at the same time, improving the overall performance of the system.

[0089] 3. The present invention can flexibly sort when caching data packets or when outputting data packets according to actual needs, and has good practicability.

[0090] Other beneficial effects of the present invention will be described by introducing specific technical features and technical solutions in the specific implementation manner. Those skilled in the art should be able to understand the beneficial technical effects brought by the described technical features and technical solutions through these introductions. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non - restrictive embodiments with reference to the accompanying drawings:

[0092] Figure 1 It is a schematic diagram of mirroring network data packets from a switch to a monitoring and analysis device in the present invention.

[0093] Figure 2 It is a schematic diagram of capturing data from a cloud environment to a monitoring and analysis device in the present invention.

[0094] Figure 3 It is a schematic diagram of the main process of the present invention.

[0095] Figure 4 It is a schematic diagram of the process of calculating pkt_slot in the present invention.

[0096] Figure 5 It is a schematic diagram of the process of judging caching, output, and discarding according to pkt_slot in the present invention.

[0097] Figure 6 It is a schematic diagram of the processing flow of caching data packets in the present invention.

[0098] Figure 7 It is a schematic diagram of the processing flow of outputting data packets in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0099] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0100] Regarding the problem of aggregating multiple data sources formed by multiple collection points in a cloud environment into a single data stream for processing, the present invention first designs a circular buffer queue of a certain size according to the maximum cache time and granularity. Then, it calculates the corresponding cache slot based on the first packet time, caches the data into the circular queue according to the slot, and records the cache start time at the same time. Then, it continuously calculates the cache slot of the newly arrived data packet according to the time of the newly arrived data packet and the cache start time, and caches the data packet into the corresponding slot. When the slot position is greater than the size of the cache queue, it drives the output of the cached data packet, and updates the start slot position and the start data packet at the same time. At the same time, according to actual needs, it can be designed to sort the data packets when caching them, or sort the data packets when outputting them.

[0101] Referring to Figure 3 As shown, a method for merging and processing multi-source data streams in a cloud environment is as follows:

[0102] Step 1: Load configuration.

[0103] The present invention includes the following two configurations:

[0104] max_timespan: The maximum cache time span. It is restricted by the maximum processing delay acceptable to the system.

[0105] time_granularity: The time granularity. It is used to round the packet time and archive the data packet into a certain time slot. Generally used granularities are seconds, minutes, hours, etc.

[0106] time_granularity is the smallest unit of the ring, which affects the processing frequency and also affects the time delay. The smaller the granularity, the higher the processing frequency and the smaller the time delay. On the other hand, it is restricted by the memory. The smaller the granularity, the more memory is required.

[0107] Step 2: Initialize the data packet cache. The data packet cache of the present invention includes the following parts:

[0108] slot_count: The value obtained by dividing max_timespan by time_granularity and then rounding up.

[0109] cache_ring: Use a circular cache, and the size of the ring is slot_count. Each storage unit of cache_ring is an array used to cache the data packets of the current slot.

[0110] All Rings use modulo operations, so the entire Ring is used cyclically. The cache_start_time can locate the starting position. Therefore, if a slot is full, the next slot will continue to be used according to the modulo operation, but the starting position will move forward.

[0111] cache_start_slot: Set the initial value to 0.

[0112] cache_start_time cache start time: Set the initial value to 0, indicating that no data packets have been received.

[0113] Step 3: Determine whether the program exits. Here, a Boolean variable, or a synchronization mechanism such as Signal or Mutex can be used to achieve this.

[0114] Step 4: Read the time of the data packet pkt_time and calculate pkt_slot.

[0115] pkt_time: That is, the time of the data packet.

[0116] pkt_slot: It is calculated based on pkt_time, cache_start_time, and time_granularity.

[0117] Refer to Figure 4 As shown below, the specific process is as follows:

[0118] Step 4.1: Read in pkt_time, cache_start_time, and time_granularity.

[0119] Step 4.2: Determine whether cache_start_time == 0. This value == 0 indicates that no data packets have been received.

[0120] Step 4.3: Update cache_start_time using pkt_time and time_granularity. Actually, it initializes cache_start_time by regularizing according to the time_granularity based on the first received data packet.

[0121] The specific calculation method is: Divide the value of cache_start_time by time_granularity, take the integer part, and then multiply by time_granularity.

[0122] Step 4.4: Calculate the time difference pkt_delta_time using pkt_time and cache_start_time. The value of pkt_delta_time is: pkt_time minus cache_start_time. This value may be negative, indicating that the new data packet is earlier than the cache time and will be processed in Step 5.

[0123] Step 4.5: Calculate the pkt_slot position using pkt_delta_time and time_granularity. The value of pkt_slot is: the integer part of the value obtained by dividing pkt_delta_time by time_granularity. Similarly, this value may be negative, indicating that the new data packet is earlier than the cache time and will be processed in Step 5.

[0124] Step 5: Determine caching, output, or discarding based on pkt_slot. The determination process also requires introducing slot_count, as Figure 5 shown.

[0125] Step 5.1: Read in pkt_slot and slot_count.

[0126] Step 5.2: Determine if pkt_slot < 0.

[0127] Yes: It means that the data packet time is earlier than the current cache start time, so "discarding" processing is required. Refer to Step 8.

[0128] No: It means that the data packet needs further determination.

[0129] Step 5.3: Determine if pkt_slot >= slot_count.

[0130] Yes: It means that the data packet time is greater than the maximum cache time range, so "output" processing is required. Refer to Step 7.

[0131] No: It means that the data packet time is within the cache time range, so only "caching" processing is required. Refer to Step 6.

[0132] Step 6: Processing of caching data packets, as Figure 6 shown.

[0133] Step 6.1: Read in pkt_slot, cache_start_slot, slot_count, and cache_ring.

[0134] Step 6.2: Use pkt_slot, cache_start_slot, and slot_count to calculate index. Index is actually the physical storage location of the slot in the cache_ring.

[0135] The value of index is: (pkt_slot + cache_start_slot) % slot_count, where % is the modulo remainder operation.

[0136] Step 6.3: Add the data packet to the array of cache_ring[index]. Use index to obtain the storage unit (data packet array) of cache_ring[index], and add the data packet to the array of the storage unit. The addition operation here can be a simple direct append operation at the end of the array, or an insertion sort can be used to keep the time order of the array. In the case of a simple append operation, the current append process has no overhead, but the data packets need to be sorted by time when they are output, which brings processing delays; but asynchronous threads can be used to reduce the burden. In the case of insertion sort, the sorting overhead is amortized to the processing process of each packet, which increases the time overhead of each packet processing, but the delay of outputting the data packet is small. The choice of the two addition operations can be selected according to the actual scenario.

[0137] Step 7: Processing of output data packets, such as Figure 7 shown.

[0138] Step 7.1: Read in pkt_slot, slot_count, cache_ring, cache_start_slot, pkt_time, time_granularity.

[0139] Step 7.2: Calculate export_count based on pkt_slot, slot_count and.

[0140] export_count takes the following values:

[0141] (pkt_slot - slot_count + 1) % slot_count, where % is the modulo remainder operation.

[0142] Step 7.3: Output export_count units of the cache_ring starting from cache_slot_start. Using the iterative method of a standard circular queue, export_count storage unit data can be obtained starting from cache_slot_start. The data of each unit is a list of data packets, so the obtained data is a set of lists of data packets. If Step 6.3 is an insertion sort operation, sorting is no longer required at this time, and only these lists of data packets need to be concatenated into a large list in order; if Step 6.3 is a simple append operation, sorting is required at this time.

[0143] Step 7.4: Update according to cache_slot_star, export_count, and slot_count.

[0144] cache_start_slot.

[0145] The value of cache_start_slot is: (cache_slot_star + export_count) % slot_count.

[0146] Step 7.5: Update cache_start_time using pkt_time and time_granularity. This step has the same calculation method as Step 4.3.

[0147] Step 8: Processing of discarded data packets. For discarded data packets, the present invention does not perform special processing, and conventional operations such as counting and outputting logs can be performed.

[0148] The present invention also provides a system for merging and processing multi-source data streams in a cloud environment. The system for merging and processing multi-source data streams in a cloud environment can be implemented by executing the process steps of the method for merging and processing multi-source data streams in a cloud environment, that is, those skilled in the art can understand the method for merging and processing multi-source data streams in a cloud environment as a preferred implementation manner of the system for merging and processing multi-source data streams in a cloud environment.

[0149] Specifically, a system for merging and processing multi-source data streams in a cloud environment includes:

[0150] Module M1: Initialize a circular buffer queue according to configuration parameters;

[0151] The configuration parameters include the maximum cache time and time granularity;

[0152] The circular buffer queue is used to store data packets from multiple collection points;

[0153] Module M2: Calculate the corresponding starting slot position according to the starting data packet, cache the starting data packet into the circular buffer queue by slot, and record the cache start time at the same time;

[0154] Module M3: Calculate the cache slot of the data packet according to the time of the newly arrived data packet and the cache start time, and cache the data packet into the corresponding slot; when the slot position is greater than the cache queue size, drive the output of the cached data packet, and update the starting slot position and the starting data packet at the same time.

[0155] It also includes:

[0156] Module M4: When caching or outputting data packets, select whether to sort the data packets according to requirements.

[0157] The time granularity is used to round the data packet time and archive the data packet into a certain time slot.

[0158] The calculation formula of the circular buffer queue is:

[0159] slot_count = max_timespan / time_granularity;

[0160] Where max_timespan is the maximum cache time span and time_granularity is the time granularity.

[0161] The calculation process of the cache slot of the data packet includes:

[0162] Read in the time pkt_time of the data packet, the cache start time cache_start_time, and the time granularity time_granularity;

[0163] Judge cache_start_time == 0, if so, it means that no data packet has been received;

[0164] Update cache_start_time using pkt_time and time_granularity:

[0165] The value of cache_start_time divided by time_granularity, after rounding, is then multiplied by time_granularity;

[0166] Calculate the time difference pkt_delta_time using pkt_time and cache_start_time, and the value of pkt_delta_time is: pkt_time minus cache_start_time;

[0167] Calculate the slot position pkt_slot using pkt_delta_time and time_granularity. The value range of pkt_slot is as follows:

[0168] The integer part of the quotient obtained by dividing pkt_delta_time by time_granularity.

[0169] After classifying the data packets into cached data packets, output data packets, and discarded data packets according to the value of pkt_slot, process them separately:

[0170] Read pkt_slot and slot_count. If pkt_slot < 0, classify the data packet as a discarded data packet for processing;

[0171] If pkt_slot ≥ 0 and pkt_slot ≥ slot_count, classify the data packet as an output data packet for processing; otherwise, classify the data packet as a cached data packet for processing.

[0172] The processing process of the cached data packet includes:

[0173] Read pkt_slot, cache_start_slot, slot_count, and cache_ring;

[0174] Among them, cache_start_slot is the starting slot of the cache;

[0175] Calculate index using pkt_slot, cache_start_slot, and slot_count. index is actually the physical storage position of the slot in the circular cache cache_ring. The value range of index is as follows:

[0176] (pkt_slot + cache_start_slot) % slot_count;

[0177] Among them, % is the modulo operation;

[0178] Add the data packet to the array of cache_ring[index]; use index to obtain the storage unit of cache_ring[index], and add the data packet to the array of this storage unit.

[0179] The processing process of the output data packet includes:

[0180] Read in pkt_slot, slot_count, cache_ring, cache_start_slot, pkt_time, and time_granularity;

[0181] Calculate export_count based on pkt_slot, slot_count, and export_count takes the value of:

[0182] (pkt_slot - slot_count + 1) % slot_count;

[0183] where % is the modulo operation;

[0184] Output export_count units of cache_ring starting from cache_slot_start;

[0185] Update cache_start_slot according to cache_slot_star, export_count, and slot_count, and cache_start_slot takes the value of:

[0186] (cache_slot_star + export_count) % slot_count;

[0187] Update cache_start_time using pkt_time and time_granularity.

[0188] Those skilled in the art know that in addition to implementing the system and its various devices, modules, and units provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system and its various devices, modules, and units provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc. to achieve the same functions. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a kind of hardware component, and the devices, modules, and units included therein for implementing various functions can also be regarded as the structure within the hardware component; the devices, modules, and units for implementing various functions can also be regarded as either software modules for implementing the method or the structure within the hardware component.

[0189] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A method for merging and processing multi-source data streams in a cloud environment, characterized in that: include: Step S1: Initialize a ring buffer queue according to configuration parameters; The configuration parameters include maximum cache time and time granularity; The circular buffer queue is used to store data packets from multiple collection points; Step S2: Calculate the corresponding starting slot position according to the starting data packet, cache the starting data packet into the ring cache queue according to the slot, and record the cache start time; Step S3: Calculate the cache slot of the data packet according to the time of the newly arrived data packet and the cache start time, and cache the data packet into the corresponding slot; when the slot position is larger than the cache queue size, drive the output of the cached data packet, and update the start slot position and the start data packet at the same time; The cache slot calculation process for a data packet includes: The time pkt_time of reading the data packet, the cache start time cache_start_time and the time granularity time_granularity; Check cache_start_time == 0. If so, it means no data packet has been received. Update cache_start_time using pkt_time and time_granularity: Divide cache_start_time by the value of time_granularity, round up, and then multiply by time_granularity; Use pkt_time and cache_start_time to calculate the time difference pkt_delta_time. The value of pkt_delta_time is: pkt_time minus cache_start_time. Use pkt_delta_time and time_granularity to calculate the slot position pkt_slot. The value of pkt_slot is: pkt_delta_time is divided by the value of time_granularity and rounded up.

2. The method for merging and processing multi-source data streams in a cloud environment according to claim 1, characterized in that: Also includes: Step S4: When caching or outputting data packets, choose whether to sort the data packets according to demand.

3. The method for merging and processing multi-source data streams in a cloud environment according to claim 1, characterized in that: The time granularity is used to round the time of a data packet and file the data packet into a certain time slot.

4. The method for merging and processing multi-source data streams in a cloud environment according to claim 1, characterized in that: The calculation formula of the ring buffer queue is: slot_count = max_timespan / time_granularity; Among them, max_timespan is the maximum cache time span, and time_granularity is the time granularity.

5. The method for merging and processing multi-source data streams in a cloud environment according to claim 1, characterized in that: According to the value of pkt_slot, the data packets are divided into cached data packets, output data packets and discarded data packets, and then processed separately: Read pkt_slot, slot_count. If pkt_slot < 0, the data packet is classified as a discarded data packet for processing; If pkt_slot ≥ 0 and pkt_slot ≥ slot_count, the data packet is classified as an output data packet for processing; otherwise, the data packet is classified as a cached data packet for processing.

6. The method for merging and processing multi-source data streams in a cloud environment according to claim 5, characterized in that: The process of caching data packets includes: Read in pkt_slot, cache_start_slot, slot_count and cache_ring; Among them, cache_start_slot is the cache starting slot; Use pkt_slot, cache_start_slot, slot_count to calculate index. Index is actually the physical storage location of the slot in the ring cache cache_ring. The value of index is: (pkt_slot + cache_start_slot) % slot_count; Among them, % is the modulo remainder operation; Add the data packet to the array of cache_ring[index]. Use index to get the storage unit of cache_ring[index] and add the data packet to the array of the storage unit.

7. The method for merging and processing multi-source data streams in a cloud environment according to claim 6, characterized in that: The processing of outgoing packets includes: Read in pkt_slot, slot_count, cache_ring, cache_start_slot, pkt_time, and time_granularity; According to pkt_slot, slot_count and calculate export_count, the value of export_count is: (pkt_slot - slot_count + 1) % slot_count; Among them, % is the modulo remainder operation; Export export_count units of cache_ring starting from cache_slot_start; Update cache_start_slot according to cache_slot_star, export_count, and slot_count. The value of cache_start_slot is: (cache_slot_star + export_count) % slot_count; Update cache_start_time using pkt_time and time_granularity.

8. A system for merging and processing multi-source data streams in a cloud environment, the system executing the method for merging and processing multi-source data streams in a cloud environment as claimed in any one of claims 1 to 7, characterized in that: include: Module M1: Initialize a ring buffer queue according to the configuration parameters; The configuration parameters include maximum cache time and time granularity; The circular buffer queue is used to store data packets from multiple collection points; Module M2: Calculate the corresponding starting slot position according to the starting data packet, cache the starting data packet into the ring cache queue according to the slot, and record the cache start time; Module M3: Calculate the cache slot of the data packet according to the time of the new data packet and the cache start time, and cache the data packet into the corresponding slot; when the slot position is larger than the cache queue size, drive the output of the cached data packet and update the start slot position and the start data packet at the same time.

9. The system for merging and processing multi-source data streams in a cloud environment according to claim 8, characterized in that: Also includes: Module M4: When caching or outputting data packets, choose whether to sort the data packets according to the needs.

Citation Information

Patent Citations

  • Network surveillance method, apparatus program, circuit failure preventing system

    CN101369869A

  • Multi-sensor data acquisition method and system based on numerical control machine tool

    CN112230603A

  • UDP-based data transmission method and device, equipment and readable storage medium

    CN112929455A

  • Universal stream computing concurrent acceleration method, system, medium and equipment

    CN115981861A