Efficient statistical processing system and method for real-time data streams
By designing an efficient statistical processing system for real-time data flows, using processing node analysis module, data shunt module, shunt mode testing module and dynamic monitoring and adjustment module, the problem that existing systems cannot flexibly adjust shunt strategies and monitor data flow changes is solved, and processing efficiency and system performance are improved.
Patent Information
- Application Number
- CN202510518764.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The existing real-time data flow statistical processing system cannot flexibly adjust the diversion strategy according to actual conditions, nor can it monitor and dynamically adjust the changes in data flow, resulting in the impact of processing efficiency and system performance.
An efficient statistical processing system for real-time data flow is designed, including processing node analysis module, data shunt module, shunt mode testing module and dynamic monitoring and adjustment module. Through these modules, the system can flexibly adjust the shunt strategy according to the different capabilities of the processing nodes, and monitor and dynamically adjust the changes in data flow.
It realizes flexible adjustment of diversion strategies and dynamic adjustment of data flow changes according to actual conditions, improving the processing efficiency and system performance of real-time data flows.
Smart Images

Figure CN120067175A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data statistics, relates to data analysis technology, and specifically is an efficient statistical processing system and method for real-time data streams. Background Art
[0002] A real-time data stream is a continuous and high-speed arriving data sequence. In today's digital age, the generation and dissemination speed of data grow exponentially. The statistical processing of real-time data streams aims to analyze continuously and rapidly arriving data and extract valuable information. The processing of real-time data streams has become crucial in many fields.
[0003] However, when dealing with high-speed, continuous, and large-scale real-time data streams, many challenges are often faced. The generation speed of real-time data streams is extremely fast, and thousands or even tens of thousands of transaction data may be generated per second. The arrival order of data may be different from the generation order, and disorderly arriving data will increase the processing difficulty and lead to inaccurate statistical results.
[0004] Existing statistical processing systems for real-time data streams often do not fully consider the capacity differences of processing nodes, cannot flexibly adjust the shunt strategy according to the actual situation, nor can they monitor and dynamically adjust the changes in data streams, affecting the overall processing efficiency and system performance.
[0005] In view of the above technical problems, this application proposes a solution. Summary of the Invention
[0006] The purpose of the present invention is to provide an efficient statistical processing system and method for real-time data streams, which are used to solve the problems that existing real-time data stream statistical processing systems cannot flexibly adjust the shunt strategy according to the actual situation, nor can they monitor and dynamically adjust the changes in data streams. The technical problem to be solved by the present invention is: how to provide an efficient statistical processing method and system for real-time data streams that can flexibly adjust the shunt strategy according to the actual situation and can also monitor and dynamically adjust the changes in data streams.
[0007] The purpose of the present invention can be achieved by the following technical solutions: An efficient statistical processing system for real-time data streams includes a statistical analysis platform, which is communicatively connected with a processing node analysis module, a data shunt module, a shunt mode test module, and a dynamic monitoring and adjustment module. The processing node analysis module is used to analyze and evaluate the capabilities of the processing nodes of real-time data streams: mark the nodes that receive real-time data streams and perform statistical processing as processing nodes, and obtain the processing coefficient CL of the processing nodes. The data shunting module is used to perform shunting processing on real-time data streams in different modes: when using processing nodes to process real-time data streams, the real-time data streams are shunted to different processing nodes; the shunting rules are based on the processing coefficient CL of the processing nodes and are divided into three modes: priority mode, balanced mode, and hybrid mode. The shunting mode testing module is used to test and analyze different shunting modes: generate a number of test cycles Tn with a fixed duration of T seconds, where n represents the number of test cycles, n = 1, 2, ……, k, and k is a positive integer; obtain the processing mean CJ and processing floating value CF of different test cycles under different shunting modes, and determine the processing mode based on the processing mean CJ and processing floating value CF. The processing modes include efficiency priority mode and stability priority mode. Perform statistical processing on real-time data streams under different processing modes, generate a monitoring cycle with a fixed duration of L2 seconds, draw a line chart of the total data input SZ - monitoring cycle, and determine whether there is an increment in the real-time data stream, generate corresponding signals, and take corresponding measures for adjustment.
[0008] Furthermore, the process of obtaining the processing coefficient CL of the processing node includes: obtaining the processing speed CS, processing delay CY, and maximum throughput TT of the processing node, and inputting the processing speed CS, processing delay CY, and maximum throughput TT into a multi-layer perceptron model to obtain the processing coefficient CL of the processing node.
[0009] Furthermore, the processing speed CS represents the average number of bytes of data processed by the processing node within a preset time; the processing delay CY represents the average time delay from when the data enters the processing node to when the processing is completed and the result is output; the maximum throughput TT represents the maximum data flow that the processing node can process.
[0010] Furthermore, the shunting rule of the priority mode is specifically: mark the area for temporarily storing data in the processing node as a buffer, arrange the processing nodes in descending order of the processing coefficient CL to obtain a node priority sequence. When shunting the real-time data stream, give priority to the processing node with a smaller serial number in the node priority sequence and with remaining storage space in the buffer.
[0011] Furthermore, the shunting rule of the balanced mode is: when shunting the real-time data stream, generate a shunting cycle with a fixed duration of L1 seconds. The real-time data stream of the first shunting cycle is given to the first processing node in the node priority sequence, the real-time data stream of the second shunting cycle is given to the second processing node in the node priority sequence, ……, and so on. Until the last processing node is assigned the real-time data stream, the real-time data stream of the next shunting cycle is given to the first processing node in the node priority sequence, ……, and so on in a cycle.
[0012] Further, the shunt rule for the hybrid mode is as follows: a buffer threshold is set for the remaining space of the buffer of the processing node. When shunting the real-time data stream, the real-time data stream is first distributed to the first processing node in the node priority sequence until the remaining space of the buffer of this processing node reaches the buffer threshold. Then, the real-time data stream is distributed to the second processing node in the node priority sequence until the remaining space of the buffer of this processing node reaches the buffer threshold. Next, the real-time data stream is distributed to the third processing node in the node priority sequence, and so on, until the remaining space of the buffer of the last processing node in the node priority sequence reaches the buffer threshold. At this time, the real-time data stream is redistributed to the first processing node in the node priority sequence, and so on, in a cycle.
[0013] Further, the processes for obtaining the processing mean CJ and the processing floating value CF include: in the first test cycle T1, the fourth test cycle T4, …, and the (3m + 1)-th test cycle T3m+1, the real-time data stream is shunted in the priority mode; in the second test cycle T2, the fifth test cycle T5, …, and the (3m + 2)-th test cycle T3m+2, the real-time data stream is shunted in the balanced mode; in the third test cycle T3, the sixth test cycle T6, …, and the (3m + 3)-th test cycle T3m+3, the real-time data stream is shunted in the hybrid mode. At the end of the test cycle, the data processing amounts in different test cycles T3m+1 in the priority mode are obtained, and the sum of the data processing amounts is averaged to calculate the processing mean CJ in the priority mode, and the variance of the data processing amounts is calculated to obtain the processing floating value CF in the priority mode. Similarly, the data processing amounts in the balanced mode and the hybrid mode are obtained and numerically calculated to obtain the corresponding processing mean CJ and processing floating value CF.
[0014] Further, the process for determining the selection of the processing mode includes: if the management adopts the efficiency-first mode for statistical processing of the real-time data stream, the shunt mode with the largest processing mean CJ is adopted; if the management adopts the stability-first mode for statistical processing of the real-time data stream, the shunt mode with the smallest processing floating value CF is adopted.
[0015] Further, when performing statistical processing on the real-time data stream, a monitoring period with a fixed duration of L2 seconds is generated, and the total data input SZ within the monitoring period is obtained. A rectangular coordinate system is established with the total data input SZ as the Y-axis of the coordinate system and the monitoring period as the X-axis of the coordinate system, and a line graph of the total data input SZ - monitoring period is plotted. The slopes of the lines in the line graph of the total data input SZ - monitoring period are compared: the slope of any one line is set as G1, and the slope of the previous line is set as G0. The slope change rate GB is obtained through the formula GB = (G1 - G0) / G0. The slope change rate GB is compared with a preset slope change threshold GBmax: if the slope change rate GB is less than the slope change threshold GBmax, no processing is required; if the slope change rate GB is greater than or equal to the slope change threshold GBmax, it is determined that there is an increment in the real-time data stream, a traffic increase signal is generated and sent to the mobile terminal of the management personnel, and the management personnel take corresponding measures for dynamic adjustment. The adjustment measures include shortening the diversion period of the real-time data stream, changing the allocation mode, adjusting the buffer threshold in the hybrid mode, and increasing the number of processing nodes.
[0016] An efficient statistical processing method for real-time data streams, comprising the following steps: Step 1: Mark the node that receives and performs statistical processing on the real-time data stream as a processing node, perform numerical calculations on the processing speed CS, processing delay CY, and maximum throughput TT to obtain the processing coefficient CL of the processing node; Step 2: When using the processing node to process the real-time data stream, divert the real-time data stream to different processing nodes; the diversion rules are divided into three modes: priority mode, balanced mode, and hybrid mode; Step 3: Generate a number of test periods Tn with a fixed duration of T seconds, obtain the processing mean CJ and processing floating value CF of different test periods under different modes, and determine the processing mode according to the processing mean CJ and processing floating value CF. The processing modes include the efficiency priority mode and the stability priority mode; Step 4: Perform statistical processing on the real-time data stream under different processing modes, generate a monitoring period with a fixed duration of L2 seconds, plot a line graph of the total data input SZ - monitoring period, and determine whether there is an increment in the real-time data stream, generate corresponding signals, and take corresponding measures for adjustment.
[0017] The present invention has the following beneficial effects: 1. By comprehensively considering factors such as processing speed, processing delay, and maximum throughput through the processing node analysis module, the actual processing capacity of the processing node is accurately evaluated, providing a scientific basis for subsequent data diversion and improving resource utilization efficiency; 2. The data shunting module provides various shunting methods such as the priority mode, the balanced mode, and the hybrid mode, which can flexibly select the most suitable shunting strategy according to different business requirements and data characteristics, and achieve the balance between data processing efficiency and system stability; 3. The shunting mode test module conducts system tests and analyzes different shunting modes, and can provide clear basis for mode selection for managers in both the efficiency - first and stability - first modes, enabling the system to achieve the best processing effect in different scenarios; 4. The dynamic monitoring and adjustment module dynamically monitors the real - time data stream. When the total amount of data input changes significantly, it can send signals to managers in a timely manner and take corresponding measures to ensure that the system always adapts to the change of data flow and guarantee the timeliness and accuracy of processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 It is the system block diagram of Embodiment 1 of the present invention; Figure 2 It is the method flowchart of Embodiment 2 of the present invention; Figure 3 It is the line chart of the total amount of data input SZ - monitoring period in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The following will clearly and completely describe the technical solutions of the present invention in combination with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0021] Embodiment 1: As Figure 1 shown, an efficient statistical processing system for real - time data stream includes a statistical analysis platform, and the statistical analysis platform is communicatively connected with a processing node analysis module, a data shunting module, a shunting mode test module, and a dynamic monitoring and adjustment module; The processing node analysis module is used to analyze and evaluate the capabilities of the processing nodes for real-time data streams: Nodes that receive real-time data streams and perform statistical processing are marked as processing nodes, and the processing speed CS, processing latency CY, and maximum throughput TT of the processing nodes are obtained; The processing speed CS represents the average number of bytes of data processed by the processing node within a preset time; The processing latency CY represents the average time delay from when the data enters the processing node to when the processing is completed and the result is output; The maximum throughput TT represents the maximum data flow that the processing node can handle. Obtain the historical monitoring data of the processing nodes, and read the historical monitoring data from a CSV file, including fields: including the processing speed CS, processing latency CY, maximum throughput TT, and processing coefficient CL. Use 70% of the sample data in the historical monitoring data as the training set and 30% of the sample data as the validation set. Construct a multi-layer perceptron (MLP) model, input the sample data in the training set into the model for model training, and optimize the parameters of the model by minimizing the loss function. Among them, the multi-layer perceptron (MLP) model includes an input layer (processing speed CS, processing latency CY, and maximum throughput TT), a hidden layer, and an output layer (processing coefficient CL). Then, verify the constructed multi-layer perceptron neural network model after training through the validation set. Use the trained model to predict the processing speed CS, processing latency CY, and maximum throughput TT of the current processing node to obtain the processing coefficient CL. Among them, using the processing speed CS, processing latency CY, and maximum throughput TT of each processing node as source data, sample data (CSi, CYi, TTi) of each processing node are constructed, where CSi represents the processing speed of the i-th processing node in the historical data, CYi represents the processing latency of the i-th processing node in the historical data, and TTi represents the maximum throughput of the i-th processing node in the historical data.
[0022] Through the processing node analysis module, comprehensively considering factors such as processing speed, processing latency, and maximum throughput, accurately evaluate the actual processing capabilities of the processing nodes, provide a scientific basis for subsequent data shunting, and improve resource utilization efficiency.
[0023] The data shunting module is used to perform shunting processing on real-time data streams in different modes: When using processing nodes to process real-time data streams, the real-time data streams are shunted to different processing nodes; The shunting rules are based on the processing coefficient CL of the processing nodes and are divided into three modes: priority mode, balanced mode, and hybrid mode. The shunting rules for the priority mode are specifically as follows: Mark the area for temporarily storing data in the processing node as the buffer. Arrange the processing nodes in descending order of the processing coefficient CL to obtain the node priority sequence. When shunting the real-time data stream, give priority to the processing nodes with a smaller serial number in the node priority sequence and with remaining storage space in the buffer. The shunting rules for the balanced mode are as follows: When shunting the real-time data stream, generate a shunting period with a fixed duration of L1 seconds. The real-time data stream in the first shunting period is given to the first processing node in the node priority sequence, the real-time data stream in the second shunting period is given to the second processing node in the node priority sequence, and so on. Until the last processing node is allocated the real-time data stream, the real-time data stream in the next shunting period is given to the first processing node in the node priority sequence, and so on, repeating in a cycle. The shunting rules for the hybrid mode are as follows: Set a buffer threshold for the remaining space in the buffer of the processing node. When shunting the real-time data stream, first give the real-time data stream to the first processing node in the node priority sequence. Until the remaining space in the buffer of this processing node reaches the buffer threshold, give the real-time data stream to the second processing node in the node priority sequence. Until the remaining space in the buffer of this processing node reaches the buffer threshold, give the real-time data stream to the third processing node in the node priority sequence, and so on. Until the remaining space in the buffer of the last processing node in the node priority sequence reaches the buffer threshold, give the real-time data stream back to the first processing node in the node priority sequence, and so on, repeating in a cycle. Through the data shunting module, multiple shunting methods such as the priority mode, the balanced mode, and the hybrid mode are provided, which can flexibly select the most suitable shunting strategy according to different service requirements and data characteristics, and achieve the balance between data processing efficiency and system stability.
[0024] The shunting mode test module is used to test and analyze different shunting modes: Before the processing node performs statistical processing on the real-time data stream, generate several test periods Tn with a fixed duration of T seconds, where n represents the number of test periods, n = 1, 2,..., k, and k is a positive integer. The first test period T1, the fourth test period T4,..., the (3m + 1)-th test period T3m+1 all use the priority mode for shunting the real-time data stream. The second test period T2, the fifth test period T5,..., the (3m + 2)-th test period T3m+2 all use the balanced mode for shunting the real-time data stream. The third test period T3, the sixth test period T6,..., the (3m + 3)-th test period T3m+3 all use the hybrid mode for shunting the real-time data stream. At the end of the test cycle, obtain the data processing volume of different test cycles T3m+1 in the priority mode, sum up the data processing volume and calculate the average value to obtain the processing mean CJ in the priority mode, and calculate the variance of the data processing volume to obtain the processing floating value CF in the priority mode; similarly, obtain the data processing volume in the balanced mode and the hybrid mode and perform numerical calculations to obtain the corresponding processing mean CJ and processing floating value CF; when the processing node performs statistical processing on the real-time data stream, the processing modes include the efficiency priority mode and the stability priority mode; if the management personnel adopt the efficiency priority mode to perform statistical processing on the real-time data stream, then adopt the shunt mode with the largest processing mean CJ; if the management personnel adopt the stability priority mode to perform statistical processing on the real-time data stream, then adopt the shunt mode with the smallest processing floating value CF; by testing and analyzing different shunt modes through the shunt mode test module, it is possible to provide a clear basis for mode selection for the management personnel under both the efficiency priority and stability priority modes, so that the system can achieve the best processing effect in different scenarios.
[0025] As Figure 3 shown, the dynamic monitoring and adjustment module is used to dynamically adjust the statistical processing according to the monitoring data of the real-time data stream: when performing statistical processing on the real-time data stream, generate a monitoring cycle with a fixed duration of L2 seconds, obtain the total data input volume SZ within the monitoring cycle, establish a rectangular coordinate system with the total data input volume SZ as the Y-axis of the coordinate system and the monitoring cycle as the X-axis of the coordinate system, and draw a line graph of the total data input volume SZ - monitoring cycle; compare the slopes of the lines in the line graph of the total data input volume SZ - monitoring cycle: set the slope of any line as G1, and the slope of the previous line as G0, and obtain the slope change rate GB through the formula GB = (G1 - G0) / G0; compare the slope change rate GB with the preset slope change threshold GBmax: if the slope change rate GB is less than the slope change threshold GBmax, no processing is required; if the slope change rate GB is greater than or equal to the slope change threshold GBmax, it is determined that there is an increment in the real-time data stream, generate a traffic increase signal and send the signal to the mobile terminal of the management personnel, and the management personnel take corresponding measures for dynamic adjustment. The adjustment measures include shortening the shunt cycle of the real-time data stream, changing the allocation mode, adjusting the buffer threshold in the hybrid mode, and increasing the number of processing nodes; by dynamically monitoring the real-time data stream through the dynamic monitoring and adjustment module, when the total data input volume changes significantly, it can timely send a signal to the management personnel and take corresponding measures to ensure that the system always adapts to the change of the data flow and guarantee the timeliness and accuracy of the processing.
[0026] Embodiment 2: As Figure 2 shown, an efficient statistical processing method for real-time data streams includes the following steps: Step 1: Mark the node that receives the real-time data stream and performs statistical processing as the processing node, calculate the numerical values of the processing speed CS, processing delay CY, and maximum throughput TT, and obtain the processing coefficient CL of the processing node. Step 2: When using the processing node to process the real-time data stream, split the real-time data stream into different processing nodes; the splitting rules are divided into three modes: priority mode, balanced mode, and hybrid mode. Step 3: Generate several test cycles Tn with a fixed duration of T seconds, obtain the processing mean CJ and processing floating value CF of different test cycles under different modes, and determine the processing mode according to the processing mean CJ and processing floating value CF. The processing modes include efficiency priority mode and stability priority mode. Step 4: Generate a monitoring cycle with a fixed duration of L2 seconds, draw a line chart of the total data input SZ - monitoring cycle, and determine whether there is an increment in the real-time data stream, generate corresponding signals, and take corresponding measures for adjustment.
[0027] An efficient statistical processing system for real-time data streams. When working, mark the node that receives the real-time data stream and performs statistical processing as the processing node, obtain the processing coefficient CL of the processing node, and split the real-time data stream into different processing nodes according to the processing coefficient CL. The splitting rules include priority mode, balanced mode, and hybrid mode; generate several test cycles Tn with a fixed duration of T seconds, obtain the processing mean CJ and processing floating value CF of different test cycles under different modes, and determine the processing mode according to the processing mean CJ and processing floating value CF. The processing modes include efficiency priority mode and stability priority mode; generate a monitoring cycle with a fixed duration of L2 seconds, draw a line chart of the total data input SZ - monitoring cycle, and determine whether there is an increment in the real-time data stream, generate corresponding signals, and take corresponding measures for adjustment.
[0028] The above content is only an example and explanation of the structure of the present invention. Those skilled in the art of this technology make various modifications or supplements to the described specific embodiments or use similar methods for substitution. As long as they do not deviate from the structure of the invention or exceed the scope defined by this claim book, they should all fall within the protection scope of the present invention.
[0029] In the description of this specification, the description with reference to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0030] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation manners. Obviously, according to the content of this specification, many modifications and variations can be made. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. An efficient statistical processing system for real-time data streams, characterized in that: include: The processing node analysis module is used to analyze and evaluate the capabilities of the processing nodes of the real-time data stream: marking the nodes that receive the real-time data stream and perform statistical processing as processing nodes, and obtaining the processing coefficient CL of the processing node; The data diversion module is used to divert real-time data streams in different modes: When a processing node is used to process a real-time data stream, the real-time data stream is split to different processing nodes; The traffic diversion rules are divided into priority mode, balance mode and mixed mode based on the processing coefficient CL of the processing node; The diversion mode test module is used to test and analyze different diversion modes: generate a number of test cycles Tn with a fixed duration of T seconds; Obtain the processing mean CJ and processing floating value CF of different test cycles under different diversion modes, and determine the processing mode according to the processing mean CJ and processing floating value CF; The processing modes include efficiency priority mode and stability priority mode; Under different processing modes, the real-time data stream is statistically processed to generate a monitoring period with a fixed duration of L2 seconds. A line graph of the total data input SZ-monitoring period is drawn to determine whether there is an increment in the real-time data stream, generate corresponding signals and take corresponding measures to make adjustments.
2. The efficient statistical processing system for real-time data stream according to claim 1, characterized in that: The process of obtaining the processing coefficient CL of the processing node includes: The processing speed CS, processing delay CY and maximum throughput TT of the processing node are obtained, and the processing speed CS, processing delay CY and maximum throughput TT are input into the multi-layer perceptron model to obtain the processing coefficient CL of the processing node.
3. The efficient statistical processing system for real-time data stream according to claim 2, characterized in that: The processing speed CS represents the average number of bytes of data processed by the processing node within a preset time; Processing delay CY represents the average time delay from when data enters a processing node to when the processing is completed and the result is output; The maximum throughput TT represents the maximum data flow that a processing node can handle.
4. The efficient statistical processing system for real-time data stream according to claim 1, characterized in that: The specific diversion rules of the priority mode are as follows: Mark the area in the processing node used for temporarily storing data as a buffer, and arrange the processing nodes in descending order of processing coefficient CL to obtain a node priority sequence; When real-time data streams are diverted, priority is given to processing nodes with smaller sequence numbers in the node priority sequence and with remaining storage space in the buffer.
5. The efficient statistical processing system for real-time data stream according to claim 1, characterized in that: The traffic distribution rules of the balanced mode are: When real-time data streams are diverted, a diversion cycle with a fixed duration of L1 seconds is generated. The real-time data stream of the first diversion cycle is distributed to the first processing node in the node priority sequence, and the real-time data stream of the second diversion cycle is distributed to the second processing node in the node priority sequence, and so on. After the last processing node is assigned a real-time data stream, the real-time data stream of the next diversion cycle is distributed to the first processing node in the node priority sequence, and so on. The cycle repeats.
6. The efficient statistical processing system for real-time data stream according to claim 1, characterized in that: The traffic diversion rules of the hybrid mode are: A buffer threshold is set for the remaining space in the buffer of the processing node. When the real-time data stream is diverted, the real-time data stream is first diverted to the first processing node in the node priority sequence until the remaining space in the buffer of the processing node reaches the buffer threshold. The real-time data stream is then diverted to the second processing node in the node priority sequence. The real-time data stream is then diverted to the third processing node in the node priority sequence until the remaining space in the buffer of the processing node reaches the buffer threshold. The real-time data stream is then diverted to the first processing node in the node priority sequence again, and the cycle repeats.
7. The efficient statistical processing system for real-time data stream according to claim 1, characterized in that: The process of obtaining the processing mean CJ and the processing floating value CF includes: the first test cycle T1, the fourth test cycle T4, ..., the 3m+1 test cycle T3m+1 all adopt the priority mode to divert the real-time data flow; the second test cycle T2, the fifth test cycle T5, ..., the 3m+2 test cycle T3m+2 all adopt the balanced mode to divert the real-time data flow; the third test cycle T3, the sixth test cycle T6, ..., the 3m+3 test cycle T3m+3 all adopt the mixed mode to divert the real-time data flow; at the end of the test cycle, the data processing amount of different test cycles T3m+1 under the priority mode is obtained, and the data processing amount is summed and averaged to calculate the processing mean CJ under the priority mode; The variance of the data processing amount is calculated to obtain the processing floating value CF in the priority mode; The data processing amount in the balanced mode and the mixed mode is obtained and numerical calculation is performed to obtain the corresponding processing mean CJ and processing floating value CF.
8. The efficient statistical processing system for real-time data stream according to claim 7, characterized in that: The selection and determination process of the processing mode includes: If the management adopts the efficiency-first mode to statistically process the real-time data stream, the diversion mode with the largest processing mean CJ is adopted; If the management personnel adopt the stability priority mode to perform statistical processing on the real-time data stream, the diversion mode with the smallest processing floating value CF will be adopted.
9. The efficient statistical processing system for real-time data stream according to claim 8, characterized in that: When statistically processing the real-time data stream, a monitoring period with a fixed duration of L2 seconds is generated, the total amount of data input SZ within the monitoring period is obtained, a rectangular coordinate system is established with the total amount of data input SZ as the Y axis of the coordinate system and the monitoring period as the X axis of the coordinate system, and a line graph of the total amount of data input SZ-monitoring period is drawn; Compare the slopes of the lines in the data input total SZ-monitoring period line chart: Set the slope of any broken line to G1, and the slope of the previous broken line to G0, and obtain the slope change rate GB by the formula GB=(G1-G0) / G0; compare the slope change rate GB with the preset slope change threshold GBmax: If the slope change rate GB is less than the slope change threshold GBmax, no processing is required; If the slope change rate GB is greater than or equal to the slope change threshold GBmax, it is determined that there is an increment in the real-time data flow, and a traffic increase signal is generated and sent to the administrator's mobile terminal. The administrator takes corresponding measures to make dynamic adjustments. The adjustment measures include shortening the diversion cycle of the real-time data flow, changing the allocation mode, adjusting the buffer threshold in the mixed mode, and increasing the number of processing nodes.
10. An efficient statistical processing method for real-time data stream, characterized in that: The following steps are involved: Step 1: Mark the nodes that receive the real-time data stream and perform statistical processing as processing nodes, perform numerical calculations on the processing speed CS, processing delay CY, and maximum throughput TT, and obtain the processing coefficient CL of the processing node; Step 2: When the real-time data stream is processed by the processing node, the real-time data stream is diverted to different processing nodes; the diversion rules are divided into three modes: priority mode, balance mode and mixed mode; Step 3: Generate several test cycles Tn with a fixed duration of T seconds, obtain the processing mean CJ and processing floating value CF of different test cycles under different modes, and determine the processing mode according to the processing mean CJ and processing floating value CF. The processing modes include efficiency priority mode and stability priority mode; Step 4: Perform statistical processing on the real-time data stream under different processing modes, generate a monitoring cycle with a fixed duration of L2 seconds, draw a line graph of the total data input SZ-monitoring cycle, and determine whether there is an increment in the real-time data stream, generate corresponding signals and take corresponding measures to make adjustments.
Citation Information
Patent Citations
Network analytic system and method supporting real-time mass data processing
CN103560943A
A marine observation big data visualization analysis method based on a complex network
CN109947879A
Data processing method and device based on docking station, electronic equipment and medium
CN117591850A
Telemetry data collection and distribution method and system
CN118540227A
Optical power monitoring system and method applied to optical fiber
CN118890092A