Big Data-Based Transmission Efficiency Improvement System
By improving the system's big data transmission efficiency and dynamically adjusting the data transmission rate and network settings, the problem of low transmission efficiency in wireless environments has been solved, achieving efficient and stable big data transmission while reducing energy consumption and maintenance costs.
Patent Information
- Application Number
- CN202511326528.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-17
AI Technical Summary
In wireless environments with large data volumes and low latency scenarios, traditional layered protocols suffer from systemic bottlenecks, resulting in low transmission efficiency, long response times, and insufficient compatibility of existing technologies in complex network environments, leading to resource mismatch and increased energy consumption.
A big data-based transmission performance improvement system is adopted. The system periodically collects network information through a parameter acquisition module, and combines it with a performance evaluation module, a classification module, and a transmission optimization module to dynamically adjust the data transmission rate. It utilizes QUIC/RDMA technology, Avro serialization and Snappy compression algorithm, backpressure mechanism, and asynchronous checkpoint technology to achieve network adjustment and fault-tolerant recovery.
In wireless, high-concurrency, and low-latency scenarios, it achieves high-throughput, low-power, and highly reliable transmission, reduces resource waste, improves transmission stability and risk resistance, and lowers operation and maintenance costs.
Smart Images

Figure CN120825458B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data transmission technology, and in particular to a system for improving transmission efficiency based on big data. Background Technology
[0002] Traditional layered protocols face systemic bottlenecks in wireless environments, high-volume data, and low-latency scenarios, necessitating global optimization through cross-layer collaboration (such as QUIC / RDMA), data processing innovation (Arrow / Avro), and fine-grained resource management (backpressure / memory pools). Current trends have shifted from single-layer protocol improvements to full-stack reconstruction and dynamic parameter tuning to further enhance adaptability. At the protocol level, while QUIC reduces connection overhead, it suffers from insufficient compatibility in complex network environments and high adaptation costs with some traditional devices. SDN dynamic routing relies on accurate real-time data, which can lead to decision delays under sudden traffic surges, increasing short-term energy consumption. Regarding intelligent scheduling, AI load prediction models require extensive historical data training, resulting in increased error rates in data-sparse edge scenarios, potentially leading to resource misallocation. While the distributed architecture of edge computing reduces transmission volume, it increases energy consumption and complexity in inter-node collaboration. Furthermore, green energy scheduling is limited by geography and weather; in areas with unstable power grids, power supply fluctuations may actually reduce transmission efficiency. Summary of the Invention
[0003] To address this, the present invention provides a transmission performance improvement system based on big data, which overcomes the problems of low transmission efficiency and long response time caused by the systemic bottlenecks faced by traditional layered protocols in wireless environments, large data volumes, and low-latency scenarios.
[0004] To achieve the above objectives, the preferred technical solution for a big data-based transmission performance improvement system includes:
[0005] The parameter acquisition module is used to periodically collect the transmission rate of each network, the bandwidth usage information at different time periods, and the transmission delay information. The transmission rate includes the real-time transmission rate and the average transmission rate. The bandwidth usage information includes the usage ratio of each network at different time periods. The transmission delay information includes the round-trip delay and processing delay during data transmission.
[0006] The performance evaluation module, which is connected to the parameter acquisition module, is used to evaluate the performance level of each network bandwidth based on transmission rate, bandwidth occupancy information and transmission delay information.
[0007] The classification module, which is connected to the performance evaluation module and the parameter acquisition module, is used to classify time periods within the monitoring period according to bandwidth usage information, and to classify each network according to network stability and latency fluctuation.
[0008] The transmission optimization module, which is connected to the classification module and the performance evaluation module, is used to determine the network adjustment method in different time periods based on the time period classification results and network classification results, combined with the QUIC / RDMA technology of the protocol layer, the Avro serialization and Snappy compression algorithm of the data layer, the backpressure mechanism of the flow control layer and the asynchronous checkpoint technology of the fault tolerance layer. The network adjustment method includes network replacement strategy and transmission parameter adjustment scheme.
[0009] The dynamic adjustment module, which is connected to the transmission optimization module, is used to dynamically adjust the data transmission rate during transmission through a backpressure mechanism, achieve transmission fault tolerance and fast recovery based on asynchronous checkpoint technology, and optimize cross-system data interaction using a vectorized format.
[0010] As a preferred technical solution for a big data-based transmission efficiency improvement system, the periodic collection period is a fixed duration, and the occupancy ratio in the bandwidth occupancy information includes peak occupancy ratio and continuous occupancy ratio.
[0011] As a preferred technical solution for a big data-based transmission performance improvement system, the evaluation of the performance level of each network bandwidth includes:
[0012] If the transmission rate is at a high level, the bandwidth utilization rate is at a low level, and the transmission delay is at a low level, then the performance level is high efficiency.
[0013] If the transmission rate is at a medium level, the bandwidth utilization rate is at a medium level, and the transmission latency is at a medium level, then the performance level is medium.
[0014] If the transmission rate is at a low level, the bandwidth utilization rate is at a high level, or the transmission delay is at a high level, then the performance level is inefficient.
[0015] As a preferred technical solution for a big data-based transmission efficiency improvement system, the time periods within the monitoring period are categorized as follows:
[0016] High-frequency usage periods are those during which the peak bandwidth occupancy rate remains high for several consecutive collection cycles.
[0017] Low-frequency usage periods are those during which the peak bandwidth occupancy rate remains low for several consecutive collection cycles.
[0018] As a preferred technical solution for improving transmission efficiency based on big data, the networks are classified as follows:
[0019] A high-quality network is one in which transmission rate fluctuations are within a preset fluctuation range and latency fluctuations are within an allowable fluctuation range.
[0020] A low-quality network is a network whose transmission rate fluctuations exceed the preset fluctuation range or whose latency fluctuations exceed the allowable fluctuation range.
[0021] As a preferred technical solution for improving transmission efficiency based on big data, the protocol layer adopts QUIC / RDMA technology, including:
[0022] During the high-frequency usage period, select the high-quality network and enable RDMA technology;
[0023] When the frequency of use is low or the network is of low quality, QUIC technology is used to improve transmission stability.
[0024] As a preferred technical solution for improving the transmission efficiency of a big data-based system, in step S3, the data layer employs Avro serialization and Snappy compression algorithms, including:
[0025] Structured data is uniformly encoded using the Avro serialization format;
[0026] For serialized data, the Snappy compression strength is selected according to the data size. High-strength compression is used for large-capacity data, and basic compression is used for small-capacity data.
[0027] As a preferred technical solution for a big data-based transmission performance improvement system, the flow control layer employs a backpressure mechanism, including:
[0028] The occupancy of the receiving end's buffer is monitored in real time. When the occupancy rate is high, the backpressure signal is triggered to reduce the transmission rate of the sending end.
[0029] When the receiver's buffer occupancy is low, the backpressure signal is released to restore the transmitter's initial transmission rate.
[0030] As a preferred technical solution for improving the transmission efficiency of a big data-based system, the fault-tolerant layer employs asynchronous checkpointing technology, including:
[0031] During the high-frequency usage period, an asynchronous checkpoint is generated at a first interval;
[0032] During the low-frequency usage period, an asynchronous checkpoint is generated at a second interval, and the checkpoint data is stored in a vectorized format.
[0033] As a preferred technical solution for a big data-based transmission performance improvement system, the network replacement strategy includes:
[0034] When the current network is a low-quality network and is in a period of high-frequency use, automatically switch to a backup high-quality network;
[0035] After the switch, core data processed in vectorized format will be transmitted first.
[0036] The beneficial effects of this invention are as follows:
[0037] This system maps network time periods, link quality, data characteristics, and protocol stack capabilities in real time. It first uses periodic large-scale data analysis for assessment, then dynamically selects Avro / Snappy compression at the data layer based on capacity. The flow control layer uses backpressure to match transmit and receive buffers in real time, and the fault tolerance layer uses asynchronous vectorized checkpoints for second-level recovery. When the network degrades, it automatically switches to backup links and prioritizes sending core data. Thus, in wireless, high-concurrency, and low-latency scenarios, it minimizes link latency and bandwidth waste while avoiding receiver overload and downtime, achieving high throughput, low power consumption, high reliability, and maintenance-free operation for large data transmission.
[0038] In particular, by dynamically adapting protocol layer technology, optimizing data encoding and compression strategies, and introducing flow control and fault tolerance mechanisms, the system can simultaneously solve problems such as low transmission rates, unreasonable bandwidth usage, and large latency fluctuations. Whether in high-frequency load periods or scenarios with poor network quality, the system can achieve efficient transmission through a combination of technologies. This ensures the high-speed advantage under high-quality networks while improving transmission stability under low-quality networks or complex periods, comprehensively improving the overall performance of big data transmission.
[0039] In particular, by classifying time periods and network quality, dynamically adjusting compression intensity based on data capacity, generating checkpoints on demand, and regulating traffic using a backpressure mechanism, the system achieves precise allocation of bandwidth, computing, and storage resources. High-frequency periods focus on efficient transmission to reduce resource waste, while low-frequency periods optimize resource utilization to reduce energy consumption. Simultaneously, vectorized formats improve data interaction efficiency, reducing resource consumption during transmission from multiple dimensions and balancing performance and cost control.
[0040] In particular, by leveraging periodic parameter collection, automatic classification, and dynamic strategy adjustment, the system can adapt to different network environments and load changes without manual intervention. Features such as real-time adjustment of the backpressure mechanism and automatic adaptation of checkpoint intervals further enhance the system's responsiveness to dynamic scenarios, reduce manual maintenance costs, and make the transmission process more closely aligned with the fluctuations in actual business needs.
[0041] In particular, the asynchronous checkpointing technology enables rapid fault recovery, and the network switching strategy automatically switches to the backup network in case of anomalies, prioritizing the transmission of core data. This significantly enhances the system's resilience. Whether it's a sudden network failure or a data transmission interruption, the fault-tolerance mechanism minimizes data loss, and the critical data priority strategy ensures that core business processes remain unaffected, providing comprehensive protection for the continuity and security of large-scale data transmission. Attached Figure Description
[0042] Figure 1 This is a schematic diagram of the transmission efficiency improvement system based on big data according to an embodiment of the present invention;
[0043] Figure 2 This is a logic diagram for classifying networks according to an embodiment of the present invention;
[0044] Figure 3 This is a logic diagram of the fault-tolerant layer using asynchronous checkpointing technology in an embodiment of the present invention. Detailed Implementation
[0045] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0046] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0047] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.
[0048] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0049] Please see Figure 1 The diagram shown is a structural schematic of a big data-based transmission performance improvement system according to an embodiment of the present invention. The present invention provides a big data-based transmission performance improvement system, comprising:
[0050] The parameter acquisition module is used to periodically collect the transmission rate of each network, the bandwidth usage information at different time periods, and the transmission delay information. The transmission rate includes the real-time transmission rate and the average transmission rate. The bandwidth usage information includes the usage ratio of each network at different time periods. The transmission delay information includes the round-trip delay and processing delay during data transmission.
[0051] The performance evaluation module, which is connected to the parameter acquisition module, is used to evaluate the performance level of each network bandwidth based on transmission rate, bandwidth occupancy information and transmission delay information.
[0052] The classification module, which is connected to the performance evaluation module and the parameter acquisition module, is used to classify time periods within the monitoring period according to bandwidth usage information, and to classify each network according to network stability and latency fluctuation.
[0053] The transmission optimization module, which is connected to the classification module and the performance evaluation module, is used to determine the network adjustment method in different time periods based on the time period classification results and network classification results, combined with the QUIC / RDMA technology of the protocol layer, the Avro serialization and Snappy compression algorithm of the data layer, the backpressure mechanism of the flow control layer and the asynchronous checkpoint technology of the fault tolerance layer. The network adjustment method includes network replacement strategy and transmission parameter adjustment scheme.
[0054] The dynamic adjustment module, which is connected to the transmission optimization module, is used to dynamically adjust the data transmission rate during transmission through a backpressure mechanism, achieve transmission fault tolerance and fast recovery based on asynchronous checkpoint technology, and optimize cross-system data interaction using a vectorized format.
[0055] This system maps network time periods, link quality, data characteristics, and protocol stack capabilities in real time. It first uses periodic large-scale data analysis for assessment, then dynamically selects Avro / Snappy compression at the data layer based on capacity. The flow control layer uses backpressure to match transmit and receive buffers in real time, and the fault tolerance layer uses asynchronous vectorized checkpoints for second-level recovery. When the network degrades, it automatically switches to backup links and prioritizes sending core data. Thus, in wireless, high-concurrency, and low-latency scenarios, it minimizes link latency and bandwidth waste while avoiding receiver overload and downtime, achieving high throughput, low power consumption, high reliability, and maintenance-free operation for large data transmission.
[0056] Specifically, the periodic collection period is a fixed duration, and the occupancy ratio in the bandwidth occupancy information includes the peak occupancy ratio and the continuous occupancy ratio.
[0057] During implementation, a fixed collection period (e.g., 5 minutes / time) is set to ensure uniform time granularity of data collection; within each period, the peak and continuous occupancy ratios of network bandwidth, the ratio of average occupancy to total bandwidth within the period are recorded, and the two types of ratio data are associated with the corresponding time periods to provide basic data support for subsequent performance evaluation and time period classification.
[0058] The optimal acquisition cycle is 5 minutes per acquisition. Short-term fluctuations in network status (bandwidth usage, speed) have a certain degree of continuity. A 5-minute cycle can capture network changes in a timely manner (avoiding lag caused by an excessively long cycle) without increasing the system's acquisition and calculation overhead due to excessive frequency (such as within 1 minute), thus balancing real-time performance and resource consumption.
[0059] Specifically, the evaluation of the performance level of each network bandwidth includes:
[0060] If the transmission rate is at a high level, the bandwidth utilization rate is at a low level, and the transmission delay is at a low level, then the performance level is high efficiency.
[0061] If the transmission rate is at a medium level, the bandwidth utilization rate is at a medium level, and the transmission latency is at a medium level, then the performance level is medium.
[0062] If the transmission rate is at a low level, the bandwidth utilization rate is at a high level, or the transmission delay is at a high level, then the performance level is inefficient.
[0063] The system presets three threshold levels (high, medium, and low) for transmission rate, bandwidth utilization, and transmission delay. It compares the collected network parameters and determines whether the network is efficient if the transmission rate is higher than the high threshold, the bandwidth utilization is lower than the low threshold, and the transmission delay is lower than the low threshold. If all three are within the medium threshold range, the network is determined to be medium. If the transmission rate is lower than the low threshold, the bandwidth utilization is higher than the high threshold, or the transmission delay is higher than the high threshold, the network is determined to be inefficient. This is how the system quantitatively evaluates the bandwidth efficiency of each network.
[0064] In this implementation, the transmission rates are: High-level ≥10Gbps, Medium-level 1-10Gbps, and Low-level <1Gbps.
[0065] Bandwidth usage: High ≥80%, Medium 30%-80%, Low <30%;
[0066] Transmission latency: High ≥100ms, Medium 10-100ms, Low <10ms.
[0067] It is understandable that modern data center network speeds are mostly 10Gbps and above, bandwidth usage of more than 80% is considered congestion, latency of less than 10ms is considered low latency (such as local area network), and latency of more than 100ms is considered high latency (such as cross-regional transmission). This three-level classification can clearly distinguish the performance level.
[0068] Specifically, classifying the time periods within the monitoring period includes:
[0069] High-frequency usage periods are those during which the peak bandwidth occupancy rate remains high for several consecutive collection cycles.
[0070] Low-frequency usage periods are those during which the peak bandwidth occupancy rate remains low for several consecutive collection cycles.
[0071] In implementation, high and low thresholds for peak bandwidth occupancy ratios and the number of consecutive judgment periods are preset. For each sub-period within the monitoring period, the performance of peak bandwidth occupancy ratios within the continuous collection period is statistically analyzed. If the peak bandwidth occupancy ratio continuously reaches or exceeds the high threshold, it is classified as a high-frequency usage period; if it is continuously at or below the low threshold, it is classified as a low-frequency usage period, thus realizing the time period division based on actual load.
[0072] In this embodiment, consecutive cycles are preferred. It is understood that the high / low bandwidth usage of 1-2 cycles may be occasional fluctuations (such as burst data transmission), while 3 consecutive cycles (such as 5 minutes × 3 = 15 minutes) can reflect a stable trend and avoid misjudging high frequency or low frequency.
[0073] Please see Figure 2 As shown, this is a logic diagram for classifying networks according to an embodiment of the present invention. Classifying each network includes:
[0074] A high-quality network is one in which transmission rate fluctuations are within a preset fluctuation range and latency fluctuations are within an allowable fluctuation range.
[0075] A low-quality network is a network whose transmission rate fluctuations exceed the preset fluctuation range or whose latency fluctuations exceed the allowable fluctuation range.
[0076] In implementation, allowable ranges for transmission rate fluctuations and delay fluctuations are preset; the transmission rate and delay changes of each network are monitored in real time, and the fluctuation amplitude, such as the ratio of standard deviation to mean, is calculated; if both fluctuation amplitudes are within the allowable range, the network is judged as a high-quality network; if either fluctuation amplitude exceeds the corresponding range, the network is judged as a low-quality network, thus completing the stability-based network classification.
[0077] The fluctuation range, i.e., the coefficient of variation ≤ 20%, indicates that the rate or latency fluctuates drastically, such as sudden packet loss or a sudden drop in bandwidth, which does not meet the stability requirements of a high-quality network. This threshold can effectively distinguish network stability.
[0078] In this invention, the system achieves precise allocation of bandwidth, computing, and storage resources by classifying time periods and network quality, dynamically adjusting compression intensity based on data capacity, generating checkpoints on demand, and regulating traffic using a backpressure mechanism. High-frequency periods focus on efficient transmission to reduce resource waste, while low-frequency periods optimize resource utilization to reduce energy consumption. Simultaneously, vectorized formats improve data interaction efficiency, reducing resource consumption during transmission from multiple dimensions and balancing performance and cost control.
[0079] Specifically, the protocol layer employs QUIC / RDMA technology, including:
[0080] During the high-frequency usage period, select the high-quality network and enable RDMA technology;
[0081] When the frequency of use is low or the network is of low quality, QUIC technology is used to improve transmission stability.
[0082] In practice, during high-frequency usage periods and when the current network is of high quality, RDMA technology is enabled to improve the transmission rate by directly accessing memory by bypassing the operating system kernel; during low-frequency usage periods or when the network is of low quality, the QUIC protocol is switched to, utilizing its multiplexing, connection migration, and strong error correction capabilities to ensure transmission stability and achieve dynamic adaptation of protocol layer technologies.
[0083] For example, please refer to Table 1 to see the performance optimization effects brought about by the system's dynamic adaptation strategy in different scenarios:
[0084] Table 1 Comparison of Optimization Effects of Dynamic Adaptation Strategies
[0085]
[0086] This implementation system can dynamically combine the technology stack according to the real-time scenario, and can provide significantly better performance than traditional fixed strategies under various conditions.
[0087] In this invention, by dynamically adapting protocol layer technology (QUIC / RDMA), optimizing data encoding and compression strategies (Avro / Snappy), and introducing flow control and fault tolerance mechanisms, the system can simultaneously solve problems such as low transmission rate, unreasonable bandwidth usage, and large latency fluctuations. Whether during high-frequency load periods or in scenarios with poor network quality, efficient transmission can be achieved through this technological combination. It ensures the high-speed advantage under high-quality networks while improving transmission stability under low-quality networks or complex time periods, comprehensively improving the overall performance of big data transmission.
[0088] Specifically, the data layer employs Avro serialization and Snappy compression algorithms, including:
[0089] Structured data is uniformly encoded using the Avro serialization format;
[0090] For serialized data, the Snappy compression strength is selected according to the data size. High-strength compression is used for large-capacity data, and basic compression is used for small-capacity data.
[0091] In implementation, all structured data is serialized using the Avro format. A unified schema is defined to describe the data structure, ensuring cross-system parsing compatibility. For the serialized data, the Snappy compression strength is selected according to a preset capacity threshold. Large data exceeding the threshold is compressed with high intensity, such as the highest compression level, while small data within the threshold is compressed with basic compression, balancing compression efficiency and computational overhead.
[0092] The preferred capacity threshold for this implementation is 100MB. It is understood that the benefits of compression are significant for data larger than 100MB, with strong compression reducing the size by more than 60%. However, for data smaller than 100MB, the computational overhead (such as CPU usage) of strong compression may outweigh the benefits, and basic compression is more efficient.
[0093] Specifically, the flow control layer employs a backpressure mechanism including:
[0094] The occupancy of the receiving end's buffer is monitored in real time. When the occupancy rate is high, the backpressure signal is triggered to reduce the transmission rate of the sending end.
[0095] When the receiver's buffer occupancy is low, the backpressure signal is released to restore the transmitter's initial transmission rate.
[0096] The receiver's buffer occupancy rate (the ratio of used space to total capacity) is monitored in real time, and a high occupancy threshold and a low occupancy threshold are preset. When the occupancy rate reaches or exceeds the high occupancy threshold, the receiver sends a backpressure signal to the transmitter, and the transmitter reduces the transmission rate after receiving the signal. When the occupancy rate drops below the low occupancy threshold, the receiver sends a release signal, and the transmitter resumes the initial rate, thus achieving dynamic matching between transmission and reception capabilities.
[0097] In this implementation, the high occupancy threshold is 80%, and the low occupancy threshold is 30%. It is understandable that when the buffer occupancy rate reaches 80%, it is close to the risk of overflow, which may lead to data loss and requires triggering back pressure reduction. When it drops to 30%, there is enough remaining space to safely restore the initial rate, which is in line with the management logic of preventing buffer overflow while ensuring efficiency.
[0098] For example, please refer to Table 2 to simulate the performance of the receiver buffer under different loads and compare the cases with and without the backpressure mechanism enabled:
[0099] Table 2 Comparison of Back Pressure Effects under Different Loads
[0100]
[0101] In this implementation, the backpressure mechanism effectively avoids data loss caused by receiver overload, ensuring the reliability of data transmission.
[0102] In this invention, by leveraging periodic parameter acquisition, automatic classification, and dynamic strategy adjustments such as automatic network switching and on-demand protocol selection, the system can adapt to different network environments and load changes without manual intervention. Real-time adjustment of the backpressure mechanism and automatic adaptation of checkpoint intervals further enhance the system's responsiveness to dynamic scenarios, reduce manual maintenance costs, and make the transmission process more closely aligned with the fluctuations in actual business needs.
[0103] Please see Figure 3 As shown, it is a logic diagram of the fault-tolerant layer using asynchronous checkpointing technology in an embodiment of the present invention. The fault-tolerant layer using asynchronous checkpointing technology includes:
[0104] During the high-frequency usage period, an asynchronous checkpoint is generated at a first interval;
[0105] During the low-frequency usage period, an asynchronous checkpoint is generated at a second interval, and the checkpoint data is stored in a vectorized format.
[0106] The system presets the checkpoint generation interval for high-frequency usage periods (first interval) and the interval for low-frequency usage periods (second interval, which is longer than the first interval). During high-frequency periods, asynchronous checkpoints are automatically generated according to the first interval, and during low-frequency periods, they are generated according to the second interval. Checkpoint data is uniformly stored in a vectorized format (such as Arrow). Through columnar storage and pre-compression optimization, storage space is reduced and read / write speeds are accelerated, balancing fault tolerance and resource efficiency.
[0107] In implementation, the first interval is preferably 5 minutes; the second interval is preferably 30 minutes. During high-frequency periods, data is dense and the impact of failures is significant, so a checkpoint every 5 minutes can enable rapid recovery. During low-frequency periods, the amount of data is small and the risk is low, so a checkpoint every 30 minutes can reduce backup overhead. Checkpoint generation requires CPU / bandwidth, so it is necessary to balance fault tolerance and resource consumption.
[0108] Specifically, the network replacement strategy includes:
[0109] When the current network is a low-quality network and is in a period of high-frequency use, automatically switch to a backup high-quality network;
[0110] After the switch, core data processed in vectorized format will be transmitted first.
[0111] The system monitors the current network quality and time period in real time. When a low-quality network is identified and is in a high-frequency usage period, the system automatically triggers a network switching mechanism to connect to a preset backup high-quality network. After the switch is completed, content marked as core data (such as key business indicators) is transmitted first. Core data must be pre-processed in a vectorized format to reduce parsing latency and ensure priority delivery, thus guaranteeing the continuity of critical data transmission.
[0112] In this invention, asynchronous checkpointing technology enables rapid fault recovery, and a network switching strategy automatically switches to a backup network in case of anomalies, prioritizing the transmission of core data. This significantly enhances the system's resilience. Whether it's a sudden network failure or data transmission interruption, fault tolerance mechanisms minimize data loss, and a critical data priority strategy ensures that core business processes remain unaffected, providing comprehensive protection for the continuity and security of large-scale data transmission.
[0113] Example 1
[0114] To evaluate the system's performance, a 24-hour simulation test was conducted in the data center of a large internet company. The test data volume was approximately 500 TB, and the network environment included a high-quality local area network (LAN) and a cross-regional public network (WAN). The comparison results are shown in Figure 3.
[0115] Table 3 Comparison of Improvement Effects of Examples
[0116]
[0117] The system in this embodiment has significant improvements in transmission rate, latency, stability and resource utilization, with only negligible additional computational overhead.
[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using dedicated hardware-based apparatus to perform the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0119] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
[0120] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A system for improving transmission efficiency based on big data, characterized in that, include: The parameter acquisition module is used to periodically collect the transmission rate of each network, the bandwidth usage information at different time periods, and the transmission delay information. The transmission rate includes the real-time transmission rate and the average transmission rate. The bandwidth usage information includes the usage ratio of each network at different time periods. The transmission delay information includes the round-trip delay and processing delay during data transmission. The performance evaluation module, which is connected to the parameter acquisition module, is used to evaluate the performance level of each network bandwidth based on transmission rate, bandwidth occupancy information and transmission delay information. The classification module, which is connected to the performance evaluation module and the parameter acquisition module, is used to classify time periods within the monitoring period according to bandwidth usage information, and to classify each network according to network stability and latency fluctuation. The transmission optimization module, which is connected to the classification module and the performance evaluation module, is used to determine the network adjustment method in different time periods based on the time period classification results and network classification results, combined with the QUIC / RDMA technology of the protocol layer, the Avro serialization and Snappy compression algorithm of the data layer, the backpressure mechanism of the flow control layer and the asynchronous checkpoint technology of the fault tolerance layer. The network adjustment method includes network replacement strategy and transmission parameter adjustment scheme. The dynamic adjustment module, which is connected to the transmission optimization module, is used to dynamically adjust the data transmission rate during transmission through a backpressure mechanism, achieve transmission fault tolerance and fast recovery based on asynchronous checkpoint technology, and optimize cross-system data interaction using a vectorized format.
2. The big data-based transmission efficiency improvement system according to claim 1, characterized in that, The periodic collection period is a fixed duration, and the occupancy ratio in the bandwidth occupancy information includes the peak occupancy ratio and the continuous occupancy ratio.
3. The big data-based transmission efficiency improvement system according to claim 2, characterized in that, The evaluation of the performance level of each network bandwidth includes: If the transmission rate is at a high level, the bandwidth utilization rate is at a low level, and the transmission delay is at a low level, then the performance level is high efficiency. If the transmission rate is at a medium level, the bandwidth utilization rate is at a medium level, and the transmission latency is at a medium level, then the performance level is medium. If the transmission rate is at a low level, the bandwidth utilization rate is at a high level, or the transmission delay is at a high level, then the performance level is inefficient.
4. The big data-based transmission efficiency improvement system according to claim 3, characterized in that, The time periods within the monitoring period are categorized as follows: High-frequency usage periods are those during which the peak bandwidth occupancy rate remains high for several consecutive collection cycles. Low-frequency usage periods are those during which the peak bandwidth occupancy rate remains low for several consecutive collection cycles.
5. The big data-based transmission efficiency improvement system according to claim 4, characterized in that, The classification of networks includes: A high-quality network is one in which transmission rate fluctuations are within a preset fluctuation range and latency fluctuations are within an allowable fluctuation range. A low-quality network is a network whose transmission rate fluctuations exceed the preset fluctuation range or whose latency fluctuations exceed the allowable fluctuation range.
6. The big data-based transmission efficiency improvement system according to claim 5, characterized in that, The protocol layer employs QUIC / RDMA technology, including: During the high-frequency usage period, select the high-quality network and enable RDMA technology; When the frequency of use is low or the network is of low quality, QUIC technology is used to improve transmission stability.
7. The big data-based transmission efficiency improvement system according to claim 6, characterized in that, The data layer employs Avro serialization and Snappy compression algorithms, including: Structured data is uniformly encoded using the Avro serialization format; For serialized data, the Snappy compression strength is selected according to the data size. High-strength compression is used for large-capacity data, and basic compression is used for small-capacity data.
8. The big data-based transmission efficiency improvement system according to claim 7, characterized in that, The flow control layer employs a backpressure mechanism, including: The system monitors the occupancy of the receiver's buffer in real time. When the occupancy rate is high, it triggers a backpressure signal to reduce the transmission rate at the transmitter. When the receiver's buffer occupancy is low, the backpressure signal is released to allow the transmitter to resume its initial transmission rate.
9. The big data-based transmission efficiency improvement system according to claim 8, characterized in that, The fault-tolerant layer employs asynchronous checkpointing technology, including: During the high-frequency usage period, an asynchronous checkpoint is generated at a first interval; During the low-frequency usage period, an asynchronous checkpoint is generated at a second interval, and the checkpoint data is stored in a vectorized format.
10. The big data-based transmission efficiency improvement system according to claim 9, characterized in that, The network replacement strategy includes: When the current network is a low-quality network and is in a period of high-frequency use, automatically switch to a backup high-quality network; After the switch, core data processed in vectorized format will be transmitted first.
Citation Information
Patent Citations
Communication method, network equipment and terminal equipment
CN108282869A
Method and device for switching audio channels of communication module
CN118538247A