A special network data security detection method based on big data

By combining big data technology with a hybrid error correction strategy of forward error correction coding and automatic retransmission requests, along with deep packet inspection and abnormal traffic identification, the problem of data integrity and confidentiality in high-interference scenarios of dedicated networks is solved, and the accurate detection and dynamic protection against potential data leakage are achieved.

CN120433986BActive Publication Date: 2026-01-27河北工业职业技术大学
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510562014.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2026-01-27
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

Existing dedicated network data security detection technologies lack adaptability, making it difficult to ensure data integrity and confidentiality in high-interference or high-risk scenarios, and their protective capabilities are inadequate when faced with diverse threats of data corruption and leakage.

Method used

A hybrid error correction strategy based on big data is adopted, combining forward error correction codes and automatic retransmission requests. Through deep packet inspection and abnormal traffic identification technologies, network traffic is monitored in real time, the error correction scheme is adaptively adjusted, and the system is matched with a sensitive data fingerprint database to detect potential data leakage, generate alarm logs, and dynamically adjust network traffic control strategies.

Benefits of technology

It improves the reliability and security of data transmission, enables accurate detection of potential data breaches, reduces the risk of sensitive information leakage, and forms a closed-loop data breach protection mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120433986B_ABST
    Figure CN120433986B_ABST
Patent Text Reader

Abstract

The application discloses a special network data security detection method based on big data, belongs to the safety detection field, and comprises the following steps: acquiring real-time data packets from special network traffic, encoding the data packets through a forward error correction code technology, generating redundant check information, and obtaining encoded data packets; if bit errors are detected in the encoded data packets during transmission, requesting the sender to retransmit the damaged data packets through an automatic repeat request mechanism to obtain retransmitted data packets; according to adaptive error correction parameters, adjusting the redundancy of the forward error correction code and the trigger threshold of the automatic repeat request, generating an optimized error correction scheme, and obtaining repaired data packets; for an alarm processing instruction, automatically adjusting a network traffic control strategy, limiting the transmission bandwidth of an abnormal destination, updating a sensitive data fingerprint library, and obtaining an optimized protection configuration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of security detection, and in particular relates to a dedicated network data security detection method based on big data. Background Technology

[0002] In the information age, dedicated networks, as critical infrastructure, make the research on data security detection methods of paramount strategic importance. With the widespread application of big data technology, the amount of information carried by dedicated networks has surged. Data security not only concerns communication quality but also directly impacts national security, enterprise operations, and the protection of personal privacy. Big data-based detection methods, through the processing and analysis of massive amounts of data, provide a novel technological path for cybersecurity.

[0003] However, existing solutions still reveal significant shortcomings in practical applications. Traditional data security detection technologies often rely on single error correction or protection mechanisms, lacking adaptability to complex network environments. Furthermore, in high-interference or high-risk scenarios, data integrity and confidentiality are difficult to fully guarantee. These limitations render dedicated networks inadequate in the face of diverse threats to data corruption and leakage. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides a dedicated network data security detection method based on big data, comprising:

[0005] Real-time data packets are obtained from dedicated network traffic, and the real-time data packets are encoded using forward error correction coding technology to generate redundant verification information, thus obtaining the encoded data packets.

[0006] If a bit error is detected in the encoded data packet during transmission, the automatic retransmission request mechanism is used to request the sender to retransmit the damaged data packet, and the retransmitted data packet is obtained.

[0007] For the encoded data packet and the retransmitted data packet, a hybrid error correction strategy is adopted. By analyzing historical error correction patterns, the current network interference intensity is calculated to obtain adaptive error correction parameters.

[0008] Based on the adaptive error correction parameters, the redundancy of the forward error correction code and the trigger threshold of the automatic retransmission request are adjusted to generate an optimized error correction scheme and obtain the repaired data packet.

[0009] Network egress traffic is extracted from the repaired data packets. Through deep packet inspection technology, combined with content-aware analysis and context analysis, the header and content of the data packets are parsed to obtain traffic feature data.

[0010] For the aforementioned traffic characteristic data, an abnormal traffic pattern recognition algorithm is used, combined with pre-established historical baseline data, to calculate the traffic deviation value, determine abnormal traffic behavior, and obtain an abnormal traffic identifier;

[0011] If the abnormal traffic identifier indicates a potential data leak, the data content is extracted from the traffic feature data, and by matching it with the sensitive data fingerprint database, it is determined whether it contains sensitive information, thus obtaining the data leak detection result.

[0012] Based on the data leakage detection results, an alarm log containing abnormal traffic identifiers and sensitive information matching details is generated and transmitted to the security management platform through a preset encrypted channel to obtain alarm processing instructions;

[0013] In response to the alarm handling instructions, the network traffic control policy is automatically adjusted to limit the transmission bandwidth of abnormal destinations and update the sensitive data fingerprint database to obtain an optimized protection configuration.

[0014] Preferably, the process of obtaining the encoded data packet includes:

[0015] Real-time data packets are acquired through a dedicated network interface, and initial error correction information is generated using forward error correction coding technology to obtain encoded data packets.

[0016] Redundant verification information in the encoded data packet is verified using a preset error correction code rule to determine the verified data packet.

[0017] If packet loss is detected in the verified data packet, the data is reconstructed using redundant verification information to obtain the reconstructed data packet;

[0018] Based on the integrity of the reconstructed data packets, traffic analysis tools are used to detect data consistency and determine the consistency results.

[0019] By adjusting the encoding parameters based on the consistency results, an optimized error-correcting code is generated, resulting in an optimized data packet.

[0020] For optimized data packets, obtain their traffic characteristics during transmission and determine the changing trend of traffic characteristics;

[0021] Based on the changing trends of traffic characteristics, the encoding strategy is adjusted through machine learning algorithms to obtain the final encoded data packets.

[0022] Preferably, the process of obtaining the retransmitted data packet includes:

[0023] By analyzing bit errors in data packet transmission through an error detection mechanism, it can be determined whether there are erroneous bits and obtain the error detection result.

[0024] If the error detection result indicates that there is a bit error, a retransmission request is generated and sent to the sender through the request mechanism to obtain the sender's response status;

[0025] Based on the response status of the sending end, receive the retransmitted data packet, determine the integrity of the data packet, and obtain the retransmitted data packet;

[0026] A data integrity verification algorithm is used to verify the retransmitted data packets to determine whether bit errors still exist and to obtain the verification result.

[0027] If the verification result shows no bit errors, the retransmitted data packets are integrated into the transmission sequence to determine the continuity of the transmission process and obtain the updated data stream.

[0028] The updated data stream is verified a second time using a cyclic redundancy check algorithm to determine the integrity of the data stream and obtain the final verification result.

[0029] Based on the final verification results, the complete data packet sequence is output to obtain the retransmitted data packet.

[0030] Preferably, the process of obtaining the adaptive error correction parameters includes:

[0031] By storing historical error correction data, the distribution of error correction patterns is obtained, and the pattern analysis results are obtained. Using the pattern analysis results, data packet analysis features are extracted from encoded data and retransmitted data to obtain feature acquisition output.

[0032] The output is obtained based on the characteristics, the statistical value of network interference is calculated, and the interference intensity value is obtained. If the interference intensity value exceeds the preset threshold, the calculation strategy is adjusted through intensity analysis to obtain the optimized intensity analysis result.

[0033] Based on the optimized intensity analysis results, the adaptive parameters are calculated using a linear regression algorithm. The parameter calculation output is obtained, and combined with the current data packet analysis characteristics, the final adaptive error correction parameters are determined.

[0034] Preferably, the process of obtaining the repaired data packet includes:

[0035] Obtain adaptive error correction parameters for the transmission environment, and determine the initial forward error correction code redundancy and automatic retransmission request trigger threshold by analyzing channel noise level and packet loss rate;

[0036] If the channel noise level is higher than a preset threshold, the redundancy of the forward error correction code is increased, and the error correction code is generated by Reed-Solomon code to obtain a data packet with enhanced error correction capability.

[0037] If the packet loss rate is higher than the preset threshold, the automatic retransmission request trigger threshold is lowered, and erroneous data packets are marked by the selective retransmission protocol to obtain a retransmission request sequence.

[0038] Based on the data packets with enhanced error correction capabilities, a forward error correction decoding algorithm is used to repair the damaged data packets, resulting in a preliminary repaired data packet.

[0039] For the initially repaired data packets, residual errors are detected through cyclic redundancy check to obtain error detection results;

[0040] If the error detection result shows residual errors, the corresponding data packets are retransmitted using the automatic retransmission request mechanism according to the retransmission request sequence to obtain the final repaired data packets.

[0041] Preferably, the process of obtaining traffic characteristic data includes:

[0042] Deep packet inspection technology is used to extract network outbound traffic from the repaired data packets, and the packet headers and contents are parsed to obtain the original traffic data.

[0043] Content-aware analysis technology is used to parse the raw traffic data and extract the protocol type, source address and destination address from the data packets to obtain content feature data.

[0044] By using context analysis technology, correlation analysis is performed on content feature data, and context feature data is obtained by combining the timestamp of data packets and session identifiers;

[0045] If the protocol type in the context feature data is a preset protocol of interest, then the support vector machine algorithm is used to classify the context feature data, determine whether the traffic is abnormal, and obtain the abnormal traffic identifier.

[0046] Based on the abnormal traffic identifier, the context feature data is filtered, and the data packet content corresponding to the abnormal traffic is extracted to obtain the abnormal traffic feature data.

[0047] The key features are determined by ranking the abnormal traffic feature data according to the importance of the features using the random forest algorithm, thus obtaining the key traffic feature data.

[0048] Using a preset threshold judgment method, if the frequency of the source address in the key traffic feature data exceeds the threshold, it is marked as high-risk traffic, and high-risk traffic feature data is obtained.

[0049] Preferably, the process of obtaining the abnormal traffic identifier includes:

[0050] Traffic characteristic data is acquired, and the data is cleaned and standardized through a preprocessing module to obtain a standardized traffic dataset;

[0051] The random forest algorithm is used in conjunction with pre-established historical baseline data to perform feature analysis on the normalized traffic dataset, generating traffic size deviation, time deviation, and destination deviation.

[0052] The deviation value generation module calculates the difference between each deviation value and the historical baseline data to obtain the deviation feature vector.

[0053] If the magnitude of the deviation feature vector is greater than the preset threshold, the abnormal traffic identification algorithm is used to classify the traffic data and determine abnormal traffic behavior.

[0054] Based on the classification results of abnormal traffic behavior, abnormal traffic identifiers are generated and stored in the abnormal behavior database.

[0055] Preferably, the process of obtaining the data breach detection result includes:

[0056] If the traffic monitoring system detects a traffic anomaly, it obtains traffic characteristic data from the network traffic, extracts the data content through packet parsing technology, and obtains the first data content.

[0057] A feature extraction algorithm is used to extract the structured features of the data content from the first data content, and a first feature set is generated.

[0058] The first feature set is matched with a pre-established sensitive data fingerprint database. If the matching degree exceeds a preset threshold, it is determined that the data contains sensitive information, and the first matching result is obtained.

[0059] Based on the first matching result, the decision tree algorithm is used to classify the possibility of data leakage and generate data leakage detection results;

[0060] Extract contextual information of the data breach event from the data breach detection results, obtain related traffic source information through log analysis technology, and generate the first set of traffic sources;

[0061] For the first set of traffic sources, anomaly traffic analysis technology is used to determine whether the traffic sources are continuously abnormal, and to obtain the abnormal traffic confirmation result;

[0062] By jointly analyzing the abnormal traffic confirmation results and the first matching results, a final data breach detection report is generated.

[0063] On the other hand, the present invention also provides an electronic device including a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the processor implements the method when executing the computing program.

[0064] On the other hand, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method.

[0065] Compared with the prior art, the present invention has the following advantages and technical effects:

[0066] This invention discloses a method for preventing network data leakage. By monitoring dedicated network traffic in real time, it employs a hybrid error correction strategy combining forward error correction coding and automatic retransmission requests to adaptively correct data packets during transmission, improving data transmission reliability. Simultaneously, deep packet inspection and abnormal traffic identification technologies are used to analyze the content and behavior of the repaired data packets, achieving accurate detection of potential data leakage. When abnormal traffic is detected, this invention automatically matches it with a sensitive data fingerprint database to determine if sensitive information leakage exists and generates detailed alarm logs. Finally, based on alarm processing instructions, network traffic control policies are dynamically adjusted and protection configurations are updated, forming a closed-loop data leakage prevention mechanism. This invention effectively enhances network security protection capabilities and reduces the risk of sensitive information leakage through multi-layered data protection and intelligent analysis. Attached Figure Description

[0067] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0068] Figure 1 This is a flowchart of a dedicated network data security detection method based on big data, according to an embodiment of the present invention. Detailed Implementation

[0069] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0070] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0071] Example 1

[0072] like Figure 1 As shown, this embodiment provides a dedicated network data security detection method based on big data, including:

[0073] Step S101: Obtain real-time data packets from dedicated network traffic, encode the data packets using forward error correction coding technology, generate redundancy check information, and obtain the encoded data packets.

[0074] Real-time data packets are acquired through a dedicated network interface. Initial error correction information is generated using forward error correction coding (FEC) technology, resulting in encoded data packets. Redundant checksums in the encoded data packets are verified using pre-defined error correction coding rules to determine the verified data packets. If packet loss is detected in the verified data packets, the data is reconstructed using the redundant checksums, resulting in reconstructed data packets. Based on the integrity of the reconstructed data packets, traffic analysis tools are used to check data consistency and determine the consistency result. Encoding parameters are adjusted based on the consistency result to generate optimized error correction codes, resulting in optimized data packets. For the optimized data packets, traffic characteristics during transmission are acquired, and the changing trends of these characteristics are determined. Based on these trends, machine learning algorithms are used to adjust the encoding strategy, resulting in the final encoded data packets.

[0075] For example, when capturing real-time data packets in dedicated network traffic, a packet capture tool such as tcpdump based on the libpcap library can be used. The filtering rule can be set to "host192.168.1.100and port443" to capture data from specific IPs and HTTPS ports, with a sampling frequency of 1000 packets / second. After MTU fragmentation, each captured raw data packet is 1500 bytes long, including a 14-byte Ethernet header, a 20-byte IP header, a 32-byte TCP header, and a 1434-byte payload. Forward error correction is performed using RS(255,223) Reed-Solomon codes, dividing the 223 bytes of raw data into 16 14-byte blocks (padded with trailing zeros). A 32-byte checksum is calculated using the generator polynomial g(x) = (x-α^0)(x-α^1)(x-α^31) over the Galois domain GF(2^8), which can correct up to 16 bytes of errors. During encoding, each 1434-byte payload is processed into eight 223-byte groups, generating a total of 8 × 32 = 256 bytes of redundant check information, which is then concatenated with the original data to form a 1690-byte encoded packet. The error correction capability can be verified by calculating the Hamming distance: data can be fully recovered when the bit error rate is below 1.2% (16 / 1357). To improve efficiency, SIMD instructions are used for parallel computation of finite field multiplication, achieving an encoding throughput of 12Gbps under the Intel AVX2 instruction set. The encoded data packet is then checked for integrity using CRC32, with the check polynomial being 0x04C11DB7, generating a 4-byte checksum which is appended to the end of the packet.

[0076] In step S102, if a bit error is detected in the encoded data packet during transmission, an automatic retransmission request mechanism is used to request the sender to retransmit the damaged data packet, and the retransmitted data packet is obtained.

[0077] An error detection mechanism analyzes bit errors in data packet transmission to determine if any erroneous bits exist, yielding an error detection result. If the error detection result indicates the presence of a bit error, a retransmission request is generated and sent to the sender via a request mechanism to obtain the sender's response status. Based on the sender's response status, retransmitted data packets are received to verify their integrity, resulting in the retransmitted data packets. A data integrity verification algorithm is used to verify the retransmitted data packets, determining if bit errors still exist, and obtaining the verification result. If the verification result shows no bit errors, the retransmitted data packets are integrated into the transmission sequence to ensure the continuity of the transmission process, resulting in an updated data stream. A cyclic redundancy check (CRC) algorithm is used to perform a secondary verification on the updated data stream to determine its integrity, obtaining the final verification result. Based on the final verification result, a complete data packet sequence is output to determine the reliability of the transmission process, resulting in the final data output.

[0078] For example, during data transmission, when the receiving end detects a bit error in a data packet using the Cyclic Redundancy Check (CRC-32) algorithm, such as a mismatch between the checksum 0xEDB88320 and the expected value 0x1A2B3C4D, it triggers the Automatic Repeat Request (ARQ) mechanism. The receiving end immediately constructs a feedback frame containing a NACK signal and a sequence number (e.g., Seq=5), and sends it to the source end via the Fast Retransmit mechanism in the TCP protocol, in a network environment with an RTT of 200ms. After parsing the NACK, the source end extracts the original data packet (e.g., a 1500-byte IP packet) from the sending buffer, performs forward error correction using Hamming(7,4) encoding, and appends a 16-bit checksum 0xABCD. It then adjusts the window size to 4 data packets using the sliding window protocol for retransmission. If a NACK is still received after three consecutive retransmissions (with a threshold set to MAX_RETRY=3), the Selective Repeat (SACK) mechanism is activated, retransmitting only the lost segment (e.g., offset 256-512 bytes). The receiving end performs a secondary check on the retransmitted data packets. If the CRC check passes and the sequence numbers are consecutive, it delivers the data to the application layer and returns an ACK signal (e.g., ACKSeq=5), while simultaneously updating the receive window to sequence number 6. Throughout the process, the retransmission interval is adjusted using an exponential backoff algorithm, with an initial timeout of 500ms and a backoff factor of 2, ensuring network congestion control. For scenarios with high real-time requirements, forward error correction codes (e.g., Reed-Solomon (255,223)) can be used to embed 32 bytes of redundant data into the data packets, enabling a 95% error correction success rate even with a 5% random bit error rate.

[0079] Step S103: For the encoded data packet and the retransmitted data packet, a hybrid error correction strategy is adopted. By analyzing historical error correction patterns, the current network interference intensity is calculated to obtain adaptive error correction parameters.

[0080] By using stored historical error correction data, the distribution of error correction patterns is obtained, and pattern analysis results are derived. Using these results, packet analysis features are extracted from encoded and retransmitted data, yielding feature acquisition outputs. Based on these outputs, statistical values ​​of network interference are calculated, resulting in interference intensity values. If the interference intensity value exceeds a preset threshold, the calculation strategy is adjusted through intensity analysis, resulting in optimized intensity analysis results. Using these optimized results, a linear regression algorithm is employed to calculate adaptive parameters, yielding parameter calculation outputs. These outputs, combined with current packet analysis features, determine the final adaptive error correction parameters. Based on these final adaptive error correction parameters, the error correction strategies for encoded and retransmitted data are adjusted, resulting in an updated data processing scheme.

[0081] For example, the Reed-Solomon (255, 223) encoding scheme is used when encoding data packets. Each data packet contains a 223-byte payload and a 32-byte checksum, which can correct up to 16 bytes of errors. When a data packet transmission fails and needs to be retransmitted, the system records historical error correction patterns. For example, in the last 10 transmissions, an average of 12 bytes of errors need to be corrected each time, with a standard deviation of 3 bytes. Based on this historical data, an exponentially weighted moving average algorithm (α = 0.2) is used to calculate the current network interference intensity index. When this index exceeds the threshold of 0.75, an adaptive adjustment mechanism is triggered. The adjustment process uses the least squares method to fit the relationship curve between historical error correction requirements and network conditions: y = 1.2x + 5 (where x is the interference intensity and y is the required redundant bytes). Based on the current interference intensity of 0.8, it is calculated that 14.6 bytes of additional protection are needed, and the system automatically increases the error correction capability to 18 bytes. Simultaneously, the system dynamically adjusts the mixing ratio of Forward Error Correction (FEC) and Automatic Repeat Request (ARQ). When the packet loss rate is below 5%, a 7:3 FEC priority strategy is adopted; when the packet loss rate is between 5% and 15%, a 5:5 balanced strategy is adopted; and when the packet loss rate is above 15%, a 3:7 ARQ priority mode is switched. The historical database is updated after each adjustment to ensure that subsequent decisions are based on the latest network status information. To optimize performance, the system recalculates network parameters every 30 seconds, using a sliding window mechanism to maintain the transmission records of the most recent 100 data packets as the basis for analysis.

[0082] Step S104: Based on the adaptive error correction parameters, adjust the redundancy of the forward error correction code and the trigger threshold of the automatic retransmission request to generate an optimized error correction scheme and obtain the repaired data packet.

[0083] The system acquires adaptive error correction parameters for the transmission environment. By analyzing channel noise levels and packet loss rates, it determines the initial forward error correction code redundancy and the automatic retransmission request trigger threshold. If the channel noise level exceeds the preset threshold, the forward error correction code redundancy is increased. Error correction codes are generated using Reed-Solomon codes (where n represents the total number of symbols, k represents the number of data symbols, and nk represents the number of redundant symbols), resulting in data packets with enhanced error correction capabilities. If the packet loss rate exceeds the preset threshold, the automatic retransmission request trigger threshold is decreased. Erroneous data packets are marked using a selective repeat protocol, resulting in a retransmission request sequence. Based on the data packets with enhanced error correction capabilities, a forward error correction decoding algorithm (Berlekamp-Massey algorithm, inputting the position of the erroneous symbol and outputting the error correction polynomial) is used to repair damaged data packets, resulting in initially repaired data packets. For the initially repaired data packets, cyclic redundancy check (Cyclic Redundancy Check, inputting the data packet bitstream and outputting the checksum) is used to detect residual errors, resulting in error detection results. If the error detection result shows residual errors, the corresponding data packets are retransmitted using the automatic retransmission request mechanism according to the retransmission request sequence to obtain the final repaired data packet. By analyzing the error correction efficiency and transmission reliability of the final repaired data packet, the adaptive error correction parameters are dynamically adjusted, and the forward error correction code redundancy and automatic retransmission request trigger threshold are updated to obtain an optimized error correction scheme.

[0084] For example, during the adaptive error correction parameter adjustment process, the channel quality indicators (such as bit error rate BER = 1.2 × 10⁻⁶) are first monitored in real time. -4The redundancy of the forward error correction code is dynamically calculated (SNR = 18dB). Using Reed-Solomon (255, 223) codes as the basis, the number of redundant symbols is adaptively expanded according to the channel state. When the packet loss rate exceeds 0.5%, the coding block is automatically adjusted to (255, 215), increasing the redundant data by 12.8%. Simultaneously, based on the sliding window algorithm, the number of consecutive erroneous frames is counted. When three CRC check failures are detected within 10ms, the automatic retransmission request (ARQ) threshold is adjusted, shortening the retransmission timeout from 200ms to 150ms. A Markov chain model is constructed to predict channel degradation trends. When the predicted probability of a bit error rate increase in the next cycle exceeds 65%, the constraint length of the convolutional code is increased from 7 to 9 in advance, and the soft-decision mode of the Viterbi decoder is enabled. The packet repair phase employs an iterative deletion decoding algorithm. For the received (255, 215) codeword, the syntactic equation is first calculated. If the syntactic equation is non-zero, a Qian search algorithm is initiated to locate up to 20 error positions. This is then combined with the Berlekamp-Massey algorithm to complete error correction. For residual erroneous packets, the Hamming distance between the ARQ retransmission version and the original packet is compared. The version with the smallest difference is selected for weighted merging, and the finally repaired packet is output. Throughout the process, the channel parameter database is continuously updated, and an exponentially weighted moving average method (α = 0.3) is used to smooth historical measurements, ensuring the stability of parameter adjustments.

[0085] Step S105: Extract network egress traffic from the repaired data packets. Using deep packet inspection technology, combined with content-aware analysis and context analysis, parse the header and content of the data packets to obtain traffic feature data.

[0086] Deep packet inspection (DPI) technology is used to extract network outbound traffic from repaired data packets, parsing the packet headers and content to obtain raw traffic data. Content-aware analysis (CNA) technology is then used to parse the raw traffic data, extracting protocol types, source addresses, and destination addresses from the packets to obtain content feature data. Context analysis is then used to perform correlation analysis on the content feature data, combining packet timestamps and session identifiers to obtain context feature data. If the protocol type in the context feature data matches a pre-defined protocol of interest, a support vector machine (SVM) algorithm is used to classify the context feature data to determine if the traffic is abnormal, resulting in an abnormal traffic identifier. Based on the abnormal traffic identifier, the context feature data is filtered to extract the packet content corresponding to the abnormal traffic, resulting in abnormal traffic feature data. A random forest algorithm is used to rank the abnormal traffic feature data by feature importance, identifying key features to obtain key traffic feature data. Finally, a pre-defined threshold method is used; if the frequency of the source address in the key traffic feature data exceeds a certain threshold, it is marked as high-risk traffic, resulting in high-risk traffic feature data.

[0087] For example, in the repaired data packets, deep packet inspection is first used to extract network outbound traffic. Regular expressions are then used to match header information for specific protocols such as HTTP, for example, matching "GET / index.htmlHTTP / 1.1" to identify the request type. Next, content-aware analysis is used, employing the TF-IDF algorithm to calculate the weight of keywords in the data packets. For instance, "login" appears 50 times in 1000 data packets, with a weight of 0.05. Combined with contextual analysis, this identifies the traffic as potentially involving user login behavior. The data packet headers are further parsed to extract source IP addresses such as 192.168.1.1 and destination IP addresses such as 203.0.113.1, combined with port numbers such as 80, to determine the traffic direction. By analyzing the data packet content, the K-means clustering algorithm is used to categorize the traffic into different classes, such as dividing 1000 data packets into 5 classes, each containing 200 data packets, identifying abnormal traffic such as DDoS attacks. Finally, by combining traffic characteristic data, such as an average packet size of 1500 bytes and an average transmission rate of 1Mbps, a traffic report is generated to provide a basis for network security management.

[0088] Step S106: For traffic characteristic data, an abnormal traffic pattern recognition algorithm is used, combined with pre-established historical baseline data, to calculate the deviation values ​​of traffic size, time, and destination, determine abnormal traffic behavior, and obtain abnormal traffic identifiers.

[0089] Traffic characteristic data is acquired, and the data is cleaned and standardized through a preprocessing module to obtain a normalized traffic dataset. A random forest algorithm, combined with pre-established historical baseline data, is used to perform feature analysis on the normalized traffic dataset, generating traffic size deviation, time deviation, and destination deviation values. The deviation value generation module calculates the difference between each deviation value and the historical baseline data to obtain a deviation feature vector. If the magnitude of the deviation feature vector is greater than a preset threshold, an abnormal traffic identification algorithm is used to classify the traffic data and determine abnormal traffic behavior. Based on the classification results of abnormal traffic behavior, abnormal traffic identifiers are generated and stored in an abnormal behavior database. Abnormal traffic identifiers are extracted from the abnormal behavior database, and time series analysis tools are used to generate the frequency and trend of abnormal behavior. Based on the trend, a dynamic distribution map of abnormal traffic behavior is generated through a visualization module to determine the abnormal traffic monitoring results.

[0090] For example, in abnormal traffic detection, network traffic data is first collected at 5-minute intervals using a sliding window algorithm. For instance, the average traffic to the destination IP 192.168.1.100 within a certain time period is 120 Mbps, with a standard deviation of 15 Mbps. An Exponentially Weighted Moving Average (EWMA) algorithm is used to establish a baseline model, with a smoothing coefficient α = 0.3. When the real-time traffic reaches 180 Mbps, its Z-score is calculated to be (180-120) / 15 = 4, exceeding the 3σ threshold. For time-dimensional anomalies, Time Series Decomposition (STL) is used to detect periodic deviations. If the traffic during weekday morning peak hours decreases by 40% compared to the baseline, an alarm is triggered. For destination anomalies, a clustering algorithm (such as DBSCAN, with eps = 0.5 and min_samples = 10) is used to discover that a certain IP suddenly connects to 50 new destination ports, while the historical baseline only shows 3-5. Finally, combining a multi-dimensional decision tree model, when the traffic volume Z-score > 3, the time deviation > 30%, and the destination dispersion > 20, it is marked as an abnormal traffic ID: ALERT_202. This traffic is then correlated with a threat intelligence database to match known attack characteristics. If the similarity to the CVE-2023-1234 vulnerability exploitation pattern reaches 85%, a defense rule is automatically generated to block the traffic. Throughout the process, Kalman filtering is used to dynamically update the baseline data, ensuring the model adapts to changes in the network environment.

[0091] Step S107: If the abnormal traffic identifier indicates a potential data leak, extract the data content from the traffic feature data, and determine whether it contains sensitive information by matching it with the sensitive data fingerprint database, thereby obtaining the data leak detection result.

[0092] If the traffic monitoring system detects an anomaly in network traffic, it extracts traffic characteristic data from the network traffic and extracts the data content using packet parsing technology to obtain the first data content. From the first data content, a feature extraction algorithm is used to obtain the structured features of the data content, generating a first feature set. This first feature set is matched against a pre-established sensitive data fingerprint database. If the matching degree exceeds a preset threshold, it is determined that sensitive information is contained, resulting in a first matching result. Based on the first matching result, a decision tree algorithm is used to classify the probability of data leakage, generating a data leakage detection result. The contextual information of the leakage event is extracted from the data leakage detection result, and related traffic source information is obtained through log analysis technology, generating a first traffic source set. For the first traffic source set, anomaly traffic analysis technology is used to determine whether the traffic sources are continuously abnormal, obtaining an anomaly traffic confirmation result. Through joint analysis of the anomaly traffic confirmation result and the first matching result, a final data leakage detection report is generated.

[0093] For example, when a network traffic monitoring system detects that a server suddenly generates more than 1GB of abnormal outbound traffic within 10 minutes (baseline value is 100MB / 10 minutes), the traffic analysis module first extracts the raw data packets in the TCP / UDP payload and uses Wireshark to parse the application layer protocol content. For HTTP traffic, it matches credit card number formats using regular expressions (such as `\d{18}\|\d{4}[-\s]\d{4}[-\s]\d{4}`). For database traffic, it uses an SQL syntax analyzer to extract field values ​​from the WHERE condition. The extracted text content is then hashed using SHA-256 and compared with a predefined sensitive data fingerprint database (containing 500,000 hash values, such as a combination hash of the first 6 digits of an ID card number + birth date + last 4 digits). The SimHash algorithm is used to calculate the Hamming distance, and data with a distance less than 3 is considered sensitive data. For example, `3e9d7a5f2b...` is detected. If the hash value (3e9d7a5f2b..) matches exactly with the hash of customer information stored in the fingerprint database, and this data is transmitted via an unauthorized port (such as 6667), a data breach alarm is triggered. The system automatically generates a security event report containing the data type of the breach (such as bank card number), the amount of data (128 records), and the destination IP (192.168.1.100). The entire process is processed in real time using a Kafka message queue, with an average latency controlled within 200 milliseconds.

[0094] Step S108: Based on the data leakage detection results, generate an alarm log containing abnormal traffic identifiers and sensitive information matching details, and transmit it to the security management platform through a preset encrypted channel to obtain alarm processing instructions.

[0095] The data leakage detection module performs real-time analysis of network traffic, employing anomaly detection algorithms to identify abnormal traffic patterns and generate abnormal traffic identifiers. If the abnormal traffic identifiers exceed a preset threshold, sensitive information matching details are extracted from the traffic data based on sensitive information matching rules, generating alarm logs. A log formatting tool is used to structure the alarm logs, resulting in standardized alarm log data. This standardized alarm log data is then encrypted using an encrypted transmission protocol via a preset encrypted channel and transmitted to the security management platform. On the security management platform, a classification algorithm prioritizes the received alarm log data to determine the urgency of the alarms. Based on the urgency, a preset command response mechanism is triggered, generating alarm handling commands. These commands are then sent back to the data leakage detection module via the encrypted channel to update detection rules and optimize real-time threat detection.

[0096] For example, during data breach detection, the system monitors network traffic in real time through a traffic analysis module. Employing an anomaly detection algorithm based on entropy calculation, when an IP address sends more than 5000 requests within 10 minutes, and the request content contains a large number of duplicate data packets, the system marks it as abnormal traffic and generates an abnormal traffic identifier "ABN-20231012-001". Subsequently, the system uses a regular expression matching algorithm to screen the data in the abnormal traffic for sensitive information. It finds strings conforming to the format of an ID card number, such as "110105199003071234", and strings conforming to the format of a bank card number, such as "6222020200001234567". These strings successfully match the preset sensitive information rule base. The system encapsulates the detection results into a JSON-formatted alarm log, including the abnormal traffic identifier, the matched sensitive information type, and the specific content. After encryption using the AES-256 encryption algorithm, the log is transmitted to the security management platform through a preset SSL / TLS encrypted channel. After receiving the alarm log, the security management platform decrypts it using the pre-shared key and automatically generates the processing command "BLOCK-IP-192.168.1.100" according to the preset alarm processing rules. This command instructs the firewall to immediately block access from the abnormal IP address and triggers an alarm notification, sending the relevant information to the security administrator for further investigation.

[0097] Step S109: In response to the alarm handling command, automatically adjust the network traffic control policy, limit the transmission bandwidth of abnormal destinations, update the sensitive data fingerprint database, and obtain the optimized protection configuration.

[0098] By parsing alarm handling commands, abnormal destination information is obtained, and the characteristics of abnormal traffic are determined. If the abnormal destination matches a preset sensitive data fingerprint, a traffic control policy is generated to limit the corresponding bandwidth, resulting in a preliminary network adjustment plan. Based on the preliminary network adjustment plan, the sensitive data fingerprint database is updated, adding new abnormal traffic features to generate an updated fingerprint dataset. Using the updated fingerprint dataset, potential anomalies in network traffic are detected to determine whether further bandwidth allocation adjustments are needed, resulting in an optimized traffic policy. Based on the optimized traffic policy, a protection configuration is generated, and the security configuration management module is updated to obtain the final protection configuration. If the final protection configuration triggers a new alarm handling command, alarm information is parsed repeatedly, and the traffic policy is adjusted to obtain a continuously optimized network protection plan. Based on the continuously optimized network protection plan, bandwidth allocation rules are updated periodically to generate dynamic security configuration management policies.

[0099] For example, upon detecting an anomaly alert, the system first identifies the IP address of the abnormal destination using a traffic analysis algorithm, such as 192.168.1.100. Based on historical traffic data, it calculates that the peak traffic of this IP reaches 500Mbps, far exceeding the normal threshold of 200Mbps. Subsequently, the system automatically invokes traffic control policies, using a token bucket algorithm to limit the transmission bandwidth of the abnormal destination to 100Mbps, ensuring the reasonable allocation of network resources. Simultaneously, the system analyzes the characteristics of the abnormal traffic using a machine learning model, extracting new sensitive data fingerprints, such as the hash value of a specific data packet "a1b2c3d4," and updates them to the sensitive data fingerprint database. Finally, based on the optimized protection configuration, the system redeploys firewall rules, ensuring that all traffic undergoes strict fingerprint matching and bandwidth control upon passage, thereby effectively improving network security.

[0100] On the other hand, this embodiment also provides an electronic device, including a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the processor implements the method when executing the computing program.

[0101] On the other hand, this embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method.

[0102] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A dedicated network data security detection method based on big data, characterized in that, include: Real-time data packets are obtained from dedicated network traffic, and the real-time data packets are encoded using forward error correction coding technology to generate redundant verification information, thus obtaining the encoded data packets. If a bit error is detected in the encoded data packet during transmission, the automatic retransmission request mechanism is used to request the sender to retransmit the damaged data packet, and the retransmitted data packet is obtained. For the encoded data packet and the retransmitted data packet, a hybrid error correction strategy is adopted. By analyzing historical error correction patterns, the current network interference intensity is calculated to obtain adaptive error correction parameters. Based on the adaptive error correction parameters, the redundancy of the forward error correction code and the trigger threshold of the automatic retransmission request are adjusted to generate an optimized error correction scheme and obtain the repaired data packet. Network egress traffic is extracted from the repaired data packets. Through deep packet inspection technology, combined with content-aware analysis and context analysis, the header and content of the data packets are parsed to obtain traffic feature data. For the aforementioned traffic characteristic data, an abnormal traffic pattern recognition algorithm is used, combined with pre-established historical baseline data, to calculate the traffic deviation value, determine abnormal traffic behavior, and obtain an abnormal traffic identifier; If the abnormal traffic identifier indicates a potential data leak, the data content is extracted from the traffic feature data, and by matching it with the sensitive data fingerprint database, it is determined whether it contains sensitive information, thus obtaining the data leak detection result. Based on the data leakage detection results, an alarm log containing abnormal traffic identifiers and sensitive information matching details is generated and transmitted to the security management platform through a preset encrypted channel to obtain alarm processing instructions; In response to the alarm handling instructions, the network traffic control policy is automatically adjusted to limit the transmission bandwidth of abnormal destinations and update the sensitive data fingerprint database to obtain an optimized protection configuration.

2. The method according to claim 1, characterized in that, The process of obtaining the encoded data packet includes: Real-time data packets are acquired through a dedicated network interface, and initial error correction information is generated using forward error correction coding technology to obtain encoded data packets. Redundant verification information in the encoded data packet is verified using a preset error correction code rule to determine the verified data packet. If packet loss is detected in the verified data packet, the data is reconstructed using redundant verification information to obtain the reconstructed data packet; Based on the integrity of the reconstructed data packets, traffic analysis tools are used to detect data consistency and determine the consistency results. By adjusting the encoding parameters based on the consistency results, an optimized error-correcting code is generated, resulting in an optimized data packet. For optimized data packets, obtain their traffic characteristics during transmission and determine the changing trend of traffic characteristics; Based on the changing trends of traffic characteristics, the encoding strategy is adjusted through machine learning algorithms to obtain the final encoded data packets.

3. The method according to claim 1, characterized in that, The process of obtaining the retransmitted data packet includes: By analyzing bit errors in data packet transmission through an error detection mechanism, it can be determined whether there are erroneous bits and obtain the error detection result. If the error detection result indicates that there is a bit error, a retransmission request is generated and sent to the sender through the request mechanism to obtain the sender's response status; Based on the sender's response status, receive the retransmitted data packet, determine the integrity of the data packet, and obtain the retransmitted data packet; A data integrity verification algorithm is used to verify the retransmitted data packets to determine whether bit errors still exist and to obtain the verification result. If the verification result shows no bit errors, the retransmitted data packets are integrated into the transmission sequence to determine the continuity of the transmission process and obtain the updated data stream. The updated data stream is verified a second time using a cyclic redundancy check algorithm to determine the integrity of the data stream and obtain the final verification result. Based on the final verification results, the complete data packet sequence is output to obtain the retransmitted data packet.

4. The method according to claim 1, characterized in that, The process of obtaining adaptive error correction parameters includes: By storing historical error correction data, the distribution of error correction patterns is obtained, and the pattern analysis results are obtained. Using the pattern analysis results, data packet analysis features are extracted from encoded data and retransmitted data to obtain feature acquisition output. The output is obtained based on the characteristics, the statistical value of network interference is calculated, and the interference intensity value is obtained. If the interference intensity value exceeds the preset threshold, the calculation strategy is adjusted through intensity analysis to obtain the optimized intensity analysis result. Based on the optimized intensity analysis results, the adaptive parameters are calculated using a linear regression algorithm. The parameter calculation output is obtained, and combined with the current data packet analysis characteristics, the final adaptive error correction parameters are determined.

5. The method according to claim 1, characterized in that, The process of obtaining the repaired data packet includes: Obtain adaptive error correction parameters for the transmission environment, and determine the initial forward error correction code redundancy and automatic retransmission request trigger threshold by analyzing channel noise level and packet loss rate; If the channel noise level is higher than a preset threshold, the redundancy of the forward error correction code is increased, and the error correction code is generated by Reed-Solomon code to obtain a data packet with enhanced error correction capability. If the packet loss rate is higher than the preset threshold, the automatic retransmission request trigger threshold is lowered, and erroneous data packets are marked by the selective retransmission protocol to obtain a retransmission request sequence. Based on the data packets with enhanced error correction capabilities, a forward error correction decoding algorithm is used to repair the damaged data packets, resulting in a preliminary repaired data packet. For the initially repaired data packets, residual errors are detected through cyclic redundancy check to obtain error detection results; If the error detection result shows residual errors, the corresponding data packets are retransmitted using the automatic retransmission request mechanism according to the retransmission request sequence to obtain the final repaired data packets.

6. The method according to claim 1, characterized in that, The process of obtaining traffic characteristic data includes: Deep packet inspection technology is used to extract network outbound traffic from the repaired data packets, and the packet headers and contents are parsed to obtain the original traffic data. Content-aware analysis technology is used to parse the raw traffic data and extract the protocol type, source address and destination address from the data packets to obtain content feature data. By using context analysis technology, correlation analysis is performed on content feature data, and context feature data is obtained by combining the timestamp of data packets and session identifiers; If the protocol type in the context feature data is a preset protocol of interest, then the support vector machine algorithm is used to classify the context feature data, determine whether the traffic is abnormal, and obtain the abnormal traffic identifier. Based on the abnormal traffic identifier, the context feature data is filtered, and the data packet content corresponding to the abnormal traffic is extracted to obtain the abnormal traffic feature data. The key features are determined by ranking the abnormal traffic feature data according to the importance of the features using the random forest algorithm, thus obtaining the key traffic feature data. Using a preset threshold judgment method, if the frequency of the source address in the key traffic feature data exceeds the threshold, it is marked as high-risk traffic, and high-risk traffic feature data is obtained.

7. The method according to claim 1, characterized in that, The process of obtaining abnormal traffic identifiers includes: Traffic characteristic data is acquired, and the data is cleaned and standardized through a preprocessing module to obtain a standardized traffic dataset; The random forest algorithm is used in conjunction with pre-established historical baseline data to perform feature analysis on the normalized traffic dataset, generating traffic size deviation, time deviation, and destination deviation. The deviation value generation module calculates the difference between each deviation value and the historical baseline data to obtain the deviation feature vector. If the magnitude of the deviation feature vector is greater than the preset threshold, the abnormal traffic identification algorithm is used to classify the traffic data and determine abnormal traffic behavior. Based on the classification results of abnormal traffic behavior, abnormal traffic identifiers are generated and stored in the abnormal behavior database.

8. The method according to claim 1, characterized in that, The process of obtaining data breach detection results includes: If the traffic monitoring system detects a traffic anomaly, it obtains traffic characteristic data from the network traffic, extracts the data content through packet parsing technology, and obtains the first data content. A feature extraction algorithm is used to extract the structured features of the data content from the first data content, and a first feature set is generated. The first feature set is matched with a pre-established sensitive data fingerprint database. If the matching degree exceeds a preset threshold, it is determined that the data contains sensitive information, and the first matching result is obtained. Based on the first matching result, the decision tree algorithm is used to classify the possibility of data leakage and generate data leakage detection results; Extract contextual information of the data breach event from the data breach detection results, obtain related traffic source information through log analysis technology, and generate the first set of traffic sources; For the first set of traffic sources, anomaly traffic analysis technology is used to determine whether the traffic sources are continuously abnormal, and to obtain the abnormal traffic confirmation result; By jointly analyzing the abnormal traffic confirmation results and the first matching results, a final data breach detection report is generated.

9. An electronic device comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, characterized in that, When the processor executes the computing program, it implements the method of any one of claims 1-8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1-8.

Citation Information

Patent Citations

  • Vehicle bus intrusion detection system and method

    CN113612786A

  • Computer network information security control device

    CN118118257A