A method for identifying abnormal data transmission in network communication devices

By establishing a baseline range of transmission status parameters in network communication equipment, calculating longitudinal and lateral differences, and combining multidimensional feature data matching, the ambiguity and misjudgment problems of anomaly identification in existing technologies are solved, and accurate location and type identification of data transmission anomalies are achieved.

CN122093233BActive Publication Date: 2026-07-17SHANDONG CLOUD SKY SECURITY TECH CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG CLOUD SKY SECURITY TECH CO LTD
Filing Date
2026-04-27
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies cannot effectively combine the vertical changes in node status with the horizontal correlation of data flow paths when identifying abnormal data transmission in network communication devices, resulting in fuzzy or misjudged abnormality types and the inability to handle situations with multiple types of matching or no type matching.

Method used

By establishing a baseline range of node transmission status parameters, vertical anomaly marking is performed, the actual transmission path of the data flow is reconstructed, horizontal difference calculation is performed, and multidimensional feature data is matched with historical fault data to identify anomaly types.

Benefits of technology

It improves the accuracy of identifying abnormal node states, avoids the problem of inaccurate path information, reduces the false alarm rate, and achieves accurate location and type determination of data stream transmission paths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122093233B_ABST
    Figure CN122093233B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of network communication technology, specifically disclosing a method for identifying data transmission anomalies in network communication devices. The method includes: establishing a baseline interval for each node based on historical transmission status data; collecting real-time transmission status data and comparing it with the baseline interval, marking nodes with deviations as vertical anomalies; obtaining the next-hop node for each data stream from the real-time forwarding table, determining the actual transmission path, and identifying the upstream and downstream nodes of each node; performing a horizontal comparison between each node and its upstream and downstream nodes to generate horizontal difference markers; identifying nodes with both vertical and horizontal anomaly markers as anomalous nodes, collecting their multi-dimensional feature data with adjacent nodes, and matching this data with an anomaly type-feature value interval relationship constructed from historical fault data to identify the data transmission anomaly type. This invention improves the reliability of anomaly type identification through dual vertical and horizontal verification and multi-dimensional feature matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of network communication technology and relates to a method for identifying abnormal data transmission in network communication devices. Background Technology

[0002] With the continuous expansion of network scale and the increasing diversity of service types, the operational stability of network communication equipment and the reliability of data transmission have become crucial to ensuring service quality. During network operation, data transmission anomalies are difficult to completely avoid. Accurately and promptly identifying such anomalies and pinpointing their types is a technical problem that needs to be solved to improve network operation and maintenance efficiency and ensure business continuity.

[0003] Currently, industry methods for identifying network device data transmission anomalies typically employ node status parameter monitoring. For example, Chinese invention patent CN116319389B discloses a network anomaly identification method and related equipment. This method includes: when data transmission occurs via a time-sensitive network (TSN) and a configured time-sensitive switching device, calculating multiple sets of training network status parameter values ​​for the TSN; selecting multiple sets of anomalous status parameter values ​​from these training sets; generating a training dataset based on the multiple sets of anomalous status parameter values ​​and the anomaly labeling category corresponding to each set; obtaining a set of test network status parameter values ​​for the TSN; and classifying the test network status parameter values ​​according to a preset number and the training dataset to determine the network anomaly category of the TSN. This method enables rapid diagnosis of network anomalies.

[0004] However, existing technologies based on monitoring the status parameters of a single node have the following obvious limitations: First, existing technologies only identify network anomalies by changing the status parameters of a single node or device, failing to consider the complete forwarding link when the data stream passes through multiple nodes in the transmission path. At the same time, they cannot make a horizontal consistency comparison of the transmission status of the same data stream between upstream and downstream nodes, making it difficult to distinguish between node status anomalies and data stream path-level transmission anomalies.

[0005] Second, existing technologies only classify anomalies based on the node's own state parameter values, without combining the multi-dimensional feature data of the abnormal node and its neighboring nodes within the abnormal time window for comprehensive matching. At the same time, they cannot handle situations where the same node matches multiple anomaly types or cannot match any anomaly type, which leads to ambiguity or misjudgment risk in anomaly type identification.

[0006] Therefore, there is an urgent need for a comprehensive anomaly identification method that can combine the vertical changes in node status with the horizontal correlation of data flow paths, so as to achieve accurate location and type determination of data transmission anomalies and thus solve the above-mentioned technical problems. Summary of the Invention

[0007] In view of this, in order to solve the problems mentioned in the background technology, a method for identifying abnormal data transmission in network communication devices is proposed.

[0008] The objective of this invention can be achieved through the following technical solution: This invention provides a method for identifying abnormal data transmission in network communication devices, comprising: establishing a baseline range for each node's transmission status parameters based on historical transmission status data of each node in the network.

[0009] Collect real-time transmission status data from each node, compare the real-time transmission status data with the corresponding baseline interval, and mark nodes whose transmission status parameters deviate from their baseline interval as vertical anomalies.

[0010] The next-hop node for each data stream is obtained from the real-time forwarding table of each node. The next-hop nodes are connected in series to form the actual transmission path. The upstream and downstream nodes of each node are identified based on the actual transmission path.

[0011] Each node is compared horizontally with its corresponding upstream and downstream nodes. When horizontal differences exist, horizontal difference markers are generated.

[0012] Nodes that simultaneously exhibit both vertical anomalies and horizontal differences are identified as anomalous nodes. Multidimensional feature data of the anomalous nodes and their adjacent nodes are collected within the anomalous time window. By matching the multidimensional feature data with the anomaly type-feature value interval relationship constructed from historical fault data, the data transmission anomaly type of the anomalous nodes is identified.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention establishes the baseline interval of each node’s transmission status parameters based on historical transmission status data, and collects real-time transmission status data and compares it with the corresponding baseline interval. It marks the nodes with transmission status parameters that deviate from their baseline intervals vertically, thereby reducing the false alarm or missed alarm problem caused by using fixed thresholds, and thus improving the accuracy of identifying the node’s own abnormal status.

[0014] (2) By dynamically obtaining next-hop information from the real-time forwarding table of each node to reconstruct the actual transmission path of the data flow, this invention can truly reflect the current forwarding status of the data flow, effectively avoid the problem of inaccurate path information caused by network topology changes or route updates, thus providing a reliable path basis for horizontal consistency comparison and improving the accuracy of anomaly location.

[0015] (3) The present invention aligns the outbound flow characteristics of each node with the inbound flow characteristics of its downstream nodes in time, and aligns the inbound flow characteristics of each node with the outbound flow characteristics of its upstream nodes in time, calculates the downstream lateral difference and the upstream lateral difference respectively, and marks any node whose lateral difference exceeds a preset threshold, thereby effectively identifying the problem of inconsistent flow caused by node abnormalities in the data flow in the transmission path.

[0016] (4) This invention takes nodes with both vertical anomalies and horizontal differences as abnormal nodes, collects multidimensional feature data of abnormal nodes and their neighboring nodes within the abnormal time window, and matches the measured feature values ​​with the anomaly type-feature value interval relationship constructed by historical fault data to generate a candidate anomaly type set, thus avoiding false detection caused by single-dimensional judgment.

[0017] (5) This invention systematically handles various situations in anomaly type matching by directly identifying the candidate anomaly type set containing only one anomaly type as the identification result, filtering those containing multiple anomaly types based on the principle of minimum comprehensive deviation, and recording empty sets as unknown types, thereby solving the problem that the prior art cannot effectively handle multi-type matching or no-type matching. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram showing the connections between the steps of the method of the present invention; Figure 2 This is a schematic diagram showing the connection of the lateral difference marker generation steps in this invention; Figure 3 This is a schematic diagram illustrating the connection steps for identifying abnormal data transmission types in this invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] This invention achieves accurate identification and root cause localization of data transmission anomalies in network communication devices through a multi-stage collaborative mechanism involving longitudinal baseline comparison, lateral traffic consistency verification, and multi-dimensional feature matching. Specifically, the method first establishes dynamic baseline intervals for each node based on historical transmission status data, and then filters out nodes with deviating status by longitudinally comparing real-time data with the baselines. Next, it reconstructs the actual data flow transmission path based on real-time forwarding table information, identifies upstream and downstream nodes, and filters out nodes with lateral differences by performing time alignment and difference calculation on outbound and inbound traffic between upstream and downstream nodes. Finally, nodes exhibiting both longitudinal anomalies and lateral differences are identified as anomalous nodes, and their multi-dimensional feature data is collected and matched with the anomaly type-feature value interval relationship constructed from historical fault data to accurately determine the anomaly type. This method solves the technical problems of existing technologies that rely on single node status parameters, ignore path-level traffic consistency, and suffer from ambiguous anomaly type identification.

[0022] Please see Figure 1 As shown, the present invention provides a method for identifying abnormal data transmission in network communication devices, including the following steps S1 to S5.

[0023] S1. Establish the baseline range for each transmission status parameter.

[0024] Based on the historical transmission status data of each node in the network, a baseline range for each transmission status parameter of each node is established. This step aims to obtain the performance benchmark of each node under normal operating conditions, providing a basis for comparison in subsequent longitudinal anomaly detection.

[0025] For example, establishing the baseline range for each transmission status parameter includes: S1-1, obtaining historical transmission status data: obtaining transmission status data for each node's transmissions within a preset historical time period prior to the current moment from the historical transmission status data. The transmission status data includes, but is not limited to, inbound packet rate, outbound packet rate, packet loss rate, and CPU utilization. The preset historical time period is set according to the network operating cycle, for example, selecting data from the most recent 7 days or 30 days, to fully reflect the normal fluctuation range of nodes under different loads and time periods.

[0026] S1-2. Constructing a baseline interval: Sort the transmission status data of each node in each transmission within the preset historical time period according to the numerical value, and then take the preset percentile interval as the baseline interval for each transmission status parameter.

[0027] Specifically, different percentile range setting strategies are adopted for different types of transmission status parameters. For example, for parameters with large fluctuations, such as inbound and outbound packet rates, a two-sided range is used, with the lower limit set at the lower percentile (e.g., 5th percentile) and the upper limit at the higher percentile (e.g., 95th percentile) to accommodate normal business fluctuations; for parameters such as packet loss rate, which tend to be close to 0 under normal conditions, a single-sided upper limit range is used, setting only the upper limit threshold (e.g., 95th percentile); for parameters such as CPU utilization, which have a reasonable upper limit, a single-sided upper limit range is selected according to actual needs.

[0028] The specific value of the preset percentile is set based on the distribution characteristics of historical transmission status data and the stability requirements of the network operating environment. As an optional implementation, to overcome the problem that fixed percentiles may not adapt to the differences in data distribution among different nodes, a baseline interval can be dynamically determined using a statistical outlier detection method. For example, the first quartile of the historical transmission status data can be calculated. and the third and fourth quartiles and interquartile range Set the upper limit of the baseline interval to With k times The sum, with a lower limit set as and times The difference, of which This is an adjustment factor (usually 1.5 or 3) used to control the width of the baseline interval. This method adaptively adjusts the interval range based on the dispersion of the historical data itself, enabling more sensitive identification of abnormal fluctuations.

[0029] S2. Perform vertical anomaly marking.

[0030] Collect real-time transmission status data from each node, compare the real-time transmission status data with the corresponding baseline interval, and mark nodes whose transmission status parameters deviate from their baseline interval as vertical anomalies.

[0031] For example, the longitudinal anomaly marking includes: S2-1, real-time data comparison: comparing each transmission status parameter in the real-time transmission status data with the corresponding baseline interval to determine whether each transmission status parameter falls within its baseline interval range.

[0032] S2-2 Deviation Judgment: If the real-time value of any transmission status parameter exceeds the upper limit of its baseline interval or falls below the lower limit of its baseline interval, then the transmission status parameter is determined to have a deviation. For example, if the real-time value of the incoming packet rate of a node is 1500pps, while its baseline interval is [800, 1200]pps, then the incoming packet rate is determined to have a deviation.

[0033] S2-3, Node Marking: Mark nodes whose transmission status parameters are determined to be deviated as longitudinal anomalies.

[0034] S3. Determine the actual transmission path of the data stream and identify upstream and downstream nodes.

[0035] The next-hop node for each data stream is obtained from the real-time forwarding table of each node. These next-hop nodes are then concatenated to form the actual transmission path. Based on this actual transmission path, the upstream and downstream nodes of each node are identified. This step aims to reconstruct the complete forwarding link of the data stream in the network, providing an accurate topology for lateral consistency comparison.

[0036] For example, forming the actual transmission path includes: S3-1, obtaining the next-hop address: obtaining the next-hop address of each data stream on each node from the real-time forwarding table information of each node.

[0037] S3-2, Mapping the next-hop node: Based on the next-hop address of each data stream on each node, map the next-hop node corresponding to each next-hop address, and then determine the next-hop node of each data stream on each node.

[0038] S3-3, Path Concatenation: Starting from the source node of each data stream, sequentially find the next-hop node of each node until the destination node of each data stream is reached. The nodes traversed are then concatenated in sequence to form the actual transmission path of each data stream. If a loop or missing next hop occurs during the search process, the actual transmission path of that data stream cannot be constructed. This data stream is marked as a path-abnormal data stream, and the horizontal difference comparison step for that node is skipped. It is not included in the scope of abnormal node judgment, but its vertical abnormality mark is retained and recorded separately as a path abnormal event in subsequent steps for verification by operations and maintenance personnel.

[0039] For example, identifying the upstream and downstream nodes of each node includes: for each node on the actual transmission path, if the node is not the source node, then the node preceding the source node is taken as the upstream node.

[0040] If the node is not the destination node, then the next adjacent node of the node is taken as the downstream node.

[0041] S4. Generate horizontal difference markers.

[0042] Please see Figure 2 As shown, each node is compared horizontally with its corresponding upstream and downstream nodes. When there are horizontal differences, horizontal difference markers are generated.

[0043] For example, the generation of lateral difference markers includes: S4-1, obtaining traffic characteristics: obtaining the inbound traffic characteristics and outbound traffic characteristics of each node from the real-time transmission status data of each node. Wherein, the inbound traffic characteristic is the inbound packet rate of each node per unit time, and the outbound traffic characteristic is the outbound packet rate of each node per unit time.

[0044] S4-2. Calculate the downstream lateral variability: Align the outbound flow characteristics of each node with the inbound flow characteristics of its downstream nodes in terms of time. When the outbound flow characteristic is not zero, the ratio of the absolute value of the difference between the two to the outbound flow characteristic is used as the downstream lateral variability. Time alignment refers to matching the outbound flow characteristics of the upstream node with the inbound flow characteristics of the downstream node using the same timestamp to ensure that the compared flow data corresponds to the same time period. When the data sampling periods or time points of the upstream and downstream nodes are inconsistent, linear interpolation or the nearest neighbor time window matching method is used to align the data to a unified time axis. For example, if the upstream node reports the outbound rate every 10 seconds and the downstream node reports the inbound rate every 15 seconds, then the inbound rate data of the downstream node is linearly interpolated based on the time point of the upstream node to calculate the estimated value at the corresponding time.

[0045] When the outbound flow characteristic is 0, if the inbound flow characteristic of the downstream node is 0, then the downstream lateral difference is 0; otherwise, it is 1. The lateral difference ranges from [0, 1].

[0046] S4-3. Calculate the upstream lateral difference: Align the inflow characteristics of each node with the outflow characteristics of its upstream node in time; when the outflow characteristics of the upstream node are not 0, use the ratio of the absolute value of the difference between the two to the outflow characteristics as the upstream lateral difference.

[0047] When the outbound flow characteristic of the upstream node is 0, if the inbound flow characteristic of the same node is 0, then the upstream lateral difference is 0; otherwise, it is 1.

[0048] For the source node, there are no adjacent nodes in its upstream direction, so only its downstream lateral difference is calculated; for the destination node, there are no adjacent nodes in its downstream direction, so only its upstream lateral difference is calculated.

[0049] S4-4, Threshold Comparison and Marking: The upstream and downstream horizontal difference degrees are compared with preset thresholds, and nodes exceeding the thresholds are marked with horizontal difference. The horizontal difference threshold is determined by collecting the traffic difference values ​​between each node and its upstream and downstream nodes within a preset historical time period (during which there are no abnormal alarm records or the abnormality is confirmed by manual review). For example, the average of the traffic difference values ​​is used as the preset horizontal difference threshold.

[0050] Among them, the vertical anomaly marker is used to identify the deviation of the node's own transmission status from the normal baseline, realizing the detection of node anomalies; the horizontal difference marker is used to identify the inconsistency of traffic forwarding on the data flow transmission path, realizing the detection of path-level anomalies; the combination of the two can effectively distinguish between node own state anomalies and data flow path-level transmission anomalies, improving the accuracy of anomaly location.

[0051] S5. Identify the data transmission anomaly type of abnormal node.

[0052] Please see Figure 3 As shown, nodes that simultaneously have both vertical anomaly and horizontal difference markers are designated as anomalous nodes. Multidimensional feature data of the anomalous nodes and their adjacent nodes are collected within the anomalous time window. The anomalous time window is a time interval formed by extending a preset duration (e.g., 5 minutes before and after) forward and backward from the moment when the vertical anomaly or horizontal difference is first marked. By matching the multidimensional feature data with the anomaly type-feature value interval relationship constructed from historical fault data, the data transmission anomaly type of the anomalous node is identified.

[0053] The construction of the anomaly type-feature value interval relationship includes the following steps S5-1 to S5-3: S5-1, Historical data acquisition: Obtain the original state data of the abnormal nodes and their adjacent nodes in each historical anomaly event from the historical fault data, and perform feature engineering processing on them to extract multi-dimensional feature data for characterizing the anomaly pattern. The feature engineering processing includes, but is not limited to: calculating the moving average difference of the inbound packet rate within adjacent time windows to obtain the inbound packet rate change rate; calculating the ratio of the number of TCP retransmitted packets to the total number of sent packets per unit time to obtain the outbound retransmission rate; and obtaining the CPU utilization of each node. At the same time, obtain the anomaly type label corresponding to each historical anomaly event. The anomaly type label is marked according to the operation and maintenance records. The anomaly type includes, but is not limited to, link congestion, equipment failure, configuration error, and network attack.

[0054] S5-2, Grouping and Statistical Analysis: Group the extracted multidimensional feature data according to the anomaly type label, and group all multidimensional feature data corresponding to the same anomaly type into one group.

[0055] For each group of multidimensional feature data, statistical analysis is performed on each feature to determine the feature value range of each feature under the corresponding anomaly type, thus obtaining the multidimensional feature value range corresponding to each anomaly type.

[0056] Furthermore, determining the feature value range of each feature under the corresponding anomaly type includes: sorting the historical data values ​​of each feature in each group of multidimensional feature data from high to low.

[0057] Based on the distribution patterns of each dimension of features under each anomaly type in historical fault data, statistical analysis is used to determine the abnormal performance status of each dimension of features under each anomaly type. The abnormal performance status includes three types: numerical increase, numerical decrease, and numerical deviation from the normal range, which respectively characterize the changing trend of the feature when the corresponding anomaly type occurs.

[0058] If the value of this feature increases under abnormal conditions, a preset high quantile (such as the 95th quantile) is taken as the upper limit value on one side, forming a one-sided upper limit interval, that is, the feature value interval is (-∞, upper limit value). For example, taking packet loss rate as an example, when network link congestion, insufficient device processing capacity, or queue overflow occurs, data packets are dropped because they cannot be forwarded in time, resulting in a significant increase in packet loss rate. The higher the packet loss rate, the more severe the abnormality. Therefore, the interval of this type of feature is set as a one-sided upper limit value, that is, exceeding a certain quantile is considered to be a feature pattern that conforms to this abnormal type.

[0059] If the feature value decreases under abnormal conditions, a preset low quantile (e.g., the 5th percentile) is taken as the one-sided lower limit value, forming a one-sided lower limit interval, i.e., the feature value interval is [lower limit value, +∞). For example, taking the incoming packet rate as an example, when the upstream node fails, the link is interrupted, or the network attack causes the traffic to be blocked, the number of data packets received by the node drops sharply, and the incoming packet rate value decreases significantly. The lower the rate, the more severe the abnormality. Therefore, the interval of this type of feature is set as a one-sided lower limit value, that is, below a certain quantile is considered to be a feature pattern that conforms to this abnormality type.

[0060] If this feature deviates from the normal range under abnormal conditions (it may increase or decrease), a preset low quantile is taken as the lower limit of the interval, and a preset high quantile is taken as the upper limit of the interval, forming a closed interval composed of the lower and upper limits. For example, taking CPU utilization as an example, when the device is attacked or an abnormal process occurs, the CPU utilization may increase abnormally; while when the device experiences a deadlock or the processing module hangs, the CPU utilization may decrease abnormally. Therefore, both excessively high and excessively low values ​​are considered abnormal behaviors, and upper and lower thresholds must be set simultaneously to form a two-sided interval. That is, exceeding the range of this closed interval is considered to conform to the feature pattern of this abnormal type.

[0061] S5-3. Constructing Relationships: Associate each anomaly type with its multidimensional feature value range to construct anomaly type-feature value range relationships.

[0062] The identification of abnormal data transmission anomaly types of abnormal nodes includes the following steps S5-4 to S5-7: S5-4, Candidate type generation: The measured feature values ​​of the abnormal node in the multidimensional feature data are compared with the feature value intervals of the corresponding dimensions under each anomaly type. The number of dimensions in which the measured feature values ​​fall within the corresponding feature value interval is counted as the number of matching feature dimensions, and the total number of feature dimensions contained in the multidimensional feature data is recorded as the total number of feature dimensions.

[0063] The ratio of the number of matching feature dimensions to the total number of feature dimensions is used as the total matching degree. If the total matching degree is greater than or equal to the preset matching degree threshold, the anomaly type is added to the candidate anomaly type set of the anomaly node.

[0064] In a preferred implementation, the preset matching threshold is not a fixed value, but is dynamically determined based on historical fault data of the network region where the abnormal node is located: All confirmed abnormal event records for that network region within a preset historical time period (e.g., the previous 30 days) are acquired; for each historical abnormal event, its feature matching degree under each abnormality type is calculated; and a minimum matching threshold for that event is determined—that is, the lowest matching degree value that allows the true abnormality type of the event to be included in the candidate set. Specifically, when the system is first deployed or historical fault data is insufficient, a preset default threshold (e.g., 0.7) is used as the matching threshold; the dynamic threshold determination method is switched to after sufficient historical data has been accumulated.

[0065] All historical anomaly events are sorted by their minimum matching threshold from smallest to largest, and the 90th percentile is used as the final matching threshold. This method ensures that in historical data, 90% of anomaly events have a matching degree of no less than this threshold under their actual anomaly type, thus covering the vast majority of historical anomaly scenarios and avoiding missed detections due to improper threshold settings.

[0066] This dynamic determination method can adapt to the data distribution characteristics of different network regions, avoiding the excessively high false negative rate in some areas caused by using a globally fixed threshold.

[0067] S5-5 Single-type processing: If the candidate anomaly type set contains only one anomaly type, then that anomaly type shall be used as the data transmission anomaly type of the anomaly node.

[0068] S5-6. Multi-type processing: If the candidate anomaly type set contains multiple anomaly types, the weights of each feature are calculated based on the similarity of the anomaly node and its neighboring nodes in each feature dimension. Based on the deviation of each measured feature value and the feature value interval corresponding to each candidate anomaly type and the weights, an anomaly type is selected from the candidate anomaly type set as the data transmission anomaly type of the anomaly node.

[0069] Further, the step of selecting an anomaly type from the candidate anomaly type set includes: S5-6-1, for each candidate anomaly type in the candidate anomaly type set, calculating the deviation between each measured feature value of the anomaly node and the corresponding feature value interval under the candidate anomaly type.

[0070] It should be added that the feature value intervals are respectively a one-sided upper limit interval, a one-sided lower limit interval, or a closed interval.

[0071] For a one-sided upper limit interval, obtain the measured value and the one-sided upper limit value corresponding to the one-sided upper limit interval. When the measured value is greater than the one-sided upper limit value, divide the difference between the measured value and the one-sided upper limit value by the one-sided upper limit value to obtain the deviation. When the measured value is less than or equal to the one-sided upper limit value, take 0 as the deviation.

[0072] For a one-sided lower limit interval, obtain the measured value and the one-sided lower limit value corresponding to the one-sided lower limit interval. When the measured value is less than the one-sided lower limit value, divide the difference between the one-sided lower limit value and the measured value by the one-sided lower limit value to obtain the deviation. When the measured value is greater than or equal to the one-sided lower limit value, take 0 as the deviation.

[0073] For a closed interval, the measured value, the lower limit of the interval, and the upper limit of the interval are obtained. When the measured value is between the lower limit and the upper limit, 0 is taken as the deviation. When the measured value is less than the lower limit, the difference between the lower limit and the measured value is divided by the interval width to obtain the deviation. When the measured value is greater than the upper limit, the difference between the measured value and the upper limit is divided by the interval width to obtain the deviation. The interval width is the difference between the upper limit and the lower limit, and based on this, the feature value intervals for each dimension of the feature under the corresponding anomaly type are constructed.

[0074] S5-6-2. Based on the correlation feature values ​​between the abnormal node and its neighboring nodes, calculate the similarity of each feature between the abnormal node and its neighboring nodes, and determine the weight of each feature according to the similarity.

[0075] It should be noted that the weights of each feature dimension are dynamically determined based on the correlation feature values ​​between the abnormal node and its neighboring nodes. The specific steps are as follows: Obtain the correlation feature value pairs: For the first... For each dimension of a feature, the measured value of that feature at the abnormal node and its measured values ​​at each of its neighboring nodes are obtained. If there are multiple neighboring nodes, the average measured value of the multiple neighboring nodes is calculated to obtain the average measured value of the neighboring nodes, so as to comprehensively reflect the overall state of the neighboring nodes.

[0076] Normalization: To eliminate the impact of differences in dimensions and orders of magnitude between features of different dimensions on similarity calculation, a min-max normalization method is used to linearly transform the measured values ​​of each feature and the average measured value of adjacent nodes. The normalized measured values ​​of each feature and the average measured value of adjacent nodes are then denoted as follows: and , For each dimension of the feature, , This represents the total number of dimensions in the multidimensional feature data.

[0077] The reason for choosing a distance-based metric to calculate similarity is that it can intuitively reflect the absolute difference between two state points in multidimensional space, with clear physical meaning. For cases where the number of adjacent nodes is not fixed, when multiple adjacent nodes exist, the above method transforms many-to-one comparisons into one-to-one comparisons through mean-based processing, ensuring the uniformity and stability of the weight calculation logic across different network topologies.

[0078] Calculate similarity: Use distance metric to calculate the first similarity. The consistency between the anomaly node and its neighboring nodes is determined by a feature and converted into similarity. The calculation formula is as follows: .

[0079] In the formula, For the first The similarity of the dimensional features ranges from (0, 1). The closer the similarity is to 1, the more consistent the dimensional feature is between the abnormal node and its neighboring nodes; the closer the similarity is to 0, the greater the difference.

[0080] Weights are determined based on similarity: the similarity of each feature dimension is summed to obtain the total similarity, and the ratio of the similarity of each feature dimension to the total similarity is used as the weight of each feature dimension.

[0081] Through the above calculations, the higher the similarity of a feature dimension, the higher its weight, indicating that the feature behaves more consistently between the abnormal node and its neighboring nodes, and is more likely to be a path-level anomaly. Therefore, it should be given a greater contribution in the calculation of the overall deviation. Conversely, the lower the similarity of a feature dimension, the lower its weight, indicating that the feature may be an isolated anomaly of the node itself, and its contribution to the overall deviation is correspondingly reduced.

[0082] This dynamic weight determination method enables anomaly type identification to adaptively adjust according to the spatial distribution characteristics of each dimension of features in the actual anomaly scenario, avoiding the deviation that may be caused by preset fixed weights and improving the accuracy of anomaly type identification.

[0083] S5-6-3. Based on the weights of each feature dimension, the deviation is weighted and summed to obtain the comprehensive deviation of each candidate anomaly type.

[0084] It should be added that the formula for calculating the overall deviation is: .

[0085] In the formula, For the overall deviation, For the first Weights of dimensional features For the first The deviation corresponding to the dimensional feature.

[0086] By calculating the comprehensive deviation of each candidate anomaly type through weighted summation, on the one hand, the weight allocation can reflect the difference in the contribution of different features to the discrimination of the anomaly type, avoiding the weakening of the discriminative role of key features by treating each feature equally; on the other hand, it can integrate the deviation information of multiple dimensions into a comprehensive quantitative index, providing a unified evaluation basis for the screening of candidate anomaly types, thereby improving the accuracy of anomaly type identification.

[0087] S5-6-4. The candidate anomaly type with the smallest overall deviation is determined as the data transmission anomaly type of the anomaly node.

[0088] If multiple candidate anomaly types have the same overall deviation and are all at the minimum value, then the overall matching degree between each dimension of the corresponding candidate anomaly type and the feature value range is further compared, and the one with the highest overall matching degree is determined as the final anomaly type; if the overall matching degree is still the same, then all the candidate anomaly types are output together for manual review and confirmation. Here, the overall matching degree is the ratio of the number of matching feature dimensions to the total number of feature dimensions.

[0089] S5-7. Empty set processing: If the candidate anomaly type set is empty, the data transmission anomaly type of the anomaly node is recorded as an unknown type, and the manual review process is triggered.

[0090] Through steps S1 to S5 above, this invention constructs a complete process for identifying data transmission anomalies in network communication devices. This method filters abnormal nodes from both vertical historical benchmarks and horizontal path consistency dimensions, and combines multi-dimensional features with historical fault mode matching to complete anomaly type identification. This effectively solves the problems of inaccurate anomaly location and ambiguous type identification in existing technologies, providing reliable technical support for network operation and maintenance.

[0091] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0092] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0093] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0094] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0095] Finally, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for identifying abnormal data transmission in network communication devices, characterized in that: The method includes: Based on the historical transmission status data of each node in the network, a baseline range for each transmission status parameter of each node is established. Collect real-time transmission status data of each node, compare the real-time transmission status data with the corresponding baseline interval, and mark the nodes whose transmission status parameters deviate from their baseline interval as vertical anomalies. The next-hop node of each data stream is obtained from the real-time forwarding table of each node, the next-hop nodes are connected to form the actual transmission path, and the upstream and downstream nodes of each node are identified based on the actual transmission path. A horizontal comparison is performed between each node and its corresponding upstream and downstream nodes. When horizontal differences exist, a horizontal difference marker is generated. Nodes that simultaneously exhibit both vertical anomalies and horizontal differences are identified as anomalous nodes. Multidimensional feature data of the anomalous nodes and their adjacent nodes are collected within the anomalous time window. By matching the multidimensional feature data with the anomaly type-feature value interval relationship constructed from historical fault data, the data transmission anomaly type of the anomalous nodes is identified. The data transmission anomaly types identified by the abnormal nodes include: Each measured feature value of the abnormal node in the multidimensional feature data is compared with the feature value interval of the corresponding dimension under each abnormality type. The number of dimensions in which the measured feature values ​​fall within the corresponding feature value interval is counted as the number of matching feature dimensions. The total number of feature dimensions contained in the multidimensional feature data is recorded as the total number of feature dimensions. The ratio of the number of matching feature dimensions to the total number of feature dimensions is used as the total matching degree. If the total matching degree is greater than or equal to the preset matching degree threshold, the anomaly type is added to the candidate anomaly type set of the anomaly node. If the candidate anomaly type set contains multiple anomaly types, the weights of each feature are calculated based on the similarity of the anomaly node and its neighboring nodes in each feature dimension. Based on the deviation of each measured feature value from the feature value interval corresponding to each candidate anomaly type and the weights, an anomaly type is selected from the candidate anomaly type set as the data transmission anomaly type of the anomaly node. The step of selecting an anomaly type from the candidate anomaly type set includes: For each candidate anomaly type in the candidate anomaly type set, calculate the deviation between each measured feature value of the anomaly node and the corresponding feature value interval under the candidate anomaly type. Based on the correlation feature values ​​between the abnormal node and its neighboring nodes, the similarity of each feature dimension between the abnormal node and its neighboring nodes is calculated, and the weight of each feature dimension is determined according to the similarity. Based on the weights of each feature dimension, the deviation is weighted and summed to obtain the comprehensive deviation of each candidate anomaly type; The candidate anomaly type with the smallest overall deviation is taken as the data transmission anomaly type of the anomaly node; If multiple candidate anomaly types have the same overall deviation and are all the minimum, then the overall matching degree between each dimension feature and the feature value range corresponding to each candidate anomaly type is further compared, and the one with the highest overall matching degree is determined as the final anomaly type; if the overall matching degree is still the same, then the multiple candidate anomaly types are output together for manual review and confirmation.

2. The method for identifying abnormal data transmission in a network communication device according to claim 1, characterized in that: The baseline range for establishing each transmission state parameter includes: Obtain the transmission status data of each node for each transmission within the previous preset historical time period at the current moment from the historical transmission status data. The transmission status data of each node during each transmission within the preset historical time period are sorted by numerical value, and then a preset percentile range is taken as the baseline range for each transmission status parameter.

3. The method for identifying abnormal data transmission in a network communication device according to claim 1, characterized in that: Vertical anomaly labeling includes: Each transmission status parameter in the real-time transmission status data is compared with its corresponding baseline interval to determine whether each transmission status parameter falls within its baseline interval range. If the real-time value of any transmission status parameter exceeds the upper limit of its baseline interval or falls below the lower limit of its baseline interval, it is determined that the transmission status parameter has deviated. Nodes whose transmission status parameters are determined to be deviated are marked as longitudinal anomalies.

4. The method for identifying abnormal data transmission in a network communication device according to claim 1, characterized in that: The formation of the actual transmission path includes: Obtain the next-hop address of each data stream on each node from the real-time forwarding table information of each node; Based on the next-hop address of each data stream on each node, the next-hop node corresponding to each next-hop address is mapped, and then the next-hop node of each data stream on each node is determined. Starting from the source node of each data stream, the next-hop node of each node is found sequentially until the destination node of each data stream is reached. The nodes traversed are then connected in sequence to form the actual transmission path of each data stream.

5. The method for identifying abnormal data transmission in a network communication device according to claim 4, characterized in that: The identification of upstream and downstream nodes of each node includes: For each node on the actual transmission path, if the node is not the source node, then the node that precedes the source node is taken as the upstream node. If the node is not the destination node, then the next adjacent node of the node is taken as the downstream node.

6. The method for identifying abnormal data transmission in a network communication device according to claim 1, characterized in that: The generation of lateral difference markers includes: The inbound and outbound traffic characteristics of each node are obtained from the real-time transmission status data of each node. The outbound flow characteristics of each node are time-aligned with the inbound flow characteristics of its downstream nodes; when the outbound flow characteristic is not 0, the ratio of the absolute value of the difference between the two to the outbound flow characteristic is used as the downstream lateral difference. When the outbound flow characteristic is 0, if the inbound flow characteristic of the downstream node is 0, then the downstream lateral difference is 0; otherwise, it is 1. The inbound flow characteristics of each node are time-aligned with the outbound flow characteristics of its upstream node; when the outbound flow characteristics of the upstream node are not 0, the ratio of the absolute value of the difference between the two to the outbound flow characteristics is used as the upstream horizontal difference degree. When the outbound flow characteristic of the upstream node is 0, if the inbound flow characteristic of the same node is 0, then the upstream lateral difference is 0; otherwise, it is 1. The upstream and downstream horizontal differences are compared with preset thresholds, and nodes that exceed the thresholds are marked with horizontal differences.

7. The method for identifying abnormal data transmission in a network communication device according to claim 1, characterized in that: The construction of the anomaly type-feature value interval relationship includes: Extract multidimensional feature data of abnormal nodes and their adjacent nodes in each historical abnormal event from historical fault data, as well as the abnormal type label corresponding to each historical abnormal event; The extracted multidimensional feature data are grouped according to the anomaly type label, and all multidimensional feature data corresponding to the same anomaly type are grouped together. For each group of multidimensional feature data, statistical analysis is performed on each feature to determine the feature value range of each feature under the corresponding anomaly type, and the multidimensional feature value range corresponding to each anomaly type is obtained. Each anomaly type is associated with its multidimensional feature value range to construct an anomaly type-feature value range relationship.

8. A method for identifying abnormal data transmission in a network communication device according to claim 7, characterized in that: The determination of the feature value range of each dimension under the corresponding anomaly type includes: Sort the historical data values ​​of each feature in each group of multidimensional feature data from high to low; Based on the distribution patterns of each dimension of features under each anomaly type in historical fault data, statistical analysis is used to determine the abnormal performance of each dimension of features under each anomaly type. If the feature dimension shows an increase in value under abnormal conditions, then the preset high quantile is taken as the one-sided upper limit value to form a one-sided upper limit interval. If the feature dimension shows a decrease in value under abnormal conditions, then the preset low quantile is taken as the one-sided lower limit value to form a one-sided lower limit interval. If the feature dimension deviates from the normal range under abnormal conditions, a preset low quantile is taken as the lower limit of the interval and a preset high quantile is taken as the upper limit of the interval, forming a closed interval composed of the lower limit and the upper limit of the interval. Based on this, the feature value interval of each feature dimension under the corresponding abnormal type is constructed.

9. A method for identifying abnormal data transmission in a network communication device according to claim 1, characterized in that: If the candidate anomaly type set contains only one anomaly type, then that anomaly type shall be regarded as the data transmission anomaly type of the anomaly node. If the candidate anomaly type set is empty, then the data transmission anomaly type of the anomaly node is recorded as an unknown type.