Coal mine communication network anomaly detection method based on big data
Patent Information
- Application Number
- CN202610973562.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-15
Smart Images

Figure CN122764631A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network telemetry technology, and in particular to a method for detecting anomalies in coal mine communication networks based on big data. Background Technology
[0002] Against the backdrop of the ongoing advancement of intelligent coal mine construction, communication networks, as the core infrastructure for data acquisition, control command transmission, and emergency response, directly impact safe production and scheduling efficiency with their operational stability. In recent years, network status awareness technology based on telemetry data has been gradually applied to industrial control networks. By collecting device-level performance indicators and traffic characteristics, it enables dynamic assessment of network link quality and node health. However, coal mine communication networks are characterized by complex topologies, harsh deployment environments, and strong equipment heterogeneity. This makes it difficult for traditional telemetry analysis methods to accurately characterize the hop-by-hop propagation behavior of data packets in the network, and also hinders the effective correlation between abnormal events and physical space responsible units, thus limiting the practical application value of anomaly detection results in the operation and maintenance closed loop.
[0003] Existing network anomaly detection technologies based on big data primarily focus on traffic statistics or protocol compliance analysis, lacking the ability to structure and model path semantic information in telemetry data. This results in coarse-grained anomaly localization, making it difficult to map to specific equipment or roadway segment responsibility domains. Furthermore, high redundancy in telemetry fields and scattered critical path information easily lead to significant noise interference and weak interpretability in detection models. Most methods typically involve introducing static topology diagrams or manual rule bases for path reconstruction and responsibility allocation, but these methods rely on prior configuration and are ill-suited to the dynamic reconstruction and frequent equipment replacement realities of coal mine networks. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a big data-based method for detecting anomalies in coal mine communication networks to address the problem of adapting to the dynamic reconstruction of coal mine networks and frequent equipment replacements in real-world working conditions.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for detecting anomalies in coal mine communication networks based on big data, comprising, Construct a telemetry list for the mining network and generate an information-sensitive projection rule package, which includes a set of telemetry object fields, projection rules, and a mapping table between hop-by-hop relationships and object identifiers. Based on the information-sensitive projection rule package, information-sensitive projection is performed in the coal mine communication network to generate a hop-by-hop journey fingerprint data stream; Perform anomaly detection on the hop-by-hop journey fingerprint data stream and generate anomaly localization result packets; The abnormal location result packet is associated with the equipment responsibility domain and roadway section responsibility domain in the coal mine communication network operation and maintenance ledger to generate an abnormal alarm record. The abnormal location result packet is then archived according to the network path identifier to form an abnormal evidence archive.
[0007] As a preferred embodiment of the coal mine communication network anomaly detection method based on big data described in this invention, wherein: the generation of the information-sensitive projection rule package specifically comprises, Based on the network topology ledger and forwarding strategy ledger, determine the hop-by-hop relationship of each network path, and initially select hop-by-hop status fields on each network path to form a set of telemetry object fields. At the same time, generate an object identifier mapping table that records the globally unique identifiers of all network nodes and network links. Obtain the historical hop-by-hop state field sequence, perform multi-view conditional entropy calculation on the historical hop-by-hop state field sequence, and generate a separately distinguishable contribution value; Based on the individual contribution value, the smallest subset of fields is selected from the set of telemetry object fields; Based on the minimum subset of fields, projection rules are defined for each network path, specifying the field writing order, writing trigger conditions, and upper limit of writing length; The telemetry object field set, projection rules, hop-by-hop relationship and object identifier mapping table corresponding to each network path are encapsulated into an information-sensitive projection rule package.
[0008] As a preferred embodiment of the big data-based coal mine communication network anomaly detection method of the present invention, wherein: the network topology ledger is a standardized archive that records and manages the device connection relationships, node layout, and link configuration in the coal mine communication network; The forwarding policy ledger is a standardized archive that records and manages the forwarding rules, ACL policies, QoS configurations, and traffic routing of routers and switches in the coal mine communication network.
[0009] As a preferred embodiment of the coal mine communication network anomaly detection method based on big data described in this invention, wherein: the generation of hop-by-hop journey fingerprint data stream specifically involves, Extract the hop-by-hop status field from the telemetry object field set, and write the hop-by-hop status field into the telemetry segment carried with the business flow in the field writing order specified by the projection rules, and stop writing when the upper limit of the writing length specified by the projection rules is reached. Based on the hop-by-hop relationship, the written hop-by-hop state fields are aligned hop-by-hop and encapsulated with a unified fingerprint to generate a hop-by-hop journey fingerprint data stream.
[0010] As a preferred embodiment of the coal mine communication network anomaly detection method based on big data described in this invention, wherein: the generation of the anomaly location result packet specifically comprises, The hop-by-hop journey fingerprint data stream is aggregated according to the network path identifier and the time window identifier to generate an aggregated fingerprint dataset; Extract node stability feature vectors and link stability feature vectors from the aggregated fingerprint dataset; The node stability feature vector and link stability feature vector are compared with the constraint relationship map in the fingerprint baseline database to identify abnormal fingerprint segments; The abnormal fingerprint segment is divided into the minimum abnormal boundary along the direction of the jump sequence number in the jump sequence relationship to obtain the abnormal start jump and the abnormal end jump. Based on the hop-by-hop relationship and object identifier mapping table, abnormal starting hops and abnormal ending hops are converted into abnormal network node candidate sets and abnormal network link candidate sets, respectively. Extract the hop-by-hop state field sequence that triggers the anomaly from the abnormal fingerprint fragment as evidence fragment; The time window identifier, network path identifier, abnormal network node candidate set, abnormal network link candidate set, and evidence fragments are encapsulated into an anomaly localization result package.
[0011] As a preferred embodiment of the big data-based coal mine communication network anomaly detection method of the present invention, the network path identifier is a unique identity code assigned to each complete communication path from the starting point to the ending point in the coal mine communication network.
[0012] As a preferred embodiment of the coal mine communication network anomaly detection method based on big data described in this invention, the time window identifier is a time axis divided into segments of fixed duration, and a sequential number is generated for each time segment.
[0013] As a preferred embodiment of the coal mine communication network anomaly detection method based on big data described in this invention, the minimum anomaly boundary segmentation refers to locating and determining a continuous and shortest interval of jump numbers that can completely cover all anomaly state jumps along the jump number direction in the jump-by-jump relationship, and taking the starting jump number and ending jump number of the jump number interval as the anomaly start jump and the anomaly end jump, respectively.
[0014] As a preferred embodiment of the coal mine communication network anomaly detection method based on big data described in this invention, the generation of anomaly alarm records specifically involves: Extract candidate sets of abnormal network nodes and candidate sets of abnormal network links from the anomaly localization result packet; Associate the candidate set of abnormal network nodes and the candidate set of abnormal network links with the equipment responsibility domain and roadway section responsibility domain in the coal mine communication network operation and maintenance ledger; The associated equipment responsibility domain and roadway section responsibility domain are combined with the abnormal time window identifier, network path identifier, abnormal network node candidate set, abnormal network link candidate set and evidence fragment in the abnormal location result package to generate an abnormal alarm record.
[0015] As a preferred embodiment of the big data-based coal mine communication network anomaly detection method of the present invention, the step of associating the candidate set of abnormal network nodes and the candidate set of abnormal network links with the equipment responsibility domain and roadway section responsibility domain in the coal mine communication network operation and maintenance ledger specifically involves: Based on the network node identifiers in the candidate set of abnormal network nodes, query the equipment-responsibility domain mapping table in the coal mine communication network operation and maintenance ledger to obtain the equipment responsibility domain corresponding to the network node identifier; Based on the network link identifier in the candidate set of abnormal network links, query the network topology ledger in the coal mine communication network operation and maintenance ledger to obtain the roadway segment identifier. Based on the roadway segment identifier, query the roadway segment-responsibility domain mapping table to obtain the corresponding roadway segment responsibility domain. Add the equipment responsibility domain to the abnormal network node candidate set, add the roadway section responsibility domain to the abnormal network link candidate set, and generate the abnormal network node candidate set and abnormal network link candidate set of the associated responsibility domain.
[0016] The beneficial effects of this invention are as follows: By selecting the minimum subset of telemetry fields based on conditional entropy and embedding path semantics to construct an information-sensitive projection rule package, high-discrimination, low-overhead hop-by-hop status acquisition is achieved in anomaly detection of coal mine communication networks. This method abandons the traditional full-scale telemetry or static rule approach, quantifies the individual distinguishing contribution of each status field to the anomaly category based on historical operational data, dynamically retains the most discriminative field combination, and precisely binds it to the hop-by-hop topology of each network path; during the service flow forwarding process, only key fields are written as needed, which not only ensures the sensitivity of the journey fingerprint to abnormal behavior, but also effectively controls the length of telemetry segments, reducing the impact on normal services; the generated hop-by-hop journey fingerprint naturally possesses path alignment capability and physical location mappability, providing a clear and well-defined data foundation for subsequent anomaly localization, significantly improving detection accuracy and operation and maintenance response efficiency. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 The flowchart shows a method for detecting anomalies in coal mine communication networks based on big data. Figure 2 A flowchart for generating an information-sensitive projection rule package; Figure 3 A flowchart for generating hop-by-hop journey fingerprint data stream; Figure 4 A flowchart for generating an anomaly location result package. Detailed Implementation
[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0020] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0021] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0022] Reference Figures 1-4 This is one embodiment of the present invention, which provides a method for detecting anomalies in coal mine communication networks based on big data, including the following steps: S1. Construct a telemetry list for the mining network and generate an information-sensitive projection rule package. The information-sensitive projection rule package includes a set of telemetry object fields, projection rules, hop-by-hop relationships and object identifier mapping tables.
[0023] Based on the network topology ledger and forwarding policy ledger, the hop-by-hop traversal relationship of each network path is determined, and the hop-by-hop status fields are initially selected on each network path to form a set of telemetry object fields. At the same time, an object identifier mapping table is generated to record the globally unique identifiers of all network nodes and network links. The network topology ledger is a standardized archive for recording and managing the device connection relationships, node layout, and link configuration in the coal mine communication network. The forwarding policy ledger is a standardized archive for recording and managing the forwarding rules, ACL policies, QoS configurations, and traffic routing of routers and switches in the coal mine communication network.
[0024] Specifically, the process involves extracting the connection relationships and link configuration information between all network nodes from the network topology ledger, and extracting the forwarding rules, access control list policies, quality of service configurations, and traffic routing information of routers and switches from the forwarding policy ledger. Combining the device connection relationships in the network topology ledger with the forwarding rules in the forwarding policy ledger, the complete communication path from the starting point to the ending point in the coal mine communication network is enumerated one by one. For each network path, the network nodes and links traversed by the data packets are traced according to the forwarding rule order, recording the complete hop sequence from the starting point to the ending point, forming the hop-by-hop relationship of each network path. The forwarding rule order is the actual forwarding path order of data packets in the coal mine communication network determined by the forwarding rules, access control list policies, quality of service configurations, and traffic routing information of routers and switches. This order is obtained through path tracing simulation by combining the device connection relationships in the network topology ledger with the forwarding rules in the forwarding policy ledger.
[0025] After determining the hop-by-hop traversal relationship of each network path, hop-by-hop status fields are initially selected for each network path. Common status indicators of network nodes and network links are selected from the hop-by-hop traversal relationship of each network path, including queue length, interface error count, signal strength, bandwidth utilization, and packet loss rate, to form an initial list of hop-by-hop status fields. After deduplication of the initial list of hop-by-hop status fields for all network paths, they are merged to obtain the telemetry object field set.
[0026] In the process of determining the hop-by-hop relationship of all network paths and obtaining the set of telemetry object fields, a globally unique identifier is assigned to all network nodes in the coal mine communication network, and a globally unique identifier is assigned to all network links. The correspondence between network node names and globally unique identifiers, and the correspondence between network node pairs at both ends of a network link and globally unique identifiers are filled into the object identifier mapping table, generating an object identifier mapping table containing records of globally unique identifiers for all network nodes and network links.
[0027] Obtain the historical hop-by-hop state field sequence, perform multi-view conditional entropy calculation on the historical hop-by-hop state field sequence, and generate a separately distinguishing contribution value. The expression is: ; In the formula, To distinguish contribution values individually, a larger value indicates a stronger ability of the feature to distinguish categories (greater contribution), used to filter the smallest subset of fields. Let C be the entropy of the categorical variable, representing the overall category uncertainty. Given a feature f, the category conditional entropy represents the remaining uncertainty. C is the category variable, i.e., the operating state category of the coal mine communication network (normal state or abnormal state). f is the same feature (i.e., the hop-by-hop state field).
[0028] Based on the individual contribution value, the smallest subset of fields is selected from the set of telemetry object fields.
[0029] Specifically, for each hop-by-hop status field in the telemetry object field set, sort them from high to low according to their individual contribution values to form a hop-by-hop status field sequence in descending order of individual contribution values.
[0030] Starting from the first position of the hop-by-hop state field sequence in descending order of contribution value, hop-by-hop state fields are added to the candidate field subset sequentially. After each addition of the hop-by-hop state field with the highest current contribution value, the overall classification ability of the candidate field subset for the historical hop-by-hop state field sequence is evaluated to see if it has reached the preset coverage threshold.
[0031] When the overall class differentiation capability of the candidate field subset reaches or exceeds the preset coverage threshold for the first time, the addition of subsequent hop-by-hop status fields is stopped, and the current candidate field subset is determined as the minimum field subset.
[0032] Once all the hop-by-hop state field sequences that distinguish contribution values in descending order have been added, if the overall class distinguishing ability of the candidate field subset still does not reach the preset coverage threshold, then the candidate field subset containing all hop-by-hop state fields will be determined as the minimum field subset.
[0033] The preset coverage threshold is set based on historical experience in feature selection, and is usually set to 90% to 95% of the sum of individual distinguishing contribution values. For example, a value of 0.95 ensures that the smallest subset of fields retains most of the ability to distinguish categories while compressing the fields.
[0034] Based on the minimum field subset, information-sensitive projection rules are formulated for each network path, specifying the field writing order, writing trigger conditions, and upper limit of writing length. Specifically, the hop-by-hop status fields in the minimum field subset are sorted from high to low according to their individual distinguishing contribution values as the field writing order. The writing trigger conditions are set as changes in the status of each hop network node or sampling at fixed intervals (e.g., triggering a hop-by-hop status field write every 10 seconds or every 100 data packets processed). The upper limit of writing length is set as a fixed multiple of the number of fields in the minimum field subset, based on the number of hops in the network path and bandwidth limitations. The fixed multiple is typically 1 to 3 times the number of fields in the minimum field subset, ensuring that the total length of the telemetry segment is controlled within the bandwidth overhead allowable range of the average number of hops in the path, while covering sufficient hop-by-hop status field writing requirements. The allowable bandwidth overhead range is determined based on the link bandwidth capacity of the coal mine communication network and the normal transmission requirements of the service flow, and is typically controlled within 1% to 5% of the link bandwidth capacity, ensuring that the additional overhead of the telemetry segment does not affect the real-time performance and reliability of service flows such as safety monitoring data, control commands, or video surveillance.
[0035] The telemetry object field set, projection rules, hop-by-hop relationship and object identifier mapping table corresponding to each network path are encapsulated into an information-sensitive projection rule package.
[0036] It should be noted that by combining the network topology ledger and the forwarding strategy ledger, all network paths are enumerated and the hop-by-hop relationship of each path is determined. At the same time, hop-by-hop status fields are initially selected for each path and deduplicated and merged to form a set of telemetry object fields and an object identifier mapping table. This process is called constructing the mining network telemetry list. Based on historical data, contribution values are calculated separately, the smallest subset of fields is selected from the set of telemetry object fields, and projection rules are formulated for each network path. The set of telemetry object fields, projection rules, hop-by-hop relationships, and object identifier mapping table corresponding to each path are encapsulated into an information-sensitive projection rule package, thereby completing the generation of the mining network telemetry list into an information-sensitive projection rule package.
[0037] In simple terms, the mining network telemetry list is constructed through the process of generating a mapping table between the telemetry object field set and the object identifier. The information-sensitive projection rule package is generated by further incorporating projection rules and hop-by-hop relationships on the basis of the mining network telemetry list.
[0038] S2. Based on the information-sensitive projection rule package, perform information-sensitive projection in the coal mine communication network to generate a hop-by-hop journey fingerprint data stream.
[0039] Extract the hop-by-hop status field from the telemetry object field set, and write the hop-by-hop status field into the telemetry segment carried with the service flow in the field writing order specified by the projection rules, and stop writing when the upper limit of the writing length specified by the projection rules is reached.
[0040] Specifically, when a network node in a coal mine communication network forwards a service flow data packet, it loads the projection rule of the corresponding network path according to the information sensitive projection rule packet, and extracts the hop-by-hop status field contained in the smallest subset of the telemetry object field set as the hop-by-hop status field that needs to be written at the moment.
[0041] During the forwarding of service flow data packets along the network path, at each network node, the current network node collects the hop-by-hop status field values corresponding to its smallest subset of fields. According to the field writing order specified by the projection rules, the collected hop-by-hop status field values are sequentially appended to the header of the telemetry segment carried with the service flow or to a designated position. The designated position is the reserved extension field position or protocol header option field position in the telemetry segment carried with the service flow, used to accommodate the hop-by-hop appended status field values. Here, the service flow data packets are the data packets carried by service flows such as safety monitoring data, control commands, or video surveillance that are normally transmitted in the coal mine communication network.
[0042] During the writing process, the current network node continuously checks the cumulative length of the telemetry segments already written. When the cumulative length reaches the upper limit of the writing length specified by the projection rules, the current network node and subsequent network nodes stop adding new hop-by-hop status field values to the telemetry segments, keeping the telemetry segment length from increasing until the service flow data packets reach the destination.
[0043] Based on the hop-by-hop relationship, the written hop-by-hop state fields are aligned hop-by-hop and encapsulated with a unified fingerprint to generate a hop-by-hop journey fingerprint data stream.
[0044] Specifically, after the service flow data packet arrives at the end of the network path, the end network node reads the sequence of hop-by-hop status field values that have been written from the telemetry segment carried with the service flow.
[0045] The endpoint network node obtains the complete jump sequence from the starting point to the endpoint based on the hop-by-hop relationship of the corresponding network path in the information-sensitive projection rule package. It then matches the order of network nodes and network links in the complete jump sequence with the position of the written hop-by-hop status field value sequence to ensure that each written hop-by-hop status field value is strictly aligned to the position of the network node that generated the hop-by-hop status field value. For network node positions where no hop-by-hop status field value is written, it fills in null values to form an aligned hop-by-hop status field value sequence.
[0046] The endpoint network node combines the aligned hop-by-hop state field value sequence with the globally unique identifier in the network path identifier, timestamp, and object identifier mapping table, adds unified header information, including network path identifier, time window identifier, and sequence length, and encapsulates it into a complete hop-by-hop journey fingerprint record; all encapsulated hop-by-hop journey fingerprint records are continuously output to the big data collection channel to form a hop-by-hop journey fingerprint data stream.
[0047] S3. Perform anomaly detection on the hop-by-hop journey fingerprint data stream and generate anomaly location result packet.
[0048] The hop-by-hop journey fingerprint data stream is aggregated according to the network path identifier and the time window identifier to generate an aggregated fingerprint dataset. The network path identifier is a unique identification code assigned to each complete communication path from the start point to the end point in the coal mine communication network. The time window identifier is a sequential number generated for each time period by dividing the time axis according to a fixed duration.
[0049] Specifically, after the hop-by-hop journey fingerprint data stream enters the big data collection channel, all hop-by-hop journey fingerprint records are grouped according to the network path identifier in the hop-by-hop journey fingerprint record, forming an independent record set based on the network path identifier.
[0050] Within the independent record set corresponding to each network path identifier, the hop-by-hop journey fingerprint records are grouped a second time according to the time window identifier in the hop-by-hop journey fingerprint record. The time window identifier is generated by dividing the time axis with a fixed duration, such as generating a continuous sequential number every 5 minutes or every 10 minutes, to ensure that the hop-by-hop journey fingerprint records within the same time window identifier correspond to the same time period.
[0051] For each subset of hop-by-hop journey fingerprint records under each combination of network path identifier and time window identifier, collect all aligned hop-by-hop state field value sequences to form an aggregated fingerprint record list.
[0052] The aggregated fingerprint record list, which combines all network path identifiers and all time window identifiers, is brought together to form a complete aggregated fingerprint dataset. The aggregated fingerprint dataset contains a fully aligned sequence of hop-by-hop state field values organized by network path identifier and time window identifier.
[0053] Extract node stability feature vectors and link stability feature vectors from the aggregated fingerprint dataset.
[0054] Specifically, the aggregate fingerprint record list under each combination of network path identifier and time window identifier in the aggregate fingerprint dataset contains multiple aligned hop-by-hop state field value sequences.
[0055] Traverse all aligned hop-by-hop state field value sequences in the aggregated fingerprint record list, separating the hop-by-hop state field values corresponding to all network nodes and the hop-by-hop state field values corresponding to all network links according to the hop-by-hop index. For each network node position, collect the hop-by-hop state field values of all aligned hop-by-hop state field value sequences in the aggregated fingerprint record list at the network node position, forming a node value distribution sequence of the network node within the current time window. Take the median of the node value distribution sequence as the stable feature value of the network node, and arrange the stable feature values of all network nodes in hop-by-hop index order to form a node stable feature vector.
[0056] For each network link location, collect the hop-by-hop state field value sequence of all aligned hop-by-hop state field values in the aggregate fingerprint record list to form the link value distribution sequence of the network link in the current time window. Take the median of the link value distribution sequence as the stable feature value of the network link. Arrange all the stable feature values of the network links in hop-by-hop order to form the link stability feature vector.
[0057] The stable feature vectors of nodes and links are compared with the constraint relationship graph in the fingerprint baseline database to identify abnormal fingerprint segments. The fingerprint baseline database is constructed based on the aggregated fingerprint dataset of the coal mine communication network under historical normal operation conditions. The fingerprint baseline database pre-stores the baseline range of the stable feature vectors of nodes and links corresponding to each network path identifier under normal operation conditions, as well as the constraint relationship graph between the stable feature vectors of nodes and links. The constraint relationship graph records the association rules of stable feature values of adjacent network node positions and network link positions under normal conditions. The stable feature value association rule is a rule that records the expected association relationship between the stable feature values of adjacent network node positions and network link positions under normal operation conditions, and is used to detect whether the actual stable feature value combination deviates from the normal pattern.
[0058] Specifically, for the node stable feature vector under the current network path identifier and the current time window identifier, the stable feature value of each network node position is matched and compared with the benchmark range of the node stable feature vector of the corresponding network path identifier in the fingerprint baseline database. When the stable feature value exceeds the benchmark range of the node stable feature vector, the corresponding network node position is marked as an abnormal node position. The benchmark range of the node stable feature vector is set based on the statistical characteristics of the node value distribution sequence under historical normal operation. It is usually taken as ±2 standard deviations of the median of the node value distribution sequence. The basis for the value is based on the empirical rule that, assuming that the node value distribution sequence approximately follows a normal distribution, the ±2 standard deviations of the median of the node value distribution sequence can cover about 95% of the observed values in the distribution, thereby effectively defining the boundary of normal fluctuations and identifying potential anomalies.
[0059] For the link stability feature vector under the current network path identifier and the current time window identifier, the stability feature value of each network link location is matched and compared with the baseline range of the link stability feature vector of the corresponding network path identifier in the fingerprint baseline database. When the stability feature value exceeds the baseline range of the link stability feature vector, the corresponding network link location is marked as an abnormal link location. The baseline range of the link stability feature vector is set based on the statistical characteristics of the link value distribution sequence under historical normal operation conditions. It is usually taken as ±2 standard deviations of the median of the link value distribution sequence. The basis for the value is based on the empirical rule that, assuming that the link value distribution sequence approximately follows a normal distribution, the ±2 standard deviations of the median of the link value distribution sequence can cover about 95% of the observations in the distribution, thereby effectively defining the boundary of normal fluctuations and identifying potential anomalies.
[0060] Simultaneously, the stable feature values in the current node's stable feature vector and the stable feature vector of the link are matched and checked against the association rules in the constraint relationship graph. When the combination of stable feature values of adjacent network node positions and network link positions violates the association rules in the constraint relationship graph, the corresponding network node positions and network link positions are marked as abnormal node positions and abnormal link positions; when the combination of stable feature values of adjacent network node positions and network link positions conforms to the association rules in the constraint relationship graph, the corresponding network node positions and network link positions are kept as normal positions.
[0061] For all hop sequence intervals marked as abnormal node locations and abnormal link locations, extract the aligned hop status field value sequence of the corresponding positions from the aggregated fingerprint record list to form an abnormal fingerprint fragment.
[0062] It should be noted that the fingerprint baseline database was constructed based on the aggregated fingerprint dataset under the historical normal operation state of the coal mine communication network. During the construction of the fingerprint baseline database, the node stable feature vectors and link stable feature vectors corresponding to all network path identifiers were extracted from the aggregated fingerprint dataset under the historical normal operation state. Statistical analysis was performed on the node stable feature vectors corresponding to all network path identifiers to determine the baseline range of stable feature values for each network node. Statistical analysis was also performed on the link stable feature vectors corresponding to all network path identifiers to determine the baseline range of stable feature values for each network link. At the same time, correlation analysis was performed on the node stable feature vectors and link stable feature vectors corresponding to all network path identifiers. The stable feature value combinations of adjacent network node positions and network link positions were traversed, and the stable feature value combination patterns that repeatedly appeared under normal conditions were extracted as association rules. All extracted association rules were recorded in the constraint relationship graph.
[0063] The constraint relationship graph contains stable feature value association rules for the positions of adjacent network nodes and network links under normal conditions, which are used to detect whether the combination of stable feature values deviates from the normal pattern during subsequent comparisons; the fingerprint baseline library pre-stores the reference range of stable feature vectors of nodes, reference range of stable feature vectors of links, and constraint relationship graphs corresponding to all network path identifiers.
[0064] The abnormal fingerprint segment is divided into minimum abnormal boundary segments along the direction of the jump sequence number in the jump sequence relationship to obtain the abnormal start jump and the abnormal end jump. The minimum abnormal boundary segmentation means to locate and determine a continuous and shortest jump number interval that can completely cover all abnormal state jumps along the direction of the jump sequence number in the jump sequence relationship, and take the start jump number and end jump number of the jump number interval as the abnormal start jump and abnormal end jump respectively.
[0065] Specifically, the abnormal fingerprint fragment contains a sequence of aligned hop-by-hop state field values, marked as the locations of abnormal nodes and abnormal links, arranged in ascending order of hop number.
[0066] Based on the hop-by-hop relationship of the corresponding network path in the information-sensitive projection rule package, obtain the hop-by-hop sequence number order of the complete jump sequence. Starting from the minimum hop-by-hop number, traverse all position markers in the abnormal fingerprint segment in the direction of the maximum hop-by-hop number. During the traversal, find the hop-by-hop number of the first abnormal node position or abnormal link position as the candidate abnormal starting hop. Continue traversing in the direction of increasing hop-by-hop number until the hop-by-hop number of the last abnormal node position or abnormal link position is found as the candidate abnormal ending hop.
[0067] Extend the candidate anomaly starting point jump to the previous hop sequence position and the candidate anomaly ending point jump to the next hop sequence position to ensure that the extended interval fully covers all anomaly node positions and anomaly link positions, while keeping the interval length as short as possible; determine the starting hop sequence number of the extended interval as the anomaly starting hop and the ending hop sequence number of the extended interval as the anomaly ending hop.
[0068] Among them, the abnormal starting point jump and the abnormal ending point jump correspond to the interval of the smallest consecutive jump number in the same abnormal fingerprint segment.
[0069] Based on the hop-by-hop relationship and object identifier mapping table, abnormal starting hops and abnormal ending hops are converted into abnormal network node candidate sets and abnormal network link candidate sets, respectively.
[0070] Specifically, the abnormal start jump and abnormal end jump correspond to the consecutive minimum jump interval in the same abnormal fingerprint segment, and the jump sequence number starts from the abnormal start jump and ends at the abnormal end jump.
[0071] Based on the hop-by-hop relationship of the corresponding network path in the information-sensitive projection rule package, obtain the list of network node locations and the list of network link locations in the hop interval from the abnormal starting point to the abnormal ending point in the complete hop sequence.
[0072] Match the hop sequence numbers of all network node positions within the hop interval from the abnormal starting point to the abnormal ending point with the object identifier mapping table. Query the correspondence between network node names and globally unique identifiers in the object identifier mapping table, extract the globally unique identifiers of all matching network nodes, and form a candidate set of abnormal network nodes.
[0073] Match the hop sequence numbers of all network link positions within the hop interval from the abnormal starting point to the abnormal ending point with the object identifier mapping table. Query the correspondence between the network node pairs at both ends of the network link and the globally unique identifier in the object identifier mapping table, extract the globally unique identifiers of all matching network links, and form a candidate set of abnormal network links.
[0074] The candidate set of abnormal network nodes contains globally unique identifiers of all network nodes within the range from the abnormal starting point hop to the abnormal ending point hop, and the candidate set of abnormal network links contains globally unique identifiers of all network links within the range from the abnormal starting point hop to the abnormal ending point hop.
[0075] Extract the hop-by-hop state field sequence that triggered the anomaly from the abnormal fingerprint fragment as evidence fragment.
[0076] Specifically, the abnormal fingerprint segment contains multiple aligned hop-by-hop state field value sequences, arranged in ascending order of hop number, with all marked abnormal node positions and abnormal link positions concentrated within the interval of the minimum consecutive hop count from the abnormal starting hop to the abnormal ending hop; all aligned hop-by-hop state field value sequences corresponding to the interval from the abnormal starting hop to the abnormal ending hop are selected from the abnormal fingerprint segment.
[0077] Within the hop interval from the abnormal starting point to the abnormal ending point, traverse the hop-by-hop status field values of all network node positions and network link positions, collect hop-by-hop status field values that exceed the baseline range of node stability feature vectors or link stability feature vectors, as well as hop-by-hop status field values corresponding to adjacent stable feature value combinations that violate the constraint relationship graph association rules; arrange all collected hop-by-hop status field values that exceed the baseline range or violate the association rules in the order of the original hop-by-hop sequence number to form the hop-by-hop status field sequence that triggers the abnormality.
[0078] The hop-by-hop state field sequence that triggers the anomaly serves as evidence fragment, directly corresponding to the specific field value content in the anomaly fingerprint fragment that caused the anomaly to be marked.
[0079] The time window identifier, network path identifier, abnormal network node candidate set, abnormal network link candidate set, and evidence fragments are encapsulated into an anomaly localization result package.
[0080] S4. Associate the anomaly location result packet with the equipment responsibility domain and roadway section responsibility domain in the coal mine communication network operation and maintenance ledger, generate anomaly alarm records, and archive them together with the anomaly location result packet according to the network path identifier to form anomaly evidence archive.
[0081] Extract candidate sets of abnormal network nodes and candidate sets of abnormal network links from the anomaly localization result package.
[0082] The candidate sets of abnormal network nodes and abnormal network links are associated with the equipment responsibility domain and roadway section responsibility domain in the coal mine communication network operation and maintenance ledger.
[0083] Specifically, based on the network node identifier in the candidate set of abnormal network nodes, query the equipment-responsibility domain mapping table in the coal mine communication network operation and maintenance ledger to obtain the equipment responsibility domain corresponding to the network node identifier; Specifically, the process involves: traversing the globally unique identifier of each network node in the candidate set of abnormal network nodes; for the currently traversed globally unique identifier of a network node, querying the corresponding network node name in the object identifier mapping table; using the queried network node name, matching the corresponding record in the equipment-responsibility domain mapping table of the coal mine communication network operation and maintenance ledger, and extracting the equipment responsibility domain information from the matching record; associating the extracted equipment responsibility domain information with the currently traversed globally unique identifier of a network node to form the equipment responsibility domain corresponding to the globally unique identifier of the network node; repeating the traversal and query process until all globally unique identifiers of network nodes in the candidate set of abnormal network nodes have obtained the corresponding equipment responsibility domain.
[0084] Based on the network link identifiers in the candidate set of abnormal network links, the network topology ledger in the coal mine communication network operation and maintenance ledger is queried to obtain the roadway segment identifier. Based on the roadway segment identifier, the roadway segment-responsibility domain mapping table is queried to obtain the corresponding roadway segment responsibility domain.
[0085] Specifically, the process involves: traversing each globally unique identifier of a network link in the candidate set of abnormal network links; for the currently traversed globally unique identifier of a network link, querying the corresponding network node pair at both ends of the network link in the object identifier mapping table; using the obtained network node pair at both ends of the network link, matching the corresponding record in the network topology ledger of the coal mine communication network operation and maintenance ledger, and extracting the roadway segment identifier from the matching record; using the extracted roadway segment identifier, matching the corresponding record in the roadway segment-responsibility domain mapping table of the coal mine communication network operation and maintenance ledger, and extracting the roadway segment responsibility domain information from the matching record; associating the extracted roadway segment responsibility domain information with the currently globally unique identifier of the network link to form the roadway segment responsibility domain corresponding to the globally unique identifier of the network link; repeating the traversal and query process until all globally unique identifiers of network links in the candidate set of abnormal network links have obtained the corresponding roadway segment responsibility domain.
[0086] Add the equipment responsibility domain to the abnormal network node candidate set, add the roadway section responsibility domain to the abnormal network link candidate set, and generate the abnormal network node candidate set and abnormal network link candidate set of the associated responsibility domain.
[0087] The associated equipment responsibility domain and roadway section responsibility domain are combined with the abnormal time window identifier, network path identifier, abnormal network node candidate set, abnormal network link candidate set and evidence fragment in the abnormal location result package to generate an abnormal alarm record.
[0088] The abnormal alarm records and abnormal location result packets are archived according to the network path identifier to form an abnormal evidence archive.
[0089] Specifically, the abnormal alarm records and abnormal location result packets are grouped according to the network path identifier in the abnormal alarm records and abnormal location result packets. Abnormal alarm records and abnormal location result packets with the same network path identifier are placed in the same archive set. An independent storage location is allocated to the archive set corresponding to each network path identifier, and the storage location is named according to the network path identifier.
[0090] Within the storage location corresponding to each network path identifier, abnormal alarm records and corresponding abnormal location result packets are stored sequentially according to the abnormal time window identifier, forming an archived record arranged in time series.
[0091] All archive sets corresponding to network path identifiers are brought together to form a complete anomaly evidence archive. The anomaly evidence archive contains persistent storage content organized by network path identifier for all anomaly alarm records and anomaly location result packages.
[0092] This embodiment also provides a computer device applicable to the case of an anomaly detection method for coal mine communication networks based on big data, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the anomaly detection method for coal mine communication networks based on big data as proposed in the above embodiment.
[0093] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0094] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the big data-based coal mine communication network anomaly detection method proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0095] In summary, this invention achieves high-discrimination, low-overhead hop-by-hop status acquisition in anomaly detection of coal mine communication networks by: selecting a minimum subset of telemetry fields based on conditional entropy and embedding them into path semantics to construct an information-sensitive projection rule package. This method abandons traditional full-scale telemetry or static rule-based approaches, quantifying the individual discriminative contribution of each status field to anomaly categories based on historical operational data, dynamically retaining the most discriminative field combinations, and precisely binding them to the hop-by-hop topology of each network path. During service flow forwarding, only key fields are written as needed, ensuring the sensitivity of the journey fingerprint to anomalies while effectively controlling the length of telemetry segments and reducing the impact on normal services. The generated hop-by-hop journey fingerprint naturally possesses path alignment capabilities and physical location mappability, providing a clear and well-defined data foundation for subsequent anomaly localization, significantly improving detection accuracy and operational response efficiency.
[0096] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A coal mine communication network anomaly detection method based on big data, characterized by: include, Construct a telemetry list for the mining network and generate an information-sensitive projection rule package, which includes a set of telemetry object fields, projection rules, and a mapping table between hop-by-hop relationships and object identifiers. Based on the information-sensitive projection rule package, information-sensitive projection is performed in the coal mine communication network to generate a hop-by-hop journey fingerprint data stream; Perform anomaly detection on the hop-by-hop journey fingerprint data stream and generate anomaly localization result packets; The abnormal location result packet is associated with the equipment responsibility domain and roadway section responsibility domain in the coal mine communication network operation and maintenance ledger to generate an abnormal alarm record. The abnormal location result packet is then archived according to the network path identifier to form an abnormal evidence archive.
2. The big data-based coal mine communication network anomaly detection method of claim 1, wherein: The generated information-sensitive projection rule package specifically includes: Based on the network topology ledger and forwarding strategy ledger, determine the hop-by-hop relationship of each network path, and initially select hop-by-hop status fields on each network path to form a set of telemetry object fields. At the same time, generate an object identifier mapping table that records the globally unique identifiers of all network nodes and network links. Obtain the historical hop-by-hop state field sequence, perform multi-view conditional entropy calculation on the historical hop-by-hop state field sequence, and generate a separately distinguishable contribution value; Based on the individual contribution value, the smallest subset of fields is selected from the set of telemetry object fields; Based on the minimum subset of fields, projection rules are defined for each network path, specifying the field writing order, writing trigger conditions, and upper limit of writing length; The telemetry object field set, projection rules, hop-by-hop relationship and object identifier mapping table corresponding to each network path are encapsulated into an information-sensitive projection rule package.
3. The big data-based coal mine communication network anomaly detection method of claim 2, wherein: The network topology ledger is a standardized archive that records and manages the device connection relationships, node layout, and link configuration in the coal mine communication network. The forwarding policy ledger is a standardized archive that records and manages the forwarding rules, ACL policies, QoS configurations, and traffic routing of routers and switches in the coal mine communication network.
4. The method for detecting anomalies in coal mine communication networks based on big data according to claim 1, characterized in that: The generation of the hop-by-hop journey fingerprint data stream is specifically as follows: Extract the hop-by-hop status field from the telemetry object field set, and write the hop-by-hop status field into the telemetry segment carried with the business flow in the field writing order specified by the projection rules, and stop writing when the upper limit of the writing length specified by the projection rules is reached. Based on the hop-by-hop relationship, the written hop-by-hop state fields are aligned hop-by-hop and encapsulated with a unified fingerprint to generate a hop-by-hop journey fingerprint data stream.
5. The method for detecting anomalies in coal mine communication networks based on big data according to claim 1, characterized in that: The generation of the anomaly localization result package specifically includes: The hop-by-hop journey fingerprint data stream is aggregated according to the network path identifier and the time window identifier to generate an aggregated fingerprint dataset; Extract node stability feature vectors and link stability feature vectors from the aggregated fingerprint dataset; The node stability feature vector and link stability feature vector are compared with the constraint relationship map in the fingerprint baseline database to identify abnormal fingerprint segments; The abnormal fingerprint segment is divided into the minimum abnormal boundary along the direction of the jump sequence number in the jump sequence relationship to obtain the abnormal start jump and the abnormal end jump. Based on the hop-by-hop relationship and object identifier mapping table, abnormal starting hops and abnormal ending hops are converted into abnormal network node candidate sets and abnormal network link candidate sets, respectively. Extract the hop-by-hop state field sequence that triggers the anomaly from the abnormal fingerprint fragment as evidence fragment; The time window identifier, network path identifier, abnormal network node candidate set, abnormal network link candidate set, and evidence fragments are encapsulated into an anomaly localization result package.
6. The method for detecting anomalies in coal mine communication networks based on big data according to claim 5, characterized in that: The network path identifier is a unique identification code assigned to each complete communication path from the origin to the destination in the coal mine communication network.
7. The method for detecting anomalies in coal mine communication networks based on big data according to claim 5, characterized in that: The time window identifier is a sequential number generated for each time period by dividing the time axis into segments of fixed duration.
8. The method for detecting anomalies in coal mine communication networks based on big data according to claim 5, characterized in that: The minimum anomaly boundary segmentation refers to locating and determining a continuous and shortest interval of jumps that can completely cover all anomaly jumps along the jump sequence direction in the jump sequence relationship, and taking the starting jump sequence and ending jump sequence of the jump interval as the anomaly start jump and anomaly end jump, respectively.
9. The method for detecting anomalies in coal mine communication networks based on big data according to claim 1, characterized in that: The generation of abnormal alarm records specifically refers to... Extract candidate sets of abnormal network nodes and candidate sets of abnormal network links from the anomaly localization result packet; Associate the candidate set of abnormal network nodes and the candidate set of abnormal network links with the equipment responsibility domain and roadway section responsibility domain in the coal mine communication network operation and maintenance ledger; The associated equipment responsibility domain and roadway section responsibility domain are combined with the abnormal time window identifier, network path identifier, abnormal network node candidate set, abnormal network link candidate set and evidence fragment in the abnormal location result package to generate an abnormal alarm record.
10. The method for detecting anomalies in coal mine communication networks based on big data according to claim 9, characterized in that: The step of associating the candidate set of abnormal network nodes and the candidate set of abnormal network links with the equipment responsibility domain and roadway section responsibility domain in the coal mine communication network operation and maintenance ledger specifically involves... Based on the network node identifiers in the candidate set of abnormal network nodes, query the equipment-responsibility domain mapping table in the coal mine communication network operation and maintenance ledger to obtain the equipment responsibility domain corresponding to the network node identifier; Based on the network link identifier in the candidate set of abnormal network links, query the network topology ledger in the coal mine communication network operation and maintenance ledger to obtain the roadway segment identifier. Based on the roadway segment identifier, query the roadway segment-responsibility domain mapping table to obtain the corresponding roadway segment responsibility domain. Add the equipment responsibility domain to the abnormal network node candidate set, add the roadway section responsibility domain to the abnormal network link candidate set, and generate the abnormal network node candidate set and abnormal network link candidate set of the associated responsibility domain.