A communication network operation state monitoring method and system

By collecting and transforming multi-source data and combining it with historical patterns in industrial plants, association rules are formulated to attribute anomalies in real time. Edge computing and predictive models are used to achieve closed-loop management of prediction and maintenance, which solves the problems of large errors and inability to predict faults in the monitoring of communication networks in industrial plants, and improves the accuracy of judgment and production stability.

CN120979971BActive Publication Date: 2026-01-06SHANGHAI TECH NETWORK COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511493416.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2026-01-06
Estimated Expiration
2045-10-20

AI Technical Summary

Technical Problem

Existing technologies for monitoring communication networks in industrial plants fail to effectively combine local environmental data with equipment maintenance records, resulting in large errors, an inability to predict faults in advance, and disruption to production continuity.

Method used

By collecting data from multiple sources, preprocessing and transforming data, and combining historical disconnection patterns in the plant area, we formulate association rules for signal-hardware-congestion, use edge computing for real-time anomaly attribution, and link it with the prediction model to achieve closed-loop management of prediction and maintenance.

Benefits of technology

It enables accurate judgment of the causes of communication network disconnection, reduces the false judgment rate, predicts and handles potential risks in advance, and ensures the continuity and stability of production activities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120979971B_ABST
    Figure CN120979971B_ABST
Patent Text Reader

Abstract

The application discloses a kind of communication network operating state monitoring method and system, belong to communication network technical field;The application includes determining supervision range and multi-source data acquisition, clear monitoring boundary, collect full-dimension original data, obtain the online network data of constructing original data, local environment data and equipment maintenance data;The application is significantly reduced by feature conversion and association rule analysis Abnormal misjudgment rate, improve maintenance resource utilization efficiency;Break through the limitation of traditional passive monitoring, with prediction analysis as core, maintenance linkage as support, risk disposal is upgraded from after-the-event response to beforehand prevention, effectively shorten abnormal response time, avoid the influence of break and connect on production efficiency and safety production;A whole-process management system of accurate monitoring-early warning-efficient disposal is constructed, the stability of intelligent equipment communication link is practically guaranteed, and reliable communication support is provided for plant digital production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication network technology, specifically to a method and system for monitoring the operational status of a communication network. Background Technology

[0002] Communication networks are the core support for intelligent equipment management in industrial plants. Key aspects such as sensor data transmission, collaborative operation of intelligent robots, and command interaction of the MES (Manufacturing Execution System) all rely on stable communication networks to achieve real-time data interaction. With the upgrading of industrial intelligence, the coverage of plant communication networks has extended from production workshops to multiple areas such as storage areas and quality inspection areas. The number of connected devices is growing exponentially. The network's operating status directly affects production efficiency and safety. If the communication between welding robots and the central control system is frequently interrupted, it may lead to welding accuracy deviations and cause product quality problems. If sensor data transmission in the storage area is interrupted, it will cause inventory counting disorder and affect supply chain scheduling.

[0003] In light of the above, it should be noted that the Chinese patent application CN2018115480404, which discloses a method and device for monitoring the operation status of a power communication network, solves the problems of low efficiency and delays in manual inspections through online communication data collection and analysis. However, it still has significant shortcomings in industrial plant scenarios: First, it ignores the linkage of local regional data, relying solely on online signal strength and traffic data to judge anomalies without combining local environmental data and equipment maintenance records, which can easily lead to errors due to the single data dimension; Second, it lacks a status pre-management mechanism, and can only passively monitor after an anomaly occurs, resulting in information gaps between supervision and maintenance control, making it impossible to replace faulty components in advance, ultimately causing disconnection accidents and affecting production continuity.

[0004] To address the aforementioned technical shortcomings, a solution is proposed. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for monitoring the operational status of communication networks to solve the problems mentioned above.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for monitoring the operating status of a communication network, comprising the following steps;

[0007] S1. Determine the scope of supervision and multi-source data collection, clarify the monitoring boundaries, collect raw data from all dimensions, and obtain online network data, local environmental data and equipment maintenance data to construct the raw data.

[0008] S2. After obtaining the raw data, data preprocessing and feature transformation are performed to solve the problems of inconsistent units, errors, and incomparability of the raw data. Finally, three types of risk characteristic indicators that can be compared horizontally are generated.

[0009] S3. Based on risk characteristic indicators and combined with the historical disconnection patterns in the plant area, formulate the correlation rules among signal, hardware and congestion, and accurately judge the abnormal attribution results of intermittent network disconnection through cross analysis;

[0010] S4 relies on the edge computing nodes of the factory area to process data quickly locally, make real-time judgments on the abnormal attribution results of S3, and generate early warning signals of corresponding levels.

[0011] S5. Based on historical data, predict future network risks and combine local daily maintenance information to verify the prediction results, thereby achieving closed-loop management of prediction and maintenance.

[0012] Furthermore, the process of acquiring the original data in step S1:

[0013] The core monitoring object is the communication link of intelligent equipment in the industrial plant area, as well as the communication status monitoring function of the intelligent equipment itself. The signal strength value reflecting the tightness of the signal connection between the equipment and the base station, the real-time traffic value reflecting the network usage intensity, and the base station load rate value reflecting the base station's carrying pressure are extracted and summarized into online network data.

[0014] By deploying environmental sensors in each sub-area, the impact of the environment on hardware equipment is extracted and summarized as temperature and humidity values; the dust content in the storage area and welding workshop is summarized as dust concentration values; and the metal raw materials temporarily piled up in the workshop and the newly added stacks of goods in the storage area are summarized as temporary obstacle record values, which are then combined to generate local environmental data.

[0015] Through the plant equipment maintenance and management system, the recent antenna replacement time and baseband chip test results are extracted and marked as historical maintenance and inspection records. The antenna VSWR and baseband chip signal demodulation rate are marked as real-time hardware parameters. The number of antenna failures and the cause of failures for the corresponding equipment are obtained and marked as hardware failure records. The equipment maintenance and inspection data are obtained, and the three sets of data are summarized into raw data.

[0016] Furthermore, the processing procedure for the original data in step S2 is as follows:

[0017] The monitoring time from the start time of data collection to the current end time of data collection is obtained and marked as the monitoring period. The monitoring period is divided into i period nodes, where i is a natural number greater than zero. Missing values ​​are processed for the collected online network data, local environment data, and equipment maintenance data in the starting order of the period nodes.

[0018] If signal strength data for a certain period node is missing, first check the average signal strength of the adjacent 10 seconds in the same sub-region, and then combine it with the local environmental records for that period. If the records show no temporary obstruction, supplement it with the average of the adjacent time periods; if there is obstruction, supplement it with the average of the adjacent time periods reduced by 10%, because obstacles will weaken the signal; if temperature and humidity data are missing, directly use the historical average temperature and humidity of the same area at the same time of the day to supplement it.

[0019] Furthermore, the process of generating risk characteristic indicators in step S2 is as follows:

[0020] The raw data of the three types of cleaned data are transformed into three risk characteristic indicators with a unified value range of 0-1 or a fixed interval. The actual values ​​of signal strength and base station load rate in online network data are extracted to find the intervals within the normal range, generating signal health and congestion risk values. Antenna VSWR and baseband chip signal demodulation rate in equipment maintenance data are also extracted. The qualified standards of the two types of hardware parameters are determined by the position of the converted values ​​within the 0-1 interval, generating hardware health values. The risk characteristic indicators are obtained by summarizing the three sets of values ​​from the previous analysis.

[0021] Furthermore, the process of generating association rules in step S3 is as follows:

[0022] We extracted 50 network disconnection events from the past few years of our factory, analyzed the specific values ​​of signal health, hardware health, and congestion risk at the time of each disconnection, summarized the three most common causes of disconnection, and referred to typical cases of network disconnection attribution in the communications industry to finally determine the following three core rules.

[0023] When the signal health value is at a low level and the hardware health value is also at a low level, the equipment in the weak signal area needs to frequently switch base stations to maintain the connection. Hardware failure will greatly increase the probability of switching failure, which will eventually lead to frequent network disconnection. This will be summarized and marked as Rule 1.

[0024] Furthermore, when the signal health value is at a high level, but the congestion risk value exceeds 1, the corresponding base station load rate value exceeds the overload threshold of 90%. Even if the signal is stable, the base station may experience processing delays or even be unable to respond due to the excessive number of devices it carries and the large amount of data transmission, ultimately leading to network disconnection. This is summarized and marked as Rule 2.

[0025] When the signal health value is at a low level, the hardware health value is at a normal level (above 0.6), there are no obvious hardware faults, and the congestion risk value is between 0.8 and 1, the weak signal itself will lead to a decrease in data transmission stability, and mild congestion will further increase the transmission delay. The combination of the two will increase the probability of disconnection by more than 25% compared to a single factor. This combination is summarized and marked as Rule 3.

[0026] Furthermore, retrieve the signal health value, hardware health value, and congestion risk value of the most recent period node in a certain sub-region, as well as the corresponding record of the number of disconnections in that region in the network connection status: count the total number of disconnections in the most recent period node of that sub-region; compare each of the three types of characteristic indicators at the time of each disconnection with the above three rules to determine which rule each disconnection conforms to; calculate the proportion of disconnections corresponding to each rule to determine the main and secondary causes of disconnections in that region, and obtain the anomaly attribution results.

[0027] Furthermore, the warning signal is generated in step S4 as follows:

[0028] By combining the importance / weight of the three types of characteristic indicators on network risk, an indicator reflecting the overall risk level is formed and marked as the comprehensive risk value. The influence of hardware health value accounts for 40%, signal health value accounts for 30%, and congestion risk value after balancing and adjustment accounts for 30%. The pre-stored risk thresholds are retrieved and compared with the comprehensive risk value to obtain low risk thresholds, medium risk thresholds, and high risk thresholds. Based on the risk level and the anomaly attribution results, corresponding early warning signals are generated.

[0029] Furthermore, in the closed-loop management generation process in step S5, the predicted data input is obtained based on the process of S1-S4, and the network risk prediction result is obtained by combining the pre-stored prediction model training and output. Based on the network risk prediction result of future period nodes and the corresponding equipment maintenance data, the maintenance plan is combined with the predicted risk through the linkage of prediction and maintenance, so as to achieve preventive maintenance and curb the occurrence of disconnection from the source.

[0030] A communication network operation status monitoring system includes the following steps:

[0031] Multi-source data acquisition module: Based on the selected data acquisition range, it is responsible for the acquisition and transmission of online network data, local environmental data, and equipment maintenance data;

[0032] The data preprocessing and analysis module cleans, transforms features, and attributes anomalies to the collected raw data, reducing processing latency and meeting real-time requirements.

[0033] The real-time judgment and prediction module analyzes the comprehensive risk value in real time and generates early warning signals. The prediction unit outputs future risk prediction results based on historical data. The two modules communicate with each other to ensure that the prediction is based on the real-time status.

[0034] The beneficial effects of this invention are:

[0035] 1. This invention constructs a multi-source data acquisition system that combines online network data, local environmental data, and equipment maintenance data. During the data preprocessing stage, it uses local environmental records to calibrate online signal data. After feature conversion, it performs multi-dimensional correlation analysis of signal health value, hardware health value, and congestion risk value to accurately attribute the cause of disconnection. Combined with antenna VSWR data from equipment maintenance, it further verifies the hardware status, significantly reducing misjudgments caused by single data dimensions, reducing ineffective maintenance resource investment, and improving the accuracy of anomaly detection.

[0036] 2. This invention uses a predictive model to predict the comprehensive network risk value for each time period in advance; at the same time, it links and compares the prediction results with local daily maintenance and inspection information, and then initiates millisecond-level preprocessing to realize closed-loop management of prediction and maintenance and inspection, curbing the risk of disconnection from the source, avoiding production interruptions caused by hardware failures, congestion and other problems, and ensuring the continuity of factory communication network and production activities. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 This is a flowchart of the method of the present invention;

[0039] Figure 2 This is a flowchart of the system of the present invention. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] Example 1: Please refer to Figure 1 - Figure 2As shown, this embodiment is a method and system for monitoring the operating status of a communication network, including the following steps:

[0042] S1. Determine the scope of supervision and collect multi-source data. The core of this step is to clarify the monitoring boundaries and collect raw data from all dimensions to provide a complete and reliable data source for subsequent analysis. All data must be labeled with the collection time and the sub-region to which they belong to ensure traceability.

[0043] With the communication links of intelligent devices within the industrial plant as the core monitoring object, the plant area is divided into three sub-areas according to production functions: production workshop, warehousing and logistics area, and central control and dispatch area. Production workshops include welding workshops and assembly workshops; warehousing and logistics areas include raw material warehouses and finished product warehouses; and central control and dispatch area includes the area around the central control room. For each sub-area, key intelligent devices are identified, such as welding robots in the workshop, AGV transport vehicles in the warehousing area, and production dispatch server in the central control room. The communication IP address and the base station number connected to each device are recorded to form a correspondence table of sub-area-intelligent device-connected base station, so as to avoid missing key links during monitoring.

[0044] Based on the core monitoring targets in the three sub-regions, online network data, local environmental data, and equipment maintenance data are acquired.

[0045] Online network data: Utilizing the built-in monitoring modules of each base station within the factory area, provided by the base station equipment manufacturer, real-time network parameter output is supported. This data is also acquired through the communication status monitoring functions of intelligent devices themselves, such as the network status recording modules built into robot control systems and AGV vehicle systems. The collected data is aggregated to obtain signal strength values ​​reflecting the tightness of the signal connection between the device and the base station; it records whether the device is normally connected or disconnected, recording the start and end times of each disconnection; it provides feedback on the amount of data transmitted between the device and the base station, reflecting real-time traffic values ​​indicating network usage intensity; and it shows the number of devices currently connected to the base station, reflecting the base station's load rate. Signal strength is used to subsequently determine whether signal coverage is sufficient; connection status is used to count disconnection frequency; and real-time traffic and base station load rate are used to determine if network congestion exists.

[0046] Local environmental data: Environmental sensors deployed in various sub-areas, such as temperature and humidity sensors in the workshop and dust concentration sensors in the storage area, are procured and installed by the plant's operations and maintenance department; daily records from on-site management personnel, such as the location of temporarily stored raw materials in the workshop and changes in the stacking height of goods in the storage area, are entered in real time via a mobile app; workshop temperature and humidity, reflecting the impact of the environment on hardware equipment, are summarized as temperature and humidity values; dust content in the storage area and welding workshop, to prevent dust from corroding hardware and causing malfunctions, is summarized as dust concentration values; temporarily stored metal raw materials in the workshop and newly added stacks of goods in the storage area, which may obstruct signals, are summarized as temporary obstacle records; used to calibrate online network data, such as determining whether a sudden signal weakening is due to temporary obstacle obstruction or hardware failure; and providing environmental evidence for subsequent analysis of hardware failure causes, such as excessive humidity potentially causing antenna corrosion.

[0047] Equipment maintenance data: This data is maintained by the operations and maintenance department through the plant's equipment maintenance management system, recording the historical maintenance records for each piece of equipment. Maintenance personnel use on-site testing tools, such as ODM hardware diagnostic instruments, to test the status of components like antennas and baseband chips. A mobile app for maintenance personnel is used to input on-site test results in real time. The previous antenna replacement time and baseband chip test results are marked as historical maintenance records. Antenna VSWR and baseband chip signal demodulation rate, reflecting the current hardware status, are marked as real-time hardware parameters. The number of antenna failures and their causes for a particular piece of equipment within the past year are marked as hardware failure records. This data is used to determine if the hardware is aging or malfunctioning, and is the core basis for subsequent analysis of hardware-related disconnections.

[0048] S2. Data preprocessing and feature transformation: The core of this step is to clean, denoise, and standardize the original data to solve the problems of inconsistent units, errors, and incomparability of the original data. Finally, three types of risk feature indicators that can be compared horizontally are generated, laying the foundation for subsequent correlation analysis.

[0049] The monitoring time from the start time of data collection to the current end time of data collection is obtained and marked as the monitoring period. The monitoring period is divided into i period nodes, where i is a natural number greater than zero. Missing values ​​are processed for the collected online network data, local environment data, and equipment maintenance data in the starting order of the period nodes.

[0050] If signal strength data for a certain period node is missing, first check the average signal strength of the adjacent 10 seconds in the same sub-region, and then combine it with the local environmental records for that period. If the records show no temporary obstacles, supplement it with the average of the adjacent periods.

[0051] If there are obstacles blocking the signal, the signal will be supplemented by reducing the average value of adjacent time periods by 10%, because obstacles will weaken the signal.

[0052] If temperature and humidity data are missing, historical average temperature and humidity data for the same sub-region at the same time of day will be used to supplement them. Historical data comes from long-term records of environmental sensors from step one. This fills in data gaps and avoids interruptions in subsequent analysis due to missing data. At the same time, it ensures that the supplemented data matches the actual scenario and reduces errors.

[0053] Select the online network data and equipment maintenance data, and extract any abnormal fluctuations in the two sets of data. For example, the signal strength suddenly drops below -110dBm, which is far beyond the normal range of the factory area, and the antenna VSWR suddenly rises to above 2.0, which is far beyond the acceptable range.

[0054] Mark abnormal signal strength, first check the local environmental data records for the same time period. If there are temporary obstacles, such as temporary stacking of metal raw materials in the workshop, mark them as temporary interference data and remove them to avoid affecting the subsequent signal stability judgment. At the same time, generate an anomaly reporting log and the cause of the anomaly.

[0055] If there is no temporary interference, the data is retained as a precursor to hardware failure, and a hardware abnormality warning signal is generated.

[0056] For abnormal hardware parameters, the pre-stored hardware qualification standards in the equipment maintenance system are retrieved for comparison. For example, the qualified range for antenna VSWR is 1.0-1.5. If the value is outside the range, it is marked as abnormal hardware data for subsequent hardware fault analysis. Meaningless interference data is removed, and data that truly reflects network problems is retained to ensure the accuracy of subsequent analysis.

[0057] The original data of the three types of cleaned data are transformed into three types of risk characteristic indicators with a unified value range of 0-1 or a fixed range, so that different types of data can be compared horizontally and feature transformation can be achieved.

[0058] Based on online network data, the normal signal strength range of the factory's communication network was first determined. Referring to the 5G signal standards for industrial plants in the communications industry, and combining this with signal records from the past year without interruptions, the normal range was determined to be -70dBm to -90dBm. Then, based on the actual signal strength values ​​within this range, they were converted to values ​​between 0 and 1.

[0059] If the actual signal strength value is closer to -70dBm, the upper limit of the normal range, the signal health value is obtained, and the closer it is to 1, the more stable the signal is.

[0060] If the actual signal strength value is closer to -90dBm, the lower limit of the normal range, the signal health value is obtained. The closer it is to 0, the weaker the signal. For example, if the actual signal strength is -85dBm, which is in the middle of the normal range, the corresponding signal health value is 0.5, which means that the signal is at a moderately stable level. The original negative signal strength value in dBm is converted into an intuitive value of 0-1, which makes it easier to compare with hardware and congestion-related data to judge the impact of the signal on network risks.

[0061] Based on equipment maintenance data, the acceptable standards for two hardware parameters—antenna VSWR and baseband chip demodulation rate—are first defined. The acceptable range for antenna VSWR is 1.0-1.5, with an upper limit of 2.0. The acceptable range for baseband chip demodulation rate is ≥95%. Considering the impact of these two parameters on hardware status, antenna VSWR, which affects signal reception more critically than demodulation rate, is converted into a value between 0 and 1. The closer the antenna VSWR is to the acceptable range, the higher the baseband chip demodulation rate, resulting in a hardware health value. A value closer to 1 indicates better hardware condition. For example, an antenna VSWR of 1.7 exceeds the acceptable range by 0.2 but does not reach the upper limit of 2.0. A baseband chip demodulation rate of 92% is slightly below the acceptable standard, resulting in a hardware health value of 0.55, indicating a minor risk of hardware failure. Integrating these dispersed hardware parameters, such as antenna VSWR and demodulation rate, into a unified health indicator facilitates subsequent assessment of the hardware's impact on network disconnection.

[0062] Based on online network data, the overload threshold of the base station is first determined. Referring to the parameters provided by the base station equipment manufacturer and combining the base station operation records during peak periods (such as the peak of production command transmission from 9:00 to 11:00), the overload threshold is determined to be 90%. Then, based on the ratio of the actual base station load rate to the overload threshold, it is converted into a value between 0 and 1.2.

[0063] If the actual base station load rate is low, the congestion risk value is obtained, and the closer it is to 0, the less pressure the base station is under.

[0064] If the actual base station load rate exceeds 90%, a congestion risk value is obtained. If the value is greater than 1, it indicates that the base station is overloaded and is prone to disconnection. For example, if the actual base station load rate is 72%, which is below the overload threshold, the corresponding congestion risk value is 0.8, indicating that the base station is in a low-load state. If the actual base station load rate is 99%, which exceeds the overload threshold, the corresponding congestion risk value is 1.1, indicating that the base station is overloaded. The load rate, expressed as a percentage, is converted into an intuitive congestion risk value to facilitate subsequent assessment of the impact of network congestion on disconnection. The obtained signal health value, hardware health value, and congestion risk value are summarized and marked as risk characteristic indicators.

[0065] S3. Multi-dimensional correlation analysis and anomaly attribution: Based on risk characteristic indicators and combined with the historical disconnection patterns in the plant area, we formulate correlation rules among signal, hardware and congestion. Through cross-analysis, we can accurately determine the specific causes of intermittent network disconnection and avoid misjudgment caused by single data.

[0066] We extracted 50 network disconnection events from the past two years in our factory, from the connection status records in the online network data of step one, and analyzed the specific values ​​of signal health, hardware health, and congestion risk at the time of each disconnection. We summarized the three most common causes of disconnection and related logics. At the same time, we referred to typical attribution cases of network disconnection in industrial plants in the communications industry and finally determined the following three core rules.

[0067] Rule 1: Weak signal combined with hardware failure exacerbates connection drops.

[0068] When the signal health value is at a low level, below 0.4, it corresponds to the actual signal strength being close to or below the lower limit of the normal range of the factory area -90dBm. At the same time, when the hardware health value is also at a low level, below 0.6, it corresponds to the antenna VSWR exceeding the standard and the baseband chip demodulation rate being low. In areas with weak signals, the equipment needs to frequently switch base stations to maintain the connection. Hardware failure will greatly increase the probability of switching failure, ultimately leading to frequent network disconnections. These are summarized and marked as Rule 1.

[0069] For example, workshop A once had a signal health value of 0.35, corresponding to a signal strength of -89dBm, close to the lower limit of the normal range, and a hardware health value of 0.55, corresponding to an antenna VSWR of 1.7 and a demodulation rate of 92%. At that time, there were 3 disconnections within 1 hour. After on-site investigation, it was confirmed that the hardware failure was caused by antenna corrosion, coupled with the signal being blocked by the metal equipment in the workshop, which perfectly matched the rule. This is used to determine the cause of the combined disconnection of weak signal and hardware failure, so as to avoid misjudging the problem as insufficient base station coverage just because the signal is weak and ignoring hardware problems.

[0070] Rule 2: If the signal is normal but the base station is overloaded, it will cause a disconnection.

[0071] When the signal health value is at a high level, above 0.6, the actual signal strength is close to or higher than the upper limit of the normal range of the factory area -70dBm. However, when the congestion risk value exceeds 1, the base station load rate exceeds the overload threshold of 90%. Even if the signal is stable, the base station may experience processing delays or even be unable to respond due to the large number of devices it supports and the large amount of data transmission, which will eventually lead to network disconnection. This will be summarized and marked as Rule 2.

[0072] For example, during peak production hours (9:00-11:00), the central control and dispatch area experienced a signal health value of 0.7, corresponding to a signal strength of -78dBm, indicating normal signal strength. However, the congestion risk value was 1.1, corresponding to a base station load rate of 99%, indicating overload. At that time, the central control system experienced two disconnections when sending instructions to multiple robots. Upon investigation, it was found that the instruction transmission traffic surged during peak hours, causing base station overload, which conforms to this rule. This distinguishes between disconnections caused by signal problems and congestion problems, preventing the risk of base station overload from being overlooked due to normal signal strength.

[0073] Rule 3: Weak signal combined with mild congestion increases the probability of connection loss.

[0074] When the signal health value is at a low level (below 0.4), the hardware health value is at a normal level (above 0.6), indicating no obvious hardware failure, and the congestion risk value is between 0.8 and 1, the corresponding base station load rate is 72%-90%, indicating mild congestion. Weak signal itself leads to decreased data transmission stability, and mild congestion further increases transmission delay. The combination of these two factors increases the probability of disconnection by more than 25% compared to a single factor. This is summarized and marked as Rule 3.

[0075] For example, in the warehouse area, a signal health value of 0.38 was observed, corresponding to a signal strength of -88dBm (weak signal). The hardware health value was 0.7, corresponding to an antenna VSWR of 1.4 and a demodulation rate of 96% (normal hardware). The congestion risk value was 0.85, corresponding to a base station load rate of 76.5% (mild congestion). Two disconnections occurred that day. After investigation, it was confirmed that the goods in the warehouse area were blocking the signal, which, combined with the data transmission from the AGV, caused mild congestion, consistent with this rule. This identifies hidden disconnection risks caused by multiple factors combined, but where no single factor is severe, avoiding the neglect of the combined risks due to a single factor not reaching the threshold.

[0076] The anomaly attribution analysis process clarifies the data source and the role of the results; retrieves the signal health value, hardware health value, and congestion risk value of the most recent period node in a certain sub-region, as well as the corresponding record of the number of disconnections in that region in the network connection status;

[0077] First, count the total number of disconnections within the most recent cycle node of this sub-region, such as 5 times;

[0078] For each disconnection, the three characteristic indicators are compared one by one with the three rules mentioned above to determine which rule each disconnection conforms to. For example, three disconnections conform to rule one, and two disconnections conform to rule three.

[0079] Calculate the percentage of disconnections corresponding to each rule, with rule one accounting for 60% and rule three accounting for 40%, to determine the main causes of disconnections in this area. Rule one: weak signal + hardware failure, and secondary causes, rule three: weak signal + mild congestion. Based on this, obtain the anomaly attribution results; clarify the core cause of disconnections in specific sub-areas, providing accurate basis for the generation of early warning signals in step four and the preprocessing plan in step five. For example, if the main cause is hardware failure, prioritize the development of a hardware replacement plan.

[0080] Example 2:

[0081] S4. Local real-time judgment and early warning signal generation: The core of this step is to rely on the edge computing node of the plant area, which is deployed in the central control room for local rapid data processing, to make real-time judgments on the abnormal attribution results of S3, and to generate early warning signals of the corresponding level, so as to ensure that maintenance personnel can quickly obtain risk information and control the response latency within 200 milliseconds.

[0082] Obtain signal health value, hardware health value, congestion risk value, and association rules to determine the influence weight of each indicator;

[0083] By combining the importance / weight of the three types of characteristic indicators on network risk, an indicator reflecting the overall risk level is formed and labeled as the comprehensive risk value.

[0084] Hardware health value has the greatest impact on network risk because hardware failure repair takes a long time and has a more lasting impact, so it has the highest weight.

[0085] The adjusted impacts of signal health and congestion risk are roughly equal, with the latter having a lower weight.

[0086] Specifically: the impact of hardware health value accounts for 40%, the impact of signal health value accounts for 30%, and the impact of congestion risk value after balancing and adjustment accounts for 30%, to avoid small fluctuations causing large changes in risk value;

[0087] The comprehensive risk value ranges from 0 to 1. The lower the value, the higher the network risk, and the higher the value, the more stable the network. The risk of the three dimensions of signal health value, hardware health value, and congestion risk value is integrated into a unified overall risk indicator, which facilitates quick judgment of the overall network status and avoids the tediousness of analyzing multiple indicators one by one.

[0088] Statistical analysis of network outages in the factory over the past two years revealed that 90% of outages occurred when the overall risk value was below 0.4. Given the factory's production needs, core production areas such as workshops have low tolerance for outages, requiring stricter thresholds. Non-core areas, such as office areas, have higher tolerance, allowing for more lenient thresholds. Referring to common standards for network risks in industrial plants within the telecommunications industry and referencing risk thresholds set for similar plants, a risk threshold for this study was determined, ranging from 0 to 1. Based on these defined criteria, risk assessment level boundaries were constructed.

[0089] Low risk threshold: A comprehensive risk value ≥ 0.7 indicates that the network signal is stable, the hardware is in good condition, the base station is not overloaded, and the operation status does not require intervention.

[0090] Medium risk threshold: A comprehensive risk value between 0.4 and 0.7 indicates potential network anomalies, such as slightly weak signal, slightly aging hardware, or mild congestion, which require continuous monitoring.

[0091] High-risk threshold: A comprehensive risk value < 0.4 indicates an extremely high probability of network anomalies, with the possibility of disconnection at any time, requiring emergency handling; as a benchmark for judging risk levels, the abstract comprehensive risk value is transformed into a specific risk level, facilitating the generation of corresponding early warnings.

[0092] The edge computing node receives the comprehensive risk value of each sub-region every 200 milliseconds, compares it with the set three-level threshold, and determines the risk level.

[0093] Based on the anomaly attribution results, identify the causes of the risk. For example, if the comprehensive risk value is 0.32 < 0.4, it indicates a high risk, and the cause is antenna failure + insufficient signal coverage.

[0094] Based on the risk level and anomaly attribution results, corresponding early warning signals are generated:

[0095] High-risk warnings should include the location of the sub-area, the core cause, and emergency response suggestions. For example, in workshop A, if there is an antenna failure and a weak signal, it is recommended to replace the antenna and temporarily deploy a micro base station within one hour. The warning signal is pushed in real time to the mobile terminal / mobile APP of maintenance personnel and the display screen of the central control system through the plant intranet to ensure that maintenance personnel can obtain risk information as soon as possible, quickly formulate response measures, and avoid the actual occurrence of disconnection or the expansion of its impact.

[0096] Medium-risk warnings should include the location of the sub-area, potential causes, and the time period of concern. For example, in the storage area, if there is mild congestion, the focus should be on the peak period from 9:00 to 11:00. The warning signal is pushed in real time to the mobile terminals / mobile APP of maintenance personnel and the display screen of the central control system through the plant intranet, so as to ensure that maintenance personnel can obtain risk information as soon as possible, quickly formulate response measures, and avoid the actual occurrence of disconnection or the expansion of its impact.

[0097] No warning will be generated if the risk is low.

[0098] S5. Predictive analysis and maintenance information linkage: The core of this step is to predict future network risks based on historical data and verify the prediction results by combining them with local daily maintenance information. If the two are consistent, pre-processing is initiated to avoid the actual occurrence of anomalies and to achieve closed-loop management of prediction and maintenance.

[0099] As pre-defined, after several cycle nodes, signal health values, hardware health values, and congestion risk data for each sub-region over the past three months are extracted to reflect historical risk characteristics; disconnection records for each sub-region over the past three months, including disconnection time, frequency, and cause, are used to verify prediction accuracy; combined with factory production plan data, such as production schedules and equipment usage for the next 7 days, from the production management system, are used to determine future traffic changes; corresponding environmental change data for several recent cycle nodes, such as weather forecasts for the next 7 days and workshop raw material stacking plans, from logistics department records, are used to determine the impact of the future environment on the signal, dynamically summarizing and generating prediction basis conditions; providing comprehensive historical and future influencing factors for the prediction model to ensure that the prediction results fit the actual scenario;

[0100] The pre-stored LSTM model is suitable for processing time series data and can capture long-term trends in historical data. The parameters are optimized by the factory's IT department in conjunction with the industrial network scenario.

[0101] The model uses the comprehensive risk value of each period node in the next 7 days as the prediction target. During training, the model parameters are continuously adjusted to ensure that the prediction accuracy is ≥85%, that is, the proportion of predicted risk levels that match the actual risk levels is ≥85%. For example, the comprehensive risk value of workshop A in the next 7 days from 9:00 to 10:00 on Thursday is predicted to be 0.32. A value <0.4 is considered high risk, caused by the continuous decline in antenna health / antenna aging + base station congestion due to the peak production period on Thursday. At the same time, the probability of disconnection during this period is predicted to be 88%, requiring early intervention. By identifying potential network risks in advance, the model provides a forward-looking basis for maintenance and inspection work, avoiding passive responses to disconnection.

[0102] Based on the network risk prediction results for future cycle nodes, equipment maintenance data is extracted, including daily maintenance records and hardware inspection reports from the factory maintenance management system. For example, the maintenance record of workshop A shows that when the antenna was inspected last week, slight corrosion was found, and the antenna VSWR increased from 1.3 to 1.4. Although it is still within the acceptable range, it has shown an upward trend, and replacement is planned for next week. Local actual hardware status information is provided to verify the authenticity of the prediction results and avoid the disconnect between prediction and reality.

[0103] Compare the risk causes in the network risk prediction results with the hardware status in the maintenance records. Risk causes could be antenna aging; hardware status could be slight antenna corrosion or increased VSWR. Determine if both point to the same problem.

[0104] If the problem is predicted to be with the pointing antenna, and the maintenance records also show that the antenna is corroded, then it is determined to be a high-matching safety hazard.

[0105] If the prediction points to base station congestion, and the maintenance records show that the base station has no hardware problems recently, but there is a production peak in the coming Thursday, then further verification should be carried out in combination with the production plan data to confirm whether the congestion is caused by peak traffic. If they match, it is also judged as a high-match security risk.

[0106] If a high-match-degree safety hazard is identified, an advance pre-treatment plan will be activated: If both the prediction and maintenance of Workshop A point to antenna aging, the antenna replacement originally scheduled for next week will be moved to Wednesday during off-peak hours, from 14:00 to 16:00. During this time, the equipment in Workshop A will be shut down for maintenance, which will not affect production.

[0107] Meanwhile, a spare antenna was deployed to workshop A to ensure a smooth replacement process; in response to the predicted congestion risk during the peak hours on Thursday, the IT department was coordinated in advance to dynamically adjust the base station frequency band, switching from 2.6GHz to 3.5GHz, increasing base station capacity by 20% and avoiding congestion; through the linkage between prediction and maintenance, the pre-processing plan was ensured to be targeted and realistic, avoiding blindly maintaining in advance; at the same time, the maintenance plan was combined with the predicted risks to achieve preventive maintenance and curb the occurrence of disconnection from the source.

[0108] Combining Embodiment 1 and Embodiment 2, this solution employs a dual design that enhances the accuracy of judgment through multi-source data linkage and achieves proactive control through a prediction-maintenance closed loop, thereby comprehensively optimizing the industrial plant's communication network monitoring system. On the one hand, compared to existing technologies that rely solely on online data for single analysis, this solution integrates multi-dimensional data from online sources, local environments, and equipment maintenance. Through feature transformation and association rule analysis, it significantly reduces the false alarm rate and improves the efficiency of maintenance resource utilization.

[0109] On the other hand, it breaks through the limitations of traditional passive monitoring, takes predictive analysis as the core and maintenance and inspection linkage as the support, upgrades risk handling from post-event response to pre-event prevention, effectively shortens the abnormal response time, and avoids the impact of disconnection on production efficiency and safe production; it has built a full-process management system of precise monitoring-early warning-efficient handling, effectively ensures the stability of intelligent equipment communication links, and provides reliable communication support for digital production in the factory area.

[0110] The above description is merely an example and illustration of the structure of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the structure of the invention or exceed the scope defined in the claims, all of which should fall within the protection scope of the present invention.

[0111] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0112] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A method of monitoring the operational state of a communications network, characterised by, Comprise the following steps; S1, determine the supervision range and multi-source data collection, clear monitoring boundary, collect full-dimensional original data, get online network data, local environment data and equipment maintenance data of original data construction; S2, after obtaining the original data, data preprocessing and feature conversion are carried out, the problems of non-uniform units, errors and inability to compare directly are solved, and finally three kinds of risk characteristic indexes which can be compared horizontally are generated; S3, based on the risk characteristic indexes, combined with the historical intermittent connection law of the plant area, the correlation rules of signal-hardware-congestion are formulated, and the abnormal attribution results of network intermittent connection are accurately judged through cross analysis; S4, relying on the edge computing node of the plant area, it is used for local rapid data processing, real-time judgment of the abnormal attribution results of S3, and generation of corresponding level warning signal; S5, based on historical data to predict future network risk, combined with local daily maintenance information to verify the prediction results, realize the closed-loop management of prediction-maintenance; The processing process of the original data in S2 step is as follows: The supervision time from the beginning of data collection to the end of current data collection is obtained, which is marked as supervision period, the supervision period is divided into i period nodes, i is a natural number greater than zero, along the starting order of the period node, the collected online network data, local environment data and equipment maintenance data are processed for missing value; If the signal strength data of a period node is missing, first check the average value of signal strength of adjacent 10 seconds in the same subarea, and then combine the local environment record of this period. If the record shows that there is no temporary obstacle blocking, the average value of the adjacent period is supplemented; if there is an obstacle, the average value of the adjacent period is reduced by 10% for supplementation, because the obstacle will weaken the signal; if the temperature and humidity data is missing, the historical average temperature and humidity of the same area at the same period is directly used for supplementation; The generation process of risk characteristic indexes in S2 step is as follows: The original data of the three kinds of cleaned data is converted into three kinds of risk characteristic indexes with unified value range, the range is unified in 0-1 or fixed interval, the position of the actual value of signal strength value and base station load rate value in online network data in normal range is extracted, signal health degree value and congestion risk degree value are generated, and signal demodulation rate of antenna standing wave ratio and baseband chip in equipment maintenance data is converted in 0-1 interval according to the qualified standard of two kinds of hardware parameters, hardware health degree value is generated, and three groups of values are collected to obtain risk characteristic indexes; The generation process of the correlation rules in S3 step is as follows: Intercept 50 network connection event records of the past few years in the factory, analyze the specific values of signal health degree value, hardware health degree value and congestion risk degree value at each connection, summarize the three most common connection causes correlation logic, and refer to the typical attribution cases of network connection in communication industry, finally determine the following 3 core rules; When the signal health value is at a low level, the value is below 0.4, and the hardware health value is also at a low level, the value is below 0.6, the signal weak area needs the device to frequently switch the base station to maintain the connection, and the hardware failure will greatly increase the probability of switching failure, eventually leading to frequent network disconnection, which is marked as rule one after being summarized; When the signal health value is at a high level, the value is above 0.6, but the congestion risk value exceeds 1, which corresponds to the overload threshold value of the base station load rate value exceeding 90%, even if the signal is stable, the base station will appear processing delay or even unable to respond due to too many devices and too large data transmission amount, eventually leading to network disconnection, which is marked as rule two after being summarized; When the signal health value is at a low level, the hardware health value is at a normal level, the value is above 0.6, the hardware has no obvious failure, and the congestion risk value is between 0.8 and 1, the weak signal itself will lead to the decline of data transmission stability, and the mild congestion will further increase the transmission delay, and the superposition of the two will increase the disconnection probability by more than 25% than single factor, which is marked as rule three after being summarized.

2. The method of claim 1, wherein, The S1 step includes the following steps: Determine the communication link of the intelligent device in the industrial plant as the core monitoring object, and the communication state monitoring function of the intelligent device itself, extract the signal strength value reflecting the close degree of signal connection between the device and the base station, the real-time traffic value reflecting the network usage intensity, and the base station load rate value reflecting the bearing pressure of the base station, and summarize the online network data; Through the environmental sensors deployed in each sub-area, the environmental influence on the hardware device is extracted and summarized as temperature and humidity values, the dust content in the warehouse area and welding workshop is summarized as dust concentration values, the temporary stacking of metal raw materials in the workshop and the newly added goods stacks in the warehouse area are summarized as temporary obstacle record values, and local environmental data is jointly generated; Through the plant equipment maintenance management system, the recent antenna replacement time and the detection result of the baseband chip are extracted, which are marked as historical maintenance records, the antenna standing wave ratio and the signal demodulation rate of the baseband chip are marked as real-time hardware parameters, the antenna fault times and fault reasons of the corresponding device are marked as hardware fault records, and the device maintenance data is obtained, and the three groups of data are summarized as original data.

3. The method of claim 1, wherein, The signal health value, hardware health value, and congestion risk value data of a sub-area in the latest cycle node are retrieved, and the disconnection times data of the corresponding records in the network connection state are recorded: the total number of disconnections in the sub-area in the latest cycle node is counted; the three types of characteristic indexes at the time of each disconnection are compared with the above three rules one by one to determine which rule each disconnection conforms to; the proportion of the disconnection times corresponding to each rule is calculated to determine the main and secondary causes of the disconnection of the region, and the abnormal attribution result is obtained.

4. The method of claim 1, wherein, The warning signal in the S4 step is generated as follows: The index reflecting the overall risk level is formed by combining the importance or weight of the influence of the three types of feature indicators on network risk, which is marked as a comprehensive risk value; the influence of the hardware health value accounts for 40%, the influence of the signal health value accounts for 30%, and the influence of the congestion risk value after balance adjustment accounts for 30%; the pre-stored risk threshold is compared and analyzed with the comprehensive risk value: low risk threshold, medium risk threshold and high risk threshold are obtained, and corresponding warning signals are generated according to the risk level and abnormal attribution result.

5. The method of claim 1, wherein, The closed-loop management generation process in the S5 step obtains prediction data input according to the processes of S1-S4, combines pre-stored prediction model training and output, obtains network risk prediction results, and according to the network risk prediction results of the future period node and the corresponding equipment maintenance data, through the linkage of prediction and maintenance, combines the maintenance plan with the prediction risk, realizes preventive maintenance, and controls the disconnection from the source.

6. A communication network operation state monitoring system for the communication network operation state monitoring method according to any one of claims 1 to 5, characterized by The method comprises the following steps: A multi-source data acquisition module is responsible for the acquisition and transmission of online network data, local environment data and equipment maintenance data based on the data acquisition range of the frame selection; A data preprocessing and analysis module cleans, converts features and attributes, and processes time delay of the collected raw data to meet real-time requirements; A real-time decision and prediction module analyzes the comprehensive risk value in real time and generates a warning signal, and a prediction unit outputs future risk prediction results based on historical data, and the two data are interconnected to ensure that the prediction is based on real-time status.

Citation Information

Patent Citations

  • High-risk environment operation emergency response method and system based on intelligent communication equipment

    CN120583451A

  • Fault monitoring in a communications network

    GB202011876D0