Communication network operation state monitoring method and system

By collecting and preprocessing multi-source data and combining it with historical patterns in industrial plants, we have achieved accurate monitoring and prediction of communication networks. This has solved the problems of large errors and inability to predict faults in advance in existing technologies, and ensured the stability of communication networks and production continuity in industrial plants.

CN120979971AActive Publication Date: 2025-11-18SHANGHAI TECH NETWORK COMM CO LTD

Patent Information

Application Number
CN202511493416.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-18
Estimated Expiration
2045-10-20

AI Technical Summary

Technical Problem

Existing technologies for monitoring communication networks in industrial plants fail to effectively combine local environmental data with equipment maintenance records, resulting in large errors, an inability to predict faults in advance, and disruption to production continuity.

Method used

By collecting data from multiple sources, including online network data, local environmental data, and equipment maintenance data, data preprocessing and feature transformation are performed. Combined with historical disconnection patterns in the plant area, correlation rules between signals, hardware, and congestion are formulated. Edge computing is used for real-time anomaly attribution and early warning. Combined with predictive models, closed-loop management of prediction and maintenance is achieved.

Benefits of technology

It enables accurate attribution of the causes of communication network disconnection, reduces the false alarm rate, predicts risks in advance, avoids production interruptions, and ensures the stability of the communication network and the continuity of production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120979971A_ABST
    Figure CN120979971A_ABST
Patent Text Reader

Abstract

The invention discloses a communication network operation state monitoring method and system, and belongs to the technical field of communication networks. The method comprises the steps of determining a supervision range and multi-source data acquisition, determining a monitoring boundary, collecting full-dimension original data, and obtaining online network data, local environment data and equipment maintenance data for constructing the original data; according to the method, through feature conversion and association rule analysis, the abnormal misjudgment rate is remarkably reduced, and the maintenance and inspection resource utilization efficiency is improved; the limitation of traditional passive monitoring is broken through, prediction analysis is taken as a core, maintenance and inspection linkage is taken as a support, risk disposal is upgraded from post-event response to pre-event prevention, the abnormal response time is effectively shortened, and the influence of disconnection on production efficiency and safety production is avoided; a full-process management system of accurate monitoring, early warning in advance and efficient disposal is constructed, the stability of a communication link of the intelligent equipment is practically guaranteed, and reliable communication support is provided for digital production of a factory area.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of communication network, in particular to a communication network operation state monitoring method and system. BACKGROUND

[0002] The communication network is the core support of the intelligent equipment management of the industrial plant. The key links such as sensor data transmission, intelligent robot collaborative work and MES production execution system instruction interaction in the plant all rely on stable communication network to realize real-time data interaction. With the upgrading of industrial intelligence, the communication network coverage of the plant extends from the production workshop to the warehouse area, quality inspection area and other areas. The number of equipment connections grows exponentially. The network operation state directly affects the production efficiency and safety production. If the communication between the welding robot and the central control system is frequently disconnected, it may cause welding precision deviation and lead to product quality problems. If the sensor data transmission in the warehouse area is interrupted, it will cause inventory counting disorder and affect the supply chain scheduling.

[0003] In combination with the above content, it needs to be explained that the Chinese patent with the application number CN2018115480404 discloses a power communication network operation state monitoring method and device. It solves the problems of low efficiency and time delay of manual inspection through online communication data collection and analysis. However, there are still significant defects in the industrial plant scenario. Firstly, it ignores regional local data linkage and only relies on online signal strength, traffic and other data to judge abnormalities. It does not combine local environmental data and equipment maintenance records in the plant, which is easy to cause errors due to single data dimension. Secondly, it lacks a state pre-management mechanism and can only passively monitor after an abnormality occurs, which leads to information gap in supervision and maintenance control, cannot replace faulty components in advance, and ultimately causes disconnection accidents and affects production continuity.

[0004] In view of the above technical defects, a solution is proposed. SUMMARY

[0005] The purpose of the present application is to provide a communication network operation state monitoring method and system to solve the problems.

[0006] To achieve the above purpose, the present application provides the following technical scheme: a communication network operation state monitoring method, comprising the following steps:

[0007] S1, determine the supervision range and multi-source data collection, clearly define the monitoring boundary, collect full-dimension original data, and obtain online network data, local environmental data and equipment maintenance data for constructing original data;

[0008] S2, after obtaining the original data, data preprocessing and feature conversion are performed to solve the problems of non-uniform units, errors and inability to directly compare the original data, and finally generate 3 types of risk feature indexes that can be compared horizontally;

[0009] S3, based on the risk feature index, combined with the historical disconnection law of the factory area, the correlation rules of signal-hardware-congestion are formulated, and the abnormal attribution results of intermittent disconnection between networks are accurately judged through cross analysis;

[0010] S4, relying on the edge computing nodes of the factory area, for local rapid data processing, real-time judgment of the abnormal attribution results of S3, and generation of corresponding level warning signals;

[0011] S5, based on historical data to predict future network risks, combined with local daily maintenance information to verify the prediction results, to realize the closed-loop management of prediction-maintenance.

[0012] Further, the S1 step of obtaining the original data process is as follows:

[0013] Determine the communication link of the intelligent equipment in the industrial plant as the core monitoring object, and the communication state monitoring function of the intelligent equipment itself, extract the signal strength value reflecting the tightness of the signal connection between the equipment and the base station, the real-time traffic value reflecting the network usage intensity, and the base station load rate value reflecting the bearing pressure of the base station, and summarize the online network data;

[0014] Through the environmental sensors deployed in each sub-region, the temperature and humidity values reflecting the influence of the environment on the hardware equipment are extracted, the dust content in the warehouse area and welding workshop is summarized as dust concentration value, the temporary stacking of metal raw materials in the workshop and the newly added goods stacks in the warehouse area are summarized as temporary obstacle record value, and the local environment data is generated in combination;

[0015] Through the equipment maintenance management system of the factory area, the recent antenna replacement time and the detection results of the baseband chip are extracted, marked as historical maintenance records, the antenna standing wave ratio and the signal demodulation rate of the baseband chip are marked as real-time hardware parameters, the antenna fault times and fault reasons of the corresponding equipment are marked as hardware fault records, and the equipment maintenance data is obtained. Summarize the three groups of data as original data.

[0016] Further, the S2 step of processing the original data is as follows:

[0017] Get the supervision time from the start time of data collection to the end time of current data collection, mark it as the supervision period, divide the supervision period into i period nodes, i is a natural number greater than zero, along the starting order of the period nodes, the collected online network data, local environment data and equipment maintenance data are processed for missing values;

[0018] If the signal strength data of a certain period node is missing, first check the average signal strength of the adjacent 10 seconds in the same sub-area, and then combine the local environment record of that period. If the record shows no temporary obstruction, it will be supplemented by the average value of the adjacent period. If there is an obstruction, it will be supplemented by the average value of the adjacent period minus 10%, because the obstruction will weaken the signal. If the temperature and humidity data is missing, the historical average temperature and humidity of the same area on the same day will be directly used to supplement it.

[0019] Further, the generation process of the risk feature index in the S2 step is as follows:

[0020] The original data of the three types of cleaned data is converted into risk feature indexes with uniform value ranges of 0-1 or fixed intervals. The actual value of the signal strength value and the base station load rate value in the online network data is located in the normal range interval to generate signal health degree values and congestion risk degree values. The signal demodulation rate of the antenna standing wave ratio and the baseband chip in the equipment maintenance data is located in the 0-1 interval generated by the qualified standard of the two types of hardware parameters to generate hardware health degree values. The three groups of values of the last time are summarized to obtain the risk feature index.

[0021] Further, the generation process of the association rule in the S3 step is as follows:

[0022] The 50 network disconnection event records of the factory in the past few years are intercepted, the specific values of the signal health degree value, the hardware health degree value, and the congestion risk degree value at each disconnection are analyzed, and the three most common disconnection cause association logics are summarized. At the same time, reference is made to the typical attribution cases of industrial factory network disconnection in the communication industry, and finally the following 3 core rules are determined.

[0023] When the signal health degree value is at a low level and the hardware health degree value is also at a low level, the signal weak area needs to frequently switch base stations to maintain connection, and hardware failure will greatly increase the probability of switching failure, ultimately leading to frequent network disconnection. It is summarized and marked as rule one.

[0024] Further, when the signal health degree value is at a high level, but the congestion risk degree value exceeds 1, the corresponding base station load rate value exceeds the overload threshold of 90%. Even if the signal is stable, the base station will also appear processing delay or even unable to respond due to too many devices and too large data transmission quantity, ultimately leading to network disconnection. It is summarized and marked as rule two.

[0025] When the signal health value is at a low level, the hardware health value is at a normal level, the value is above 0.6, there is no obvious fault in the hardware, and the congestion risk value is between 0.8 and 1, the weak signal itself will cause the data transmission stability to decrease, and the slight congestion will further increase the transmission delay, and the superposition of the two will increase the disconnection probability by more than 25% than single factor, and the summary is marked as rule three.

[0026] Further, the signal health value, hardware health value, congestion risk value data of the last cycle node of a sub-region, and the disconnection frequency data of the corresponding record in the network connection state of the region are called: the total number of disconnections in the last cycle node of the sub-region is counted; for each disconnection, the three types of feature indicators are compared one by one with the above three rules to determine which rule each disconnection conforms to; the proportion of the number of disconnections corresponding to each rule is calculated to determine the main and secondary causes of the disconnection of the region, and the abnormal attribution result is obtained.

[0027] Further, the warning signal in the S4 step is generated as follows:

[0028] Combined with the importance / weight of the influence of the three types of feature indicators on network risk, the index reflecting the overall risk level is formed, which is marked as the comprehensive risk value; the influence of the hardware health value accounts for 40%, the influence of the signal health value accounts for 30%, and the influence of the congestion risk value after balance adjustment accounts for 30%; the pre-stored risk threshold is called and compared with the comprehensive risk value: the low risk threshold, the medium risk threshold and the high risk threshold are obtained, and the corresponding warning signal is generated according to the risk level and the abnormal attribution result.

[0029] Further, the closed-loop management generation process in the S5 step inputs the prediction data obtained from the process of S1-S4, combines the pre-stored prediction model training and output, obtains the network risk prediction result, and according to the network risk prediction result of the future cycle node and the corresponding equipment maintenance data, through the linkage of prediction and maintenance, combines the maintenance plan with the prediction risk to realize preventive maintenance and prevent disconnection from the source.

[0030] A communication network operation state monitoring system, comprising the following steps:

[0031] Multi-source data acquisition module: based on the data acquisition range of the frame selection, responsible for the collection and transmission of online network data, local environment data and equipment maintenance data;

[0032] Data preprocessing and analysis module, which cleans, converts features and attributes of the collected raw data, processes the time delay, and meets the real-time requirement;

[0033] Real-time judgment and prediction module, real-time analysis of comprehensive risk value and generates early warning signal, prediction unit based on historical data output future risk prediction results, both data intercommunication, ensure that the prediction is based on real-time state.

[0034] The beneficial effects of the present application are:

[0035] 1、The present application is through constructing online network data + local environment data + equipment maintenance data multi-source data acquisition system, in the data preprocessing stage, the online signal data is calibrated by using the local environment record, after the feature conversion, the multi-dimensional correlation analysis of signal health degree value + hardware health degree value + congestion risk degree value is carried out, the accurate attribution of the disconnection cause is realized, the antenna standing wave ratio data of equipment maintenance is combined to further verify the hardware state, the misjudgment caused by single data dimension is greatly reduced, the invalid maintenance resource investment is reduced, and the accuracy of abnormal judgment is improved.

[0036] 2、The present application is through prediction model, early prediction of network comprehensive risk value in each period;At the same time, the prediction result is compared with the local daily maintenance information, then the millisecond level pre-processing is started, the prediction-maintenance closed loop management is realized, the disconnection risk is suppressed from the source, the production interruption caused by hardware failure, congestion and other problems is avoided, and the continuity of plant communication network and production activity is ensured. BRIEF DESCRIPTION OF DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0038] Figure 1 The method flow chart of the present application is shown in the figure;

[0039] Figure 2 The system flow chart of the present application is shown in the figure. DETAILED DESCRIPTION

[0040] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0041] Embodiment one: please refer to Figure 1 - Figure 2As shown, the embodiment is a communication network operation state monitoring method and system, comprising the following steps:

[0042] S1, determine the monitoring range and multi-source data collection, the core of this step is to clarify the monitoring boundary, collect full-dimensional original data, provide complete and reliable data source for subsequent analysis, all data need to be marked with collection time and belonging sub-area, ensure traceability;

[0043] Taking the communication link of intelligent equipment in the industrial plant as the core monitoring object, the plant is divided into three sub-regions according to production function: production workshop, warehouse logistics area, central control scheduling area; Production workshop such as welding workshop, assembly workshop; Warehouse logistics area such as raw material warehouse, finished product storage area; Central control scheduling area such as central control room; For each sub-region, comb the key intelligent equipment in it, such as welding robot in workshop, AGV carrying trolley in warehouse area, production scheduling server in central control room, record the communication IP address of each device, the connected base station number, form the corresponding relationship table of sub-area-intelligent equipment-connection base station, avoid missing key link during monitoring;

[0044] According to the core monitoring object of the three sub-regions, online network data, local environment data and equipment maintenance data are obtained,

[0045] Online network data: through the built-in monitoring module of each base station in the plant, provided by the base station equipment manufacturer, supporting real-time output of network parameters; And the communication state monitoring function of intelligent equipment itself, such as the network state recording module built in the robot control system and AGV vehicle system, joint acquisition, data collection, get the signal strength value reflecting the tightness of the signal connection between the equipment and the base station; Record the network connection state of the equipment, whether it is normal connection or disconnection, record the start and end time of each disconnection; Feedback the data volume between the equipment and the base station, reflect the real-time flow value of network use intensity; The number of devices currently connected to the base station, reflect the base station load rate value of base station bearing pressure; Among them, signal strength is used for subsequent judgment of whether the signal coverage is sufficient; Connection state is used to count the disconnection frequency; Real-time flow and base station load rate are used to judge whether there is network congestion.

[0046] Local environment data: through the deployment of environmental sensors in each sub-area, such as temperature and humidity sensors in the workshop, dust concentration sensors in the warehouse area, purchased and installed by the factory operation and maintenance department; daily records of on-site management personnel, such as the location of temporary stacking of raw materials in the workshop, the change of stacking height of goods in the warehouse area, which are input in real time through the management personnel's mobile APP; the temperature and humidity of the workshop, reflecting the influence of the environment on the hardware equipment, are summarized as temperature and humidity values; the dust content in the warehouse area and the welding workshop is avoided to erode the hardware and cause failure, which is summarized as dust concentration value; the temporary stacking of metal raw materials in the workshop and the newly added goods piles in the warehouse area may block the signal, which is summarized as temporary obstacle record value; used for calibrating online network data, such as sudden signal weakening, to judge whether it is temporary obstacle shielding or hardware failure; at the same time, it provides environmental basis for subsequent analysis of hardware failure reasons, such as humidity exceeding standard may cause antenna corrosion;

[0047] Equipment maintenance data: through the factory equipment maintenance management system, maintained by the operation and maintenance department, recording the historical maintenance records of each equipment; on-site detection tools of maintenance personnel, such as ODM hardware diagnostic instrument, used for detecting the status of antenna, baseband chip and other components; maintenance personnel's mobile APP, used for real-time input of on-site detection results; the time of the previous antenna replacement and the detection results of the baseband chip are marked as historical maintenance records; the antenna standing wave ratio and the signal demodulation rate of the baseband chip, which reflect the current state of the hardware, are marked as real-time hardware parameters; the number of antenna failures and the failure causes of a certain equipment in the past 1 year are marked as hardware failure records; used to judge whether the hardware equipment exists aging or failure, which is the core basis for subsequent analysis of hardware causing disconnection.

[0048] S2, data preprocessing and feature conversion, the core of this step is to clean and denoise the original data and unify the processing, solve the problem of non-uniform units, error and cannot be directly compared, finally generate 3 kinds of risk feature indexes which can be compared horizontally, lay the foundation for subsequent correlation analysis;

[0049] Get the supervision time from the start time of data collection to the current data collection cutoff time, mark it as the supervision period, divide the supervision period into i period nodes, i is a natural number greater than zero, along the starting order of the period nodes, process the missing values of the collected online network data, local environment data and equipment maintenance data;

[0050] If the signal strength data of a certain period node is missing, first check the average value of the signal strength of the adjacent 10 seconds in the same sub-area, and then combine the local environment record of this period, if the record shows no temporary obstacle shielding, it is supplemented by the average value of the adjacent period;

[0051] If there is an obstacle, it is supplemented by reducing the average value of the adjacent period by 10% because the obstacle will weaken the signal;

[0052] If the temperature and humidity data is missing, directly use the historical average temperature and humidity of the same sub-region on the same day and at the same time period to supplement. The historical data comes from the long-term record of the environmental sensor in step one. Fill in the data gap to avoid interruption in subsequent analysis due to data missing. At the same time, ensure that the supplemented data conforms to the actual scene and reduce errors.

[0053] Frame the online network data and equipment maintenance data, extract the abnormal fluctuations existing in the two groups of data, such as signal strength suddenly falling below-110dBm, far exceeding the normal range of the factory area, and antenna standing wave ratio suddenly rising above 2.0, far exceeding the qualified range;

[0054] Label abnormal signal strength, first check the local environmental data record at the same time period, if there is temporary obstacles, such as temporary stacking of metal raw materials in the workshop, mark it as temporary interference data and eliminate it to avoid affecting the subsequent signal stability judgment, and generate an abnormality log and the cause of the abnormality;

[0055] If there is no temporary interference, keep the data as a precursor to hardware failure and generate a hardware abnormality warning signal;

[0056] For abnormal hardware parameters, retrieve the pre-stored hardware qualification standards in the equipment maintenance system, such as the antenna standing wave ratio qualified range which can take values of 1.0-1.5, and mark it as hardware abnormal data if it exceeds the range, which is used for subsequent hardware failure analysis; eliminate meaningless interference data and keep the data that truly reflects network problems to ensure the accuracy of subsequent analysis.

[0057] Convert the original data of the three types of cleaned data into 3 types of risk feature indicators with uniform value range, and unify the range to 0-1 or a fixed interval, so that different types of data can be compared horizontally, and feature conversion is realized;

[0058] Based on the online network data, first determine the normal signal strength range of the factory communication network, refer to the 5G signal standard of the communication industry in industrial factories, and combine the past 1 year of signal records without interruption to determine the normal range as-70dBm to-90dBm; then according to the position of the actual signal strength value in this range, it is converted into a value between 0 and 1:

[0059] If the actual signal strength value is closer to-70dBm, the upper limit of the normal range, the signal health value is obtained, and the closer to 1, the more stable the signal is;

[0060] The closer the actual signal strength value is to the lower limit of the normal range, -90dBm, the closer the signal health value is to 0, and the weaker the signal is. For example, if the actual signal strength is -85dBm, which is in the middle of the normal range, the corresponding signal health value is 0.5, which represents a medium stable level of the signal. The negative signal strength value in dBm is converted into a direct value between 0 and 1, which is convenient for subsequent comparison with hardware and congestion-related data to determine the impact of the signal on network risk.

[0061] Based on equipment maintenance data, the qualified standards of two types of hardware parameters, antenna standing wave ratio and baseband chip signal demodulation rate, are first determined. The antenna standing wave ratio qualified range is 1.0-1.5, and the upper limit of the over-standard is 2.0. The baseband chip demodulation rate qualified range is ≥95%. Then, the influence of these two parameters on the hardware state is combined. The antenna standing wave ratio has a greater impact on signal reception than the demodulation rate. The higher the baseband chip demodulation rate, the closer the antenna standing wave ratio is to the qualified range, and the closer the hardware health value is to 1, which represents a better hardware state. For example, if the antenna standing wave ratio is 1.7, which is 0.2 higher than the qualified range but does not reach the upper limit of 2.0, and the baseband chip demodulation rate is 92%, which is slightly lower than the qualified standard, the corresponding hardware health value is 0.55, which represents a mild fault risk of the hardware. The scattered hardware parameters such as antenna standing wave ratio and demodulation rate are integrated into a unified health index, which is convenient for subsequent determination of the impact of hardware on network disconnection.

[0062] Based on online network data, the overload threshold of the base station is first determined. Referring to the parameters provided by the base station equipment manufacturer and combining the base station operation records during the peak period (such as 9:00-11:00 production instruction transmission peak), the overload threshold is determined to be 90%. Then, according to the ratio of the actual base station load rate value to the overload threshold, a value between 0 and 1.2 is converted.

[0063] If the actual base station load rate value is low, the congestion risk value is obtained, and the closer it is to 0, the smaller the pressure on the base station is.

[0064] If the actual base station load rate value exceeds 90%, the congestion risk value is obtained, and the closer it is to 1, the more overloaded the base station is, which is prone to disconnection. For example, if the actual base station load rate value is 72%, which is lower than the overload threshold, the corresponding congestion risk value is 0.8, which represents a low load state of the base station. If the actual base station load rate value is 99%, which exceeds the overload threshold, the corresponding congestion risk value is 1.1, which represents an overloaded base station. The load rate in % is converted into a direct congestion risk value, which is convenient for subsequent determination of the impact of network congestion on disconnection. The signal health value, hardware health value, and congestion risk value obtained are summarized as risk feature indicators.

[0065] S3. Multi-dimensional correlation analysis and anomaly attribution: Based on risk characteristic indicators and combined with the historical disconnection patterns in the plant area, we formulate correlation rules among signal, hardware and congestion. Through cross-analysis, we can accurately determine the specific causes of intermittent network disconnection and avoid misjudgment caused by single data.

[0066] We extracted 50 network disconnection events from the past two years in our factory, from the connection status records in the online network data of step one, and analyzed the specific values ​​of signal health, hardware health, and congestion risk at the time of each disconnection. We summarized the three most common causes of disconnection and related logics. At the same time, we referred to typical attribution cases of network disconnection in industrial plants in the communications industry and finally determined the following three core rules.

[0067] Rule 1: Weak signal combined with hardware failure exacerbates connection drops.

[0068] When the signal health value is at a low level, below 0.4, it corresponds to the actual signal strength being close to or below the lower limit of the normal range of the factory area -90dBm. At the same time, when the hardware health value is also at a low level, below 0.6, it corresponds to the antenna VSWR exceeding the standard and the baseband chip demodulation rate being low. In areas with weak signals, the equipment needs to frequently switch base stations to maintain the connection. Hardware failure will greatly increase the probability of switching failure, ultimately leading to frequent network disconnections. These are summarized and marked as Rule 1.

[0069] For example, workshop A once had a signal health value of 0.35, corresponding to a signal strength of -89dBm, close to the lower limit of the normal range, and a hardware health value of 0.55, corresponding to an antenna VSWR of 1.7 and a demodulation rate of 92%. At that time, there were 3 disconnections within 1 hour. After on-site investigation, it was confirmed that the hardware failure was caused by antenna corrosion, coupled with the signal being blocked by the metal equipment in the workshop, which perfectly matched the rule. This is used to determine the cause of the combined disconnection of weak signal and hardware failure, so as to avoid misjudging the problem as insufficient base station coverage just because the signal is weak and ignoring hardware problems.

[0070] Rule 2: If the signal is normal but the base station is overloaded, it will cause a disconnection.

[0071] When the signal health value is at a high level, above 0.6, the actual signal strength is close to or higher than the upper limit of the normal range of the factory area -70dBm. However, when the congestion risk value exceeds 1, the base station load rate exceeds the overload threshold of 90%. Even if the signal is stable, the base station may experience processing delays or even be unable to respond due to the large number of devices it supports and the large amount of data transmission, which will eventually lead to network disconnection. This will be summarized and marked as Rule 2.

[0072] For example, during peak production hours (9:00-11:00), the central control and dispatch area experienced a signal health value of 0.7, corresponding to a signal strength of -78dBm, indicating normal signal strength. However, the congestion risk value was 1.1, corresponding to a base station load rate of 99%, indicating overload. At that time, the central control system experienced two disconnections when sending instructions to multiple robots. Upon investigation, it was found that the instruction transmission traffic surged during peak hours, causing base station overload, which conforms to this rule. This distinguishes between disconnections caused by signal problems and congestion problems, preventing the risk of base station overload from being overlooked due to normal signal strength.

[0073] Rule 3: Weak signal combined with mild congestion increases the probability of connection loss.

[0074] When the signal health value is at a low level (below 0.4), the hardware health value is at a normal level (above 0.6), indicating no obvious hardware failure, and the congestion risk value is between 0.8 and 1, the corresponding base station load rate is 72%-90%, indicating mild congestion. Weak signal itself leads to decreased data transmission stability, and mild congestion further increases transmission delay. The combination of these two factors increases the probability of disconnection by more than 25% compared to a single factor. This is summarized and marked as Rule 3.

[0075] For example, in the warehouse area, a signal health value of 0.38 was observed, corresponding to a signal strength of -88dBm (weak signal). The hardware health value was 0.7, corresponding to an antenna VSWR of 1.4 and a demodulation rate of 96% (normal hardware). The congestion risk value was 0.85, corresponding to a base station load rate of 76.5% (mild congestion). Two disconnections occurred that day. After investigation, it was confirmed that the goods in the warehouse area were blocking the signal, which, combined with the data transmission from the AGV, caused mild congestion, consistent with this rule. This identifies hidden disconnection risks caused by multiple factors combined, but where no single factor is severe, avoiding the neglect of the combined risks due to a single factor not reaching the threshold.

[0076] The anomaly attribution analysis process clarifies the data source and the role of the results; retrieves the signal health value, hardware health value, and congestion risk value of the most recent period node in a certain sub-region, as well as the corresponding record of the number of disconnections in that region in the network connection status;

[0077] First, count the total number of disconnections within the most recent cycle node of this sub-region, such as 5 times;

[0078] For each disconnection, the three characteristic indicators are compared one by one with the three rules mentioned above to determine which rule each disconnection conforms to. For example, three disconnections conform to rule one, and two disconnections conform to rule three.

[0079] Calculate the percentage of disconnections corresponding to each rule, with rule one accounting for 60% and rule three accounting for 40%, to determine the main causes of disconnections in this area. Rule one: weak signal + hardware failure, and secondary causes, rule three: weak signal + mild congestion. Based on this, obtain the anomaly attribution results; clarify the core cause of disconnections in specific sub-areas, providing accurate basis for the generation of early warning signals in step four and the preprocessing plan in step five. For example, if the main cause is hardware failure, prioritize the development of a hardware replacement plan.

[0080] Example 2:

[0081] S4. Local real-time judgment and early warning signal generation: The core of this step is to rely on the edge computing node of the plant area, which is deployed in the central control room for local rapid data processing, to make real-time judgments on the abnormal attribution results of S3, and to generate early warning signals of the corresponding level, so as to ensure that maintenance personnel can quickly obtain risk information and control the response latency within 200 milliseconds.

[0082] Obtain signal health value, hardware health value, congestion risk value, and association rules to determine the influence weight of each indicator;

[0083] By combining the importance / weight of the three types of characteristic indicators on network risk, an indicator reflecting the overall risk level is formed and labeled as the comprehensive risk value.

[0084] Hardware health value has the greatest impact on network risk because hardware failure repair takes a long time and has a more lasting impact, so it has the highest weight.

[0085] The adjusted impacts of signal health and congestion risk are roughly equal, with the latter having a lower weight.

[0086] Specifically: the impact of hardware health value accounts for 40%, the impact of signal health value accounts for 30%, and the impact of congestion risk value after balancing and adjustment accounts for 30%, to avoid small fluctuations causing large changes in risk value;

[0087] The comprehensive risk value ranges from 0 to 1. The lower the value, the higher the network risk, and the higher the value, the more stable the network. The risk of the three dimensions of signal health value, hardware health value, and congestion risk value is integrated into a unified overall risk indicator, which facilitates quick judgment of the overall network status and avoids the tediousness of analyzing multiple indicators one by one.

[0088] Statistical analysis of network outages in the factory over the past two years revealed that 90% of outages occurred when the overall risk value was below 0.4. Given the factory's production needs, core production areas such as workshops have low tolerance for outages, requiring stricter thresholds. Non-core areas, such as office areas, have higher tolerance, allowing for more lenient thresholds. Referring to common standards for network risks in industrial plants within the telecommunications industry and referencing risk thresholds set for similar plants, a risk threshold for this study was determined, ranging from 0 to 1. Based on these defined criteria, risk assessment level boundaries were constructed.

[0089] Low risk threshold: A comprehensive risk value ≥ 0.7 indicates that the network signal is stable, the hardware is in good condition, the base station is not overloaded, and the operation status does not require intervention.

[0090] Medium risk threshold: A comprehensive risk value between 0.4 and 0.7 indicates potential network anomalies, such as slightly weak signal, slightly aging hardware, or mild congestion, which require continuous monitoring.

[0091] High-risk threshold: A comprehensive risk value < 0.4 indicates an extremely high probability of network anomalies, with the possibility of disconnection at any time, requiring emergency handling; as a benchmark for judging risk levels, the abstract comprehensive risk value is transformed into a specific risk level, facilitating the generation of corresponding early warnings.

[0092] The edge computing node receives the comprehensive risk value of each sub-region every 200 milliseconds, compares it with the set three-level threshold, and determines the risk level.

[0093] Based on the anomaly attribution results, identify the causes of the risk. For example, if the comprehensive risk value is 0.32 < 0.4, it indicates a high risk, and the cause is antenna failure + insufficient signal coverage.

[0094] Based on the risk level and anomaly attribution results, corresponding early warning signals are generated:

[0095] High-risk warnings should include the location of the sub-area, the core cause, and emergency response suggestions. For example, in workshop A, if there is an antenna failure and a weak signal, it is recommended to replace the antenna and temporarily deploy a micro base station within one hour. The warning signal is pushed in real time to the mobile terminal / mobile APP of maintenance personnel and the display screen of the central control system through the plant intranet to ensure that maintenance personnel can obtain risk information as soon as possible, quickly formulate response measures, and avoid the actual occurrence of disconnection or the expansion of its impact.

[0096] Medium-risk warnings should include the location of the sub-area, potential causes, and the time period of concern. For example, in the storage area, if there is mild congestion, the focus should be on the peak period from 9:00 to 11:00. The warning signal is pushed in real time to the mobile terminals / mobile APP of maintenance personnel and the display screen of the central control system through the plant intranet, so as to ensure that maintenance personnel can obtain risk information as soon as possible, quickly formulate response measures, and avoid the actual occurrence of disconnection or the expansion of its impact.

[0097] No warning will be generated if the risk is low.

[0098] S5. Predictive analysis and maintenance information linkage: The core of this step is to predict future network risks based on historical data and verify the prediction results by combining them with local daily maintenance information. If the two are consistent, pre-processing is initiated to avoid the actual occurrence of anomalies and to achieve closed-loop management of prediction and maintenance.

[0099] As pre-defined, after several cycle nodes, signal health values, hardware health values, and congestion risk data for each sub-region over the past three months are extracted to reflect historical risk characteristics; disconnection records for each sub-region over the past three months, including disconnection time, frequency, and cause, are used to verify prediction accuracy; combined with factory production plan data, such as production schedules and equipment usage for the next 7 days, from the production management system, are used to determine future traffic changes; corresponding environmental change data for several recent cycle nodes, such as weather forecasts for the next 7 days and workshop raw material stacking plans, from logistics department records, are used to determine the impact of the future environment on the signal, dynamically summarizing and generating prediction basis conditions; providing comprehensive historical and future influencing factors for the prediction model to ensure that the prediction results fit the actual scenario;

[0100] The pre-stored LSTM model is suitable for processing time series data and can capture long-term trends in historical data. The parameters are optimized by the factory's IT department in conjunction with the industrial network scenario.

[0101] The model uses the comprehensive risk value of each period node in the next 7 days as the prediction target. During training, the model parameters are continuously adjusted to ensure that the prediction accuracy is ≥85%, that is, the proportion of predicted risk levels that match the actual risk levels is ≥85%. For example, the comprehensive risk value of workshop A in the next 7 days from 9:00 to 10:00 on Thursday is predicted to be 0.32. A value <0.4 is considered high risk, caused by the continuous decline in antenna health / antenna aging + base station congestion due to the peak production period on Thursday. At the same time, the probability of disconnection during this period is predicted to be 88%, requiring early intervention. By identifying potential network risks in advance, the model provides a forward-looking basis for maintenance and inspection work, avoiding passive responses to disconnection.

[0102] Based on the network risk prediction results for future cycle nodes, equipment maintenance data is extracted, including daily maintenance records and hardware inspection reports from the factory maintenance management system. For example, the maintenance record of workshop A shows that when the antenna was inspected last week, slight corrosion was found, and the antenna VSWR increased from 1.3 to 1.4. Although it is still within the acceptable range, it has shown an upward trend, and replacement is planned for next week. Local actual hardware status information is provided to verify the authenticity of the prediction results and avoid the disconnect between prediction and reality.

[0103] Compare the risk causes in the network risk prediction results with the hardware status in the maintenance records. Risk causes could be antenna aging; hardware status could be slight antenna corrosion or increased VSWR. Determine if both point to the same problem.

[0104] If the problem is predicted to be with the pointing antenna, and the maintenance records also show that the antenna is corroded, then it is determined to be a high-matching safety hazard.

[0105] If the prediction points to base station congestion, and the maintenance records show that the base station has no hardware problems recently, but there is a production peak in the coming Thursday, then further verification should be carried out in combination with the production plan data to confirm whether the congestion is caused by peak traffic. If they match, it is also judged as a high-match security risk.

[0106] If a high-match-degree safety hazard is identified, an advance pre-treatment plan will be activated: If both the prediction and maintenance of Workshop A point to antenna aging, the antenna replacement originally scheduled for next week will be moved to Wednesday during off-peak hours, from 14:00 to 16:00. During this time, the equipment in Workshop A will be shut down for maintenance, which will not affect production.

[0107] Meanwhile, a spare antenna was deployed to workshop A to ensure a smooth replacement process; in response to the predicted congestion risk during the peak hours on Thursday, the IT department was coordinated in advance to dynamically adjust the base station frequency band, switching from 2.6GHz to 3.5GHz, increasing base station capacity by 20% and avoiding congestion; through the linkage between prediction and maintenance, the pre-processing plan was ensured to be targeted and realistic, avoiding blindly maintaining in advance; at the same time, the maintenance plan was combined with the predicted risks to achieve preventive maintenance and curb the occurrence of disconnection from the source.

[0108] Combining Embodiment 1 and Embodiment 2, this solution employs a dual design that enhances the accuracy of judgment through multi-source data linkage and achieves proactive control through a prediction-maintenance closed loop, thereby comprehensively optimizing the industrial plant's communication network monitoring system. On the one hand, compared to existing technologies that rely solely on online data for single analysis, this solution integrates multi-dimensional data from online sources, local environments, and equipment maintenance. Through feature transformation and association rule analysis, it significantly reduces the false alarm rate and improves the efficiency of maintenance resource utilization.

[0109] On the other hand, it breaks through the limitations of traditional passive monitoring, takes predictive analysis as the core and maintenance and inspection linkage as the support, upgrades risk handling from post-event response to pre-event prevention, effectively shortens the abnormal response time, and avoids the impact of disconnection on production efficiency and safe production; it has built a full-process management system of precise monitoring-early warning-efficient handling, effectively ensures the stability of intelligent equipment communication links, and provides reliable communication support for digital production in the factory area.

[0110] The above description is merely an example and illustration of the structure of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the structure of the invention or exceed the scope defined in the claims, all of which should fall within the protection scope of the present invention.

[0111] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0112] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A method for monitoring the operational status of a communication network, characterized in that, Includes the following steps; S1. Determine the scope of supervision and multi-source data collection, clarify the monitoring boundaries, collect raw data from all dimensions, and obtain online network data, local environmental data and equipment maintenance data to construct the raw data. S2. After obtaining the raw data, data preprocessing and feature transformation are performed to solve the problems of inconsistent units, errors, and incomparability of the raw data. Finally, three types of risk characteristic indicators that can be compared horizontally are generated. S3. Based on risk characteristic indicators and combined with the historical disconnection patterns in the plant area, formulate the correlation rules among signal, hardware and congestion, and accurately judge the abnormal attribution results of intermittent network disconnection through cross analysis. S4 relies on the edge computing nodes of the factory area to process data quickly locally, make real-time judgments on the abnormal attribution results of S3, and generate early warning signals of corresponding levels. S5. Predict future network risks based on historical data, and verify the prediction results by combining local daily maintenance and inspection information to achieve closed-loop management of prediction and maintenance and inspection.

2. The method for monitoring the operating status of a communication network according to claim 1, characterized in that, The process of obtaining the original data in step S1: The core monitoring object is the communication link of intelligent equipment in the industrial plant area, as well as the communication status monitoring function of the intelligent equipment itself. The signal strength value reflecting the tightness of the signal connection between the equipment and the base station, the real-time traffic value reflecting the network usage intensity, and the base station load rate value reflecting the base station's carrying pressure are extracted and summarized into online network data. By deploying environmental sensors in each sub-area, the impact of the environment on hardware equipment is extracted and summarized as temperature and humidity values; the dust content in the storage area and welding workshop is summarized as dust concentration values; and the metal raw materials temporarily piled up in the workshop and the newly added stacks of goods in the storage area are summarized as temporary obstacle record values, which are then combined to generate local environmental data. Through the plant equipment maintenance and management system, the recent antenna replacement time and baseband chip test results are extracted and marked as historical maintenance and inspection records. The antenna VSWR and baseband chip signal demodulation rate are marked as real-time hardware parameters. The number of antenna failures and the cause of failures for the corresponding equipment are obtained and marked as hardware failure records. The equipment maintenance and inspection data are obtained, and the three sets of data are summarized into raw data.

3. The method for monitoring the operating status of a communication network according to claim 1, characterized in that, The processing procedure for the original data in step S2 is as follows: The monitoring time from the start time of data collection to the current end time of data collection is obtained and marked as the monitoring period. The monitoring period is divided into i period nodes, where i is a natural number greater than zero. Missing values ​​are processed for the collected online network data, local environment data, and equipment maintenance data in the starting order of the period nodes. If signal strength data for a certain period node is missing, first check the average signal strength of the adjacent 10 seconds in the same sub-region, and then combine it with the local environmental records for that period. If the records show no temporary obstruction, supplement it with the average of the adjacent time periods; if there is obstruction, supplement it with the average of the adjacent time periods reduced by 10%, because obstacles will weaken the signal; if temperature and humidity data are missing, directly use the historical average temperature and humidity of the same area at the same time of the day to supplement it.

4. The method for monitoring the operating status of a communication network according to claim 3, characterized in that, The process of generating risk characteristic indicators in step S2 is as follows: The raw data of the three types of cleaned data are transformed into three risk characteristic indicators with a unified value range of 0-1 or a fixed interval. The actual values ​​of signal strength and base station load rate in online network data are extracted to find the intervals within the normal range, generating signal health and congestion risk values. Antenna VSWR and baseband chip signal demodulation rate in equipment maintenance data are also extracted. The qualified standards of the two types of hardware parameters are determined by the position of the converted values ​​within the 0-1 interval, generating hardware health values. The risk characteristic indicators are obtained by summarizing the three sets of values ​​from the previous analysis.

5. The method for monitoring the operating status of a communication network according to claim 1, characterized in that, The process of generating association rules in step S3 is as follows: We extracted 50 network disconnection events from the past few years of our factory, analyzed the specific values ​​of signal health, hardware health, and congestion risk at the time of each disconnection, summarized the three most common causes of disconnection, and referred to typical cases of network disconnection attribution in the communications industry to finally determine the following three core rules. When the signal health value is at a low level and the hardware health value is also at a low level, the equipment in the weak signal area needs to frequently switch base stations to maintain the connection. Hardware failure will greatly increase the probability of switching failure, which will eventually lead to frequent network disconnection. This will be summarized and marked as Rule 1.

6. The method for monitoring the operating status of a communication network according to claim 5, characterized in that, When the signal health value is at a high level, but the congestion risk value exceeds 1, the corresponding base station load rate value exceeds the overload threshold of 90%. Even if the signal is stable, the base station may experience processing delays or even be unable to respond due to the large number of devices it supports and the large amount of data transmission, which will eventually lead to network disconnection. This is summarized and marked as Rule 2. When the signal health value is at a low level, the hardware health value is at a normal level (above 0.6), there are no obvious hardware faults, and the congestion risk value is between 0.8 and 1, the weak signal itself will lead to a decrease in data transmission stability, and mild congestion will further increase the transmission delay. The combination of the two will increase the probability of disconnection by more than 25% compared to a single factor. This combination is summarized and marked as Rule 3.

7. The method for monitoring the operating status of a communication network according to claim 6, characterized in that, Retrieve the signal health value, hardware health value, and congestion risk value of the most recent period node in a certain sub-region, as well as the corresponding record of the number of disconnections in the network connection status of that region: Calculate the total number of disconnections within the most recent period node in that sub-region; compare each of the three characteristic indicators at the time of each disconnection with the above three rules to determine which rule each disconnection conforms to; calculate the proportion of disconnections corresponding to each rule to determine the primary and secondary causes of disconnections in that region, and obtain the anomaly attribution results.

8. The method for monitoring the operating status of a communication network according to claim 1, characterized in that, The warning signal is generated in step S4 as follows: By combining the importance / weight of the three types of characteristic indicators on network risk, an indicator reflecting the overall risk level is formed and marked as the comprehensive risk value. The influence of hardware health value accounts for 40%, signal health value accounts for 30%, and congestion risk value after balancing and adjustment accounts for 30%. The pre-stored risk thresholds are retrieved and compared with the comprehensive risk value to obtain low risk thresholds, medium risk thresholds, and high risk thresholds. Based on the risk level and the anomaly attribution results, corresponding early warning signals are generated.

9. The method for monitoring the operating status of a communication network according to claim 1, characterized in that, In step S5, the closed-loop management generation process obtains prediction data input based on the processes of S1-S4, combines the pre-stored prediction model training and output to obtain network risk prediction results, and based on the network risk prediction results of future periodic nodes and the corresponding equipment maintenance data, combines the maintenance plan with the predicted risks through the linkage of prediction and maintenance to achieve preventive maintenance and curb the occurrence of disconnection from the source.

10. A communication network operation status monitoring system, used in the communication network operation status monitoring method according to any one of claims 1-9, characterized in that, Includes the following steps: Multi-source data acquisition module: Based on the selected data acquisition range, it is responsible for the acquisition and transmission of online network data, local environmental data, and equipment maintenance data; The data preprocessing and analysis module cleans, transforms features, and attributes anomalies to the collected raw data, reducing processing latency and meeting real-time requirements. The real-time judgment and prediction module analyzes the comprehensive risk value in real time and generates early warning signals. The prediction unit outputs future risk prediction results based on historical data. The two modules communicate with each other to ensure that the prediction is based on the real-time status.

Citation Information

Patent Citations

  • Fault monitoring method, system and equipment for power distribution equipment and medium

    CN119884596A

  • High-risk environment operation emergency response method and system based on intelligent communication equipment

    CN120583451A

  • Method and system for evaluating time service precision of network time server

    CN120768495A

  • Fault monitoring in a communications network

    GB202011876D0

Cited By

  • Electricity utilization safety monitoring and power saving control method and platform

    CN121813690A