A fault detection method for a wireless communication network

By collecting multi-dimensional network data in wireless communication networks and performing pre-processing and feature extraction, combining clustering algorithms and fault detection strategies, the problem of insufficient fault detection accuracy in complex networks is solved, and efficient and accurate fault detection and positioning is achieved.

CN119485420BActive Publication Date: 2025-06-10SHANGHAI LANYANG NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510067669.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-06-10
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The prior art is difficult to accurately and timely detect and locate faults in complex wireless communication networks, resulting in false alarms, missed alarms and high complexity of handling.

Method used

By collecting data on base station signal strength, network delay, data packet loss rate and user equipment signal quality in wireless communication networks, preprocessing and feature extraction, and fault detection and location are used to detect and locate faults using clustering algorithms and fault detection strategies.

Benefits of technology

It improves the accuracy and timeliness of fault detection, can effectively identify and distinguish different types of fault modes, reduces false alarms and missed alarm rates, and improves network operation and maintenance efficiency and service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119485420B_ABST
    Figure CN119485420B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of fault detection for wireless communication networks, and discloses a fault detection method for wireless communication networks. By collecting data on base station signal strength, network latency, data packet loss rate, and user equipment signal quality, the problem of insufficient fault detection caused by a single indicator is solved. Multi-dimensional data collection enables the system to comprehensively evaluate network performance, improve the accuracy of fault detection, and effectively identify different fault types especially in complex fault modes. By storing data in chronological order, the problem of time series chaos is avoided, ensuring the accurate detection of time-varying faults. The comprehensive data vector generated by each sampling contains multiple network performance indicators, improving the ability of fault detection and location. Data preprocessing eliminates outliers, ensuring data quality, providing an accurate basis for subsequent analysis, and ensuring the reliability of detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fault detection for wireless communication networks, and specifically to a fault detection method for wireless communication networks. Background Art

[0002] With the continuous development of wireless communication technology, especially the increasingly complex application scenarios of 5G and future 6G networks, the reliability, stability, and efficiency of communication networks have become particularly important. In wireless communication networks, due to the influence of various factors, such as hardware failures, network topology changes, external environment changes, etc., network failures occur frequently. These failures not only affect the user experience but may even lead to the paralysis of the entire system. Therefore, the timely detection and location of network failures are of great significance for improving network operation and maintenance efficiency, reducing maintenance costs, and ensuring network service quality.

[0003] Existing fault detection methods usually rely on rule-based systems or experience-based diagnostic means. For example, some methods rely on threshold settings, and when network signals or quality indicators exceed the preset thresholds, it is considered that a failure has occurred. Although these traditional methods can provide certain help in simple network environments, with the diversification of network environments and fault modes, these methods often cannot accurately and timely detect the occurrence of complex faults. For example, in large-scale wireless communication networks, traditional fault detection methods may cause false alarms and missed alarms due to their dependence on fixed rules, especially when facing various complex network fault types, their accuracy and adaptability are relatively insufficient.

[0004] Therefore, this case aims to propose a fault detection method for wireless communication networks, mainly a data-driven fault detection and location method, aiming to solve the problems of insufficient accuracy and high processing complexity existing in the prior art. Summary of the Invention

[0005] The present invention provides a fault detection method for wireless communication networks, which helps to solve the problems mentioned in the above background art.

[0006] The present invention provides the following technical solution: A fault detection method for wireless communication networks, comprising:

[0007] Collecting data from a wireless communication network;

[0008] The data collection includes:

[0009] The base station signal strength, denoted as , representing the signal strength from the base station to the user equipment at time t, with the unit of dBm;

[0010] The network delay, denoted as , representing the latency of communication between the user equipment and the base station at time t, with the unit of ms;

[0011] The data packet loss rate, denoted as , representing the packet loss ratio of data transmission at time t, expressed as a percentage;

[0012] The signal quality of the user equipment, denoted as , representing the quality index of the received signal of the user terminal equipment at time t, with the unit of dBm;

[0013] Set the sampling period to , and the total observation time is , calculate the number of sampling points:

[0014] ;

[0015] Among them, is the total number of sampling points;

[0016] Each sampling generates a data vector, denoted as:

[0017] ;

[0018] Among them, is the comprehensive data vector at time t, including signal strength, network latency, data packet loss rate, and signal quality;

[0019] Store the sampled data into the database in chronological order;

[0020] Preprocess the collected data.

[0021] Optionally, the preprocessing of the collected data specifically includes:

[0022] Classify and extract the data in the database according to the base station signal strength, network latency, data packet loss rate, and user equipment signal quality, form the data of the same type into a data set, and sort the classified and extracted data in chronological order of storage time;

[0023] For any type of data set, perform the following steps:

[0024] S1. Use the three - standard - deviation method to detect outliers, specifically:

[0025] Calculate the average value of all data in the data set:

[0026] ;

[0027] Among them, is the mean of the data set; M is the total number of all elements in the data set; is the value of the i-th data point in the dataset;

[0028] Calculate the standard deviation:

[0029] ;

[0030] where, is the standard deviation of the dataset;

[0031] When the data point meets the condition: then mark the data point as an outlier;

[0032] where, is the absolute value of the difference between the data point and the mean ;

[0033] S2. Use linear interpolation to fill in the missing values:

[0034] When the data point has missing data, use linear interpolation to fill in the missing values, specifically:

[0035] ;

[0036] where, is the filled data value; is the data point before the missing value; is the data point after the missing value;

[0037] Perform feature extraction on the preprocessed dataset.

[0038] Optionally, the performing feature extraction on the preprocessed dataset specifically includes:

[0039] Set the sliding window size to k, and extract the following features at each time point:

[0040] S3. Extract the sliding mean:

[0041] ;

[0042] where, is the sliding mean at time t; k is the sliding window size; is the r-th data point within the window;

[0043] S4. Extract the sliding standard deviation:

[0044] ;

[0045] where, is the sliding standard deviation at time t;

[0046] S5. Extract the change rate:

[0047] ;

[0048] where is the change rate at time t; is the data point at time t; is the data point at the starting time of the window;

[0049] Combine the original data and features to form an extended feature vector:

[0050] ;

[0051] where is the extended feature vector at time t;

[0052] Based on the extended feature vector, execute the fault detection strategy.

[0053] Optionally, the step of executing the fault detection strategy based on the extended feature vector specifically includes:

[0054] Set the normal operating range of each data type index:

[0055] S6. ;

[0056] S7. ;

[0057] S8. ;

[0058] S9. ;

[0059] Set the fault label:

[0060] ;

[0061] where is the fault status label at time t;

[0062] When it indicates a fault;

[0063] When it indicates normal.

[0064] Optionally, the step of executing the fault detection strategy based on the extended feature vector specifically includes:

[0065] Calculate the similarity of the extended feature vector:

[0066] ;

[0067] Among them, is the extended feature vector and similarity; , are the i-th and j-th extended feature vectors respectively; is the Gaussian kernel width parameter;

[0068] Initialize the cluster center as a randomly selected sample point, and update iteratively:

[0069] ;

[0070] Among them, is the center of the k-th class after the (t + 1)-th iteration; is the sample set included in the k-th class; is the number of samples in the k-th class.

[0071] Optionally, the fault detection strategy is executed based on the extended feature vector, specifically including:

[0072] All extended feature vectors are randomly assigned to K initial clusters, and the cluster center is a randomly selected sample point:

[0073] ;

[0074] Among them, is the center of the initial cluster k; p is a random index;

[0075] Calculate the similarity between each sample and all cluster centers, and assign the sample to the nearest cluster:

[0076] ;

[0077] Among them, is the sample set included in the k-th cluster; is the sample to the cluster center similarity;

[0078] Update the cluster center according to the sample set of each cluster:

[0079] ;

[0080] Among them, is the center of the k-th class after the (t + 1)-th iteration; is the sample set of the k-th class; is the number of samples in the k-th class;

[0081] Set the clustering convergence threshold to ;

[0082] Check whether the change of the cluster center is lower than the threshold :

[0083] ;

[0084] If the changes of all cluster centers are lower than the threshold, the clustering process terminates;

[0085] Classify the faults.

[0086] Optionally, the classification of the faults specifically includes:

[0087] After clustering is completed, each cluster is mapped to a fault mode or a normal state, specifically:

[0088] Pre-defined types of fault modes ;

[0089] Use the majority voting method to assign a fault label to each cluster :

[0090] ;

[0091] where is the label of cluster k; is the indicator function, which takes the value of 1 if and 0 otherwise;

[0092] When there is a cluster that fails to match an existing fault mode, it is marked as an unknown fault , and its characteristics are recorded;

[0093] Locate the detected faults.

[0094] Optionally, the location of the detected faults specifically includes:

[0095] Set the set of identified fault modes, denoted as set Y;

[0096] Obtain all identified fault modes and add them to set Y;

[0097] For each fault mode , calculate the similarity with the collected data :

[0098] ;

[0099] where and respectively represent the values of the j-th feature in the fault mode and the data ; and are the mean values of the failure mode and data respectively; is the weight of the j-th feature;

[0100] By calculating the similarity of each failure mode, calculate the priority of each failure mode:

[0101] ;

[0102] wherein, represents the priority of the failure mode ; d is the total number of failure modes; is the failure mode and the current data similarity;

[0103] Arrange all failure modes in descending order according to the priority ;

[0104] Output the failure mode with the highest priority and the data with the highest similarity to the failure mode with the highest priority.

[0105] The present invention has the following beneficial effects:

[0106] 1. By collecting data on base station signal strength, network latency, data packet loss rate, and user equipment signal quality from a wireless communication network, the problem of insufficient fault detection relying on a single metric is solved. Traditional methods usually rely on only a single network performance metric, which may lead to low accuracy in fault detection. Especially when facing complex fault patterns, they cannot comprehensively reflect the network status. Collecting multi-dimensional network data enables the system to comprehensively evaluate network performance, capture finer-grained information, and helps improve the accuracy of fault detection. Especially in complex environments, it can effectively identify and distinguish different types of fault patterns. By storing data in chronological order, the problem of chaotic data time order is solved. The state of a wireless network is a dynamically changing process, and the time series nature of data is very important. Many faults may occur instantaneously and have a certain time dependence. If the data storage does not maintain the order of the time series, it may lead to incorrect timing judgments during the analysis process, affecting the accuracy of fault detection. By ensuring the chronological storage of data, the system can accurately capture the dynamic process of the network state changing over time, thereby enhancing the ability to perceive time-varying faults and effectively improving the accuracy of data analysis and fault prediction. By generating a comprehensive data vector for each sample, including base station signal strength, network latency, data packet loss rate, and user equipment signal quality, the problem of missing information that cannot comprehensively express the network state is solved. Data vectorization enables the system to use more comprehensive information for fault detection and location in subsequent processing and analysis stages, improving the system's analysis ability and judgment accuracy. Data preprocessing can effectively remove invalid data and outliers, improve data quality, and provide a clearer and more accurate basis for subsequent fault analysis and pattern recognition. By removing interference, the accuracy of subsequent steps is guaranteed, and the fault detection results are ensured to be more credible.

[0107] 2. By classifying and extracting the data in the database according to base station signal strength, network latency, data packet loss rate, and user equipment signal quality, and sorting them in the order of storage time, the problem of messy data and difficult unified analysis is solved. Different types of performance indicators in a wireless communication network have different data characteristics and time dependencies. Without differentiation and sorting, information loss or misinterpretation may occur during the data analysis process. Through classification extraction and sorting, it can be ensured that each type of data is processed in the correct time series. By using the three - standard - deviation method to detect outliers, the interference problem caused by noise data or outliers to subsequent analysis is solved. The data in a wireless communication network is affected by various factors, such as equipment failures and environmental interference, and extreme outliers may appear. If these outliers are not processed, they may mislead the fault detection model and result in incorrect fault diagnosis results. Through the detection and elimination of outliers, the quality and consistency of the data are guaranteed, providing a reliable data basis for subsequent feature extraction and fault detection. The elimination of abnormal data reduces the detection error caused by noise data interference, thereby improving the accuracy and credibility of fault detection. By using linear interpolation to fill in the missing values, the problem of incomplete analysis caused by data loss is solved. In a wireless communication network, due to reasons such as unstable signals and equipment failures, missing values may occur during the data collection process. If these missing data are not processed, it may lead to data missing in subsequent analysis and affect the overall analysis result. Linear interpolation estimates the missing values by using the values of adjacent data points, thereby smoothing the time series of the data and ensuring the continuity and integrity of the data. After filling in the missing values, the data set is more complete, ensuring the continuity of feature extraction and subsequent analysis. This not only avoids the computational difficulties caused by data missing but also improves the accuracy of the overall analysis result, ensures the full utilization of the data, and enhances the reliability of fault detection. By performing feature extraction on the pre - processed data set, the interference problem of redundant and irrelevant information in the original data is solved. The data in a wireless communication network contains a large amount of information, but not all data is useful for fault detection. Through feature extraction, the most valuable features for fault diagnosis can be selected from a large amount of original data, thereby improving the efficiency of subsequent fault location and analysis. Feature extraction makes data analysis more focused on key parameters, reduces the interference of irrelevant data, can improve the efficiency and accuracy of subsequent analysis algorithms, and further enhances the overall performance of the fault detection system.

[0108] 3. By extracting the moving average, the influence of local fluctuations and interferences in the data is addressed. The data in a wireless communication network is usually subject to instantaneous fluctuations and noises. These short-term fluctuations may not represent the actual condition of the system and are likely to mislead fault detection. Through the calculation of the moving average, the short-term fluctuations in the data can be smoothed, and a more stable trend can be extracted. The moving average helps reduce the noises and random fluctuations in the data, ensuring that the fault detection algorithm focuses on the real trend changes rather than short-term fluctuations. This processing makes the subsequent fault detection more stable and enhances the sensitivity to long-term trend changes. By extracting the moving standard deviation, the problem of measuring the amplitude and degree of data variation is solved. In a wireless communication network, the changes in indicators such as signal strength and latency are not always uniform, and some fault modes may manifest as abnormal increases in indicator fluctuations. The moving standard deviation can effectively quantify these fluctuations and help detect abnormal change amplitudes. The moving standard deviation provides a dynamic fluctuation measurement criterion, which can reflect the degree of change in network performance within a certain time window. This helps detect abnormal fluctuations in the system, identify possible faults in advance, and enhances the sensitivity and accuracy of the fault detection algorithm. By extracting the rate of change, the problem of being unable to capture rapid changes in network indicators and emergencies is solved. In a wireless communication network, some faults may manifest as rapid signal changes, and traditional static indicators may not be able to accurately capture this. The extraction of the rate of change can measure the speed of change between a certain time point and the previous time point, helping to identify sudden faults. By calculating the rate of change of the data, the system can detect the speed of data change and emergencies, which is crucial for timely identifying sudden faults in the network. The extraction of the rate of change improves the timeliness of fault detection, enabling the system to respond to rapid changes in the network and quickly locate possible fault sources. By combining the original data with the extracted features, such as the moving average, the moving standard deviation, and the rate of change, to form an extended feature vector, the problem of insufficient information in a single data point is solved. A single original data point often cannot provide enough information to accurately judge faults. Combining multiple features together to form an extended feature vector can represent the network state more comprehensively. The construction of the extended feature vector provides richer data input for subsequent fault detection. By integrating information from multiple dimensions, the detection system can make more accurate fault judgments based on a more comprehensive network state. The fusion of features enhances the overall robustness of the system and improves the diversity and accuracy of fault detection.

[0109] 4. By setting the normal operating range for each data type, such as base station signal strength, network latency, data packet loss rate, user equipment signal quality, etc., the problem of how to distinguish between normal and abnormal states is solved. In a wireless communication network, different performance metrics have their respective normal fluctuation ranges. An overly broad or inappropriate normal range may lead to inaccurate identification of faults. By presetting the normal operating range for each metric, it is possible to more accurately determine whether the data is within the normal range, thus effectively distinguishing between normal and abnormal states. This operation effectively improves the accuracy of fault detection. By clarifying the normal range, the system can more clearly distinguish between fault states and normal states, avoiding misjudgments and missed detections, and improving the stability and reliability of the fault detection system. By setting fault tags to identify the fault states of each data point, the problem of being unable to clearly identify the fault location in the data is solved. In practical applications, fault states in the network may be instantaneous, sudden, or persistent. Without a clear fault tag to identify the occurrence of a fault, subsequent location and repair work will become difficult. Through the setting of fault tags, the system can clearly identify whether a fault has occurred at each data point in time, thus providing a clear identification for fault location and recovery. The setting of fault tags makes the identification of faults more accurate and clear, providing very important support for subsequent fault location. This fault representation method based on tagging simplifies the operation process of the fault detection system and improves the speed and accuracy of fault response. Through the judgment of fault tags, the system can quickly identify the moment when a fault occurs in the real-time data stream, avoiding the defects of manual judgment or delayed judgment in traditional methods. Such a real-time fault marking and judgment mechanism improves the efficiency of network fault detection and accelerates the speed of fault response, providing instant alarm information for network operation and maintenance personnel, and further enhancing the stability of the network and the user experience.

[0110] 5. By calculating the similarity between extended feature vectors, the problem of how to identify similar fault patterns in a multi-dimensional data space is solved. Each extended feature vector contains multiple metrics, such as signal strength, delay, packet loss rate, signal quality, etc. Through similarity calculation, samples with similar characteristics can be effectively identified, thus helping the system understand which data points belong to the same type of fault pattern. The introduction of the Gaussian kernel function can make the similarity measure smoother and can handle the complex relationships between different feature dimensions. By calculating similarity with the Gaussian kernel, it is possible to effectively focus on fault patterns with similar manifestations, thereby improving the accuracy of fault detection. Compared with traditional Euclidean distance or other simple similarity calculation methods, the Gaussian kernel function can better capture the non-linear relationships and potential complex patterns in the data, effectively improving the accuracy of fault pattern recognition. By initializing the cluster centers as randomly selected sample points, the problem of how to efficiently start the clustering process is solved. In cluster analysis, the selection of the initial cluster centers has an important impact on the final clustering result. By randomly selecting some sample points from the dataset as the initial centers, it is possible to avoid over-reliance on certain prior assumptions and increase the flexibility and adaptability of the clustering process. This step enables the clustering process to start from diverse initial conditions, enhancing the diversity and exploratory nature of the clustering results and avoiding the problem of local optimal solutions caused by improper selection of the initial points. Randomly initializing the cluster centers helps improve the effectiveness of the clustering algorithm and avoid the clustering falling into local optimal solutions. By iteratively updating the cluster centers, the problem of how to continuously optimize the results during the clustering process is solved. By continuously adjusting the cluster centers, the system can better adapt to changes in the data and improve the clustering accuracy. In each iteration, the cluster centers are updated according to the current set of samples in the class, thus gradually converging to the true fault patterns. Iterative update enables the clustering results to be optimized as the data is continuously adjusted, avoiding errors caused by inaccurate initial cluster centers. The iterative update of the cluster centers greatly enhances the flexibility and accuracy of the clustering algorithm. Each update enables the cluster centers to more accurately reflect the distribution of the data, thereby improving the accuracy of fault pattern recognition. In wireless communication network fault detection, real-time data and network conditions may change at any time. Iterative update enables the clustering algorithm to dynamically adapt to these changes and maintain an efficient fault identification ability. By performing multiple iterative updates of the cluster centers, the problem of how to ensure that the clustering algorithm finally converges to the correct fault pattern is solved. In multiple iterations, the cluster centers will gradually approach the true distribution of their affiliated samples, making the clustering results more and more accurate. The convergence of the clustering ensures that the final fault detection results can reflect the true structure of the data, avoiding errors caused by insufficient iteration or early stopping. The process of cluster center convergence guarantees the stability and accuracy of the clustering results, enabling fault pattern recognition to accurately reflect the actual state of the network and reducing misjudgments caused by non-convergence or unstable results of the clustering.

[0111] 6. By randomly assigning all extended feature vectors to K initial clusters and initializing the cluster centers as randomly selected sample points, the problem of how to reasonably initialize the clustering centers is solved. In traditional clustering algorithms, the selection of the initial cluster centers is crucial for the clustering effect. Randomly selecting the cluster centers can increase the diversity of clustering, avoiding over-reliance on certain prior knowledge or assumptions. Randomly initializing the cluster centers makes the clustering process more flexible, capable of exploring more potential fault patterns and avoiding the bias caused by over-reliance on preliminary data analysis. Through this initialization method, the algorithm can better adapt to the complex and variable fault patterns in the network, improving the exploratory and accuracy of clustering. By calculating the similarity between each sample and all cluster centers and assigning the sample to the closest cluster, the problem of how to reasonably classify samples according to their characteristics is solved. By calculating the similarity between the sample and the cluster centers, the sample can be classified into its most likely fault pattern, ensuring the accurate attribution of the fault pattern. This process can efficiently aggregate similar fault samples together, avoiding misclassification and false judgment. In a wireless communication network, different fault patterns may have similar manifestations. Clustering similarity calculation can help the algorithm quickly identify and accurately group them, thereby improving the accuracy of fault detection. By updating the cluster centers according to the sample sets of each cluster, the problem of how to iteratively optimize the clustering results is solved. By continuously updating the cluster centers, the clustering can gradually converge to a more accurate state, thus better reflecting the actual distribution of the samples. This step improves the adaptability and accuracy of the clustering algorithm, enabling the clustering to be optimized as the data changes. By updating the cluster centers, the system can continuously optimize the clustering results, making the fault pattern represented by each cluster more accurate. The updated cluster centers can better capture the changes in the data, increasing the stability of the clustering and ensuring the accurate identification of real-time fault patterns in the wireless communication network. By setting a clustering convergence threshold and checking whether the change in the cluster centers is lower than the threshold, the problem of how to determine whether the clustering algorithm has reached convergence is solved. By setting the threshold, it is possible to automatically determine whether the clustering is stable at the end of the clustering process, avoiding unnecessary calculations and iterations and improving the efficiency of the algorithm. Setting the clustering convergence threshold enables the clustering process to terminate in a timely manner after reaching a stable state, avoiding the computational waste caused by over-iteration and ensuring the accuracy and efficiency of the clustering results. In wireless communication network fault detection, quickly terminating the clustering process helps improve the real-time performance and response speed of the system. By classifying the faults after the clustering process is completed, the problem of how to extract specific fault patterns from the clustered samples is solved. By classifying the samples within each cluster, each fault pattern can be accurately located and marked, providing clear guidance for subsequent fault troubleshooting and handling.

[0112] 7. Solved the problem of how to convert the clustering results into specific fault mode labels. After the clustering step is completed, by corresponding each cluster to a predefined fault mode or normal state, the clustering results can be converted into specific fault types, thus providing clear guidance for subsequent processing. This step ensures that each clustering cluster can be accurately labeled as a specific fault mode, avoiding the ambiguity or misunderstanding of the clustering results. This step enables the clustering results to be docked with the actual fault modes, enhancing the practical application value of fault detection. By corresponding to the actual fault modes, each network problem can be accurately labeled, enhancing the interpretability and operability of the detection, and providing clearer fault information for network maintenance personnel. Solved the problem of how to handle the fault label assignment for each cluster. In practical applications, some clusters may contain multiple different fault modes. Using the majority voting method can effectively avoid the influence of human intervention on the labels, making the label assignment more objective and fair. The majority voting method selects the label with the highest frequency of occurrence as the final cluster label by counting the labels of each sample in the cluster, thus ensuring the reliability of the cluster label. This step effectively reduces the randomness and error in label assignment, improving the accuracy of classification. Especially when facing diverse fault samples, the majority voting method can ensure an accurate reflection of the main fault modes within the cluster, enhancing the robustness of the system. Solved the problem of how to handle unknown faults. In a wireless communication network, due to the diverse and rapidly changing fault types, there may be some unknown fault modes. By identifying the clusters that do not match the predefined fault modes during the classification process and labeling them as "unknown faults", new and unforeseen fault types can be discovered in a timely manner. This step provides the necessary flexibility for the comprehensive monitoring of network faults. This step effectively improves the adaptability of fault detection, ensuring that the system can handle and record unknown faults, avoiding potential problems being overlooked. Over time and with the discovery of new faults, the characteristics of these unknown fault clusters can be further accumulated, ultimately providing data support for building a more comprehensive fault mode library, thereby continuously improving the accuracy and comprehensiveness of the system.

[0113] 8. By setting the set Y of identified fault modes and obtaining all fault modes, the problem of how to systematically manage and organize the identified fault modes is solved. In a wireless communication network, the number of fault modes is huge and may increase over time. Therefore, storing all identified fault modes centrally in the set Y can help the system effectively track historical fault modes and provide a complete fault mode reference for subsequent fault location. The maintenance and update of this set ensure that the system can continuously optimize and enrich the fault mode library as the network state changes. By constructing a complete set Y of fault modes, the system can flexibly handle new fault types and modes, and at the same time use the existing fault modes for quick comparative analysis, improving the accuracy and efficiency of fault location. The management of historical fault modes helps reduce the detection and resolution time of future similar problems. By calculating the similarity between each fault mode and the collected data, the problem of how to compare fault modes with real-time collected data to accurately determine the fault type is solved. In an actual network, fault modes are diverse, and the characteristics of each fault mode may vary at different time points or environmental conditions. By calculating the similarity between each fault mode and the collected data, the fault mode most relevant to the current data can be efficiently identified, avoiding simple rule matching or overly rough judgment methods. This step makes the matching between fault modes and real-time data more accurate by quantifying the similarity index. The similarity score between each fault mode and the data provides a clear basis for subsequent priority calculation, further improving the accuracy of fault detection. This method can reduce the false alarm rate and missed alarm rate, and improve the intelligent level of fault location. By calculating the priority of each fault mode, the problem of how to evaluate the severity and urgency of fault modes based on similarity is solved. In a wireless communication network, different fault modes may have different impacts on the network. Some faults may cause a major collapse of the system, while others may be small-scale problems. By calculating the priority of each fault mode, it can be ensured that the fault location system can give priority to faults with greater impact on the network, thus improving the response speed and processing efficiency of network operation and maintenance. This step makes the system able to automatically judge and focus on the most critical faults by quantifying the priority of each fault mode, avoiding waste of resources and redundancy in processing. Through priority sorting, network operation and maintenance personnel can quickly handle high-priority faults, reduce the network interruption time, and improve the availability and stability of the network. By sorting all fault modes from high to low according to priority, the problem of how to sort multiple fault modes and select the most urgent fault for processing is solved. Since multiple fault modes may be detected simultaneously during the fault detection process, how to sort these fault modes and determine the processing order directly affects the efficiency and effect of fault handling. By sorting fault modes according to priority, the system can ensure that the most critical problems are solved first, thus reducing the impact on network services.The sorting mechanism enables network operation and maintenance to quickly identify the most important faults and respond within the shortest time. In this way, the system can allocate resources more efficiently in the face of multiple faults, improve the overall fault repair efficiency, reduce the network fault recovery time, and optimize the network service quality. By outputting the fault mode with the highest priority and the data with the highest similarity, the problem of how to associate the detected faults with specific data characteristics and further analyze and locate them is solved. After sorting the fault modes, outputting the fault mode with the highest priority and the collected data most similar to it helps to accurately lock the time and specific problems when the fault occurs, providing a more intuitive basis for fault troubleshooting. This step helps the operation and maintenance personnel quickly locate the cause and background of the fault by outputting the fault mode with the highest priority and the most similar data. Through the precise docking with the data, the operation and maintenance personnel can quickly make fault repair decisions based on specific historical data and similar network conditions, improving the efficiency and accuracy of fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0114] Figure 1 This is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0115] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0116] Embodiment, refer to Figure 1 A fault detection method for a wireless communication network, comprising:

[0117] Collect data from a wireless communication network;

[0118] The collected data includes:

[0119] The base station signal strength, denoted as , representing the signal strength from the base station to the user equipment at time t, with the unit of dBm;

[0120] The network latency, denoted as , representing the latency of communication between the user equipment and the base station at time t, with the unit of ms;

[0121] The data packet loss rate, denoted as , representing the packet loss ratio of data transmission at time t, expressed as a percentage;

[0122] The signal quality of the user equipment, denoted as , representing the quality metric of the signal received by the user terminal device at time t, with the unit of dBm;

[0123] Set the sampling period to , and the total observation time is , calculate the number of sampling points:

[0124] ;

[0125] Among them, is the total number of sampling points;

[0126] Each sampling generates a data vector, denoted as:

[0127] ;

[0128] Among them, is the comprehensive data vector at time t, including signal strength, network latency, data packet loss rate, and signal quality;

[0129] Store the sampling data into the database in chronological order;

[0130] Preprocess the collected data.

[0131] By collecting data on base station signal strength, network latency, data packet loss rate, and user equipment signal quality from a wireless communication network, the problem of insufficient fault detection relying only on a single metric is solved. Traditional methods usually rely only on a single network performance metric, which may lead to lower accuracy in fault detection, especially when facing complex fault patterns and unable to comprehensively reflect the network status. Collecting multi-dimensional network data enables the system to comprehensively evaluate network performance, capture finer-grained information, and helps improve the accuracy of fault detection, especially in complex environments where it can effectively identify and distinguish different types of fault patterns. By storing data in chronological order, the problem of chaotic data time order is solved. The state of a wireless network is a dynamically changing process, and the time series nature of data is very important. Many faults may occur instantaneously and have a certain time dependence. If the data storage does not maintain the order of the time series, it may lead to incorrect timing judgments during the analysis process, affecting the accuracy of fault detection. By ensuring the chronological storage of data, the system can accurately capture the dynamic process of network status changing over time, thereby enhancing the perception ability of time-varying faults and effectively improving the accuracy of data analysis and fault prediction. By generating a comprehensive data vector for each sampling, including base station signal strength, network latency, data packet loss rate, and user equipment signal quality, the problem of missing information that cannot comprehensively express the network status is solved. Data vectorization enables the system to use more comprehensive information for fault detection and location in subsequent processing and analysis stages, improving the analysis ability and judgment accuracy of the system. Data preprocessing can effectively eliminate invalid data and outliers, improve data quality, and provide a clearer and more accurate basis for subsequent fault analysis and pattern recognition. By removing interference, the accuracy of subsequent steps is ensured, and the reliability of fault detection results is guaranteed.

[0132] The preprocessing of the collected data specifically includes:

[0133] Classify and extract the data in the database according to base station signal strength, network latency, data packet loss rate, and user equipment signal quality, form data sets with the same type of data, and sort the classified and extracted data in the order of storage time;

[0134] For any data set of a certain type, perform the following steps:

[0135] S1. Use the three-sigma method to detect outliers, specifically:

[0136] Calculate the average value of all data in the data set:

[0137] ;

[0138] Among them, is the mean of the dataset, reflecting the overall average level of the data; M is the total number of all elements in the dataset; is the value of the i-th data point in the dataset;

[0139] Calculate the standard deviation:

[0140] ;

[0141] where, is the standard deviation of the dataset, representing the fluctuation or dispersion degree of the data;

[0142] When the data point meets the condition: then the data point is marked as an outlier;

[0143] where, is the absolute value of the difference between the data point and the mean ;

[0144] S2. Use linear interpolation to fill in the missing values:

[0145] When the data point has missing data, use linear interpolation to fill in the missing values, specifically:

[0146] ;

[0147] where, is the filled data value; is the data point before the missing value; is the data point after the missing value;

[0148] Perform feature extraction on the preprocessed dataset.

[0149] By classifying and extracting the data in the database according to the base station signal strength, network latency, data packet loss rate, and user equipment signal quality, etc., and sorting them in the order of storage time, the problem of messy data and difficult unified analysis is solved. Different types of performance indicators in a wireless communication network have different data characteristics and time dependencies. Without distinction and sorting, information loss or misinterpretation may occur during the data analysis process. Through classification extraction and sorting, it can be ensured that each type of data is processed in the correct time series. By using the three-standard-deviation method to detect outliers, the interference problem caused by noise data or outliers to subsequent analysis is solved. The data in a wireless communication network is affected by various factors, such as equipment failures and environmental interference, and extreme outliers may occur. If these outliers are not processed, they may mislead the fault detection model and result in incorrect fault diagnosis results. Through the detection and elimination of outliers, the quality and consistency of the data are guaranteed, providing a reliable data basis for subsequent feature extraction and fault detection. The elimination of abnormal data reduces the detection error caused by noise data interference, thereby improving the accuracy and credibility of fault detection. By using the linear interpolation method to fill in the missing values, the problem of incomplete analysis caused by data loss is solved. In a wireless communication network, due to reasons such as unstable signals and equipment failures, missing values may occur during the data collection process. If these missing data are not processed, it may lead to data missing in subsequent analysis and affect the overall analysis result. The linear interpolation method estimates the missing values by using the values of adjacent data points, thereby smoothing the time series of the data and ensuring the continuity and integrity of the data. After filling in the missing values, the data set is more complete, ensuring the continuity of feature extraction and subsequent analysis. This not only avoids the computational difficulties caused by data missing but also improves the accuracy of the overall analysis result, ensures the full utilization of the data, and enhances the reliability of fault detection. By performing feature extraction on the preprocessed data set, the interference problem of redundant and irrelevant information in the original data is solved. The data in a wireless communication network contains a large amount of information, but not all data is useful for fault detection. Through feature extraction, the most valuable features for fault diagnosis can be screened out from a large amount of original data, thereby improving the efficiency of subsequent fault location and analysis. Feature extraction makes data analysis more focused on key parameters, reduces the interference of irrelevant data, and can improve the efficiency and accuracy of subsequent analysis algorithms, further enhancing the overall performance of the fault detection system.

[0150] The performing of feature extraction on the preprocessed data set specifically includes:

[0151] Set the sliding window size to k, and extract the following features at each time point:

[0152] S3. Extract the sliding mean:

[0153] ;

[0154] Among them, is the moving average at time t; k is the moving window size; is the r-th data point within the window;

[0155] S4. Extract the moving standard deviation:

[0156] ;

[0157] Among them, is the moving standard deviation at time t;

[0158] S5. Extract the change rate:

[0159] ;

[0160] Among them, is the change rate at time t; is the data point at time t; is the data point at the starting time of the window;

[0161] Combine the original data and features to form an extended feature vector:

[0162] ;

[0163] Among them, is the extended feature vector at time t;

[0164] Execute a fault detection strategy based on the extended feature vector.

[0165] By extracting the moving average, the impact of local fluctuations and interferences in the data is resolved. Data in wireless communication networks is usually subject to instantaneous fluctuations and noise. These short-term fluctuations may not represent the actual condition of the system and can easily mislead fault detection. Through the calculation of the moving average, short-term fluctuations in the data can be smoothed, and a more stable trend can be extracted. The moving average helps reduce noise and random fluctuations in the data, ensuring that the fault detection algorithm focuses on real trend changes rather than short-term fluctuations. This processing makes subsequent fault detection more stable and enhances the sensitivity to long-term trend changes. By extracting the moving standard deviation, the problem of measuring the amplitude and degree of data variation is resolved. In wireless communication networks, the changes in indicators such as signal strength and latency are not always uniform, and certain fault patterns may manifest as abnormal increases in indicator fluctuations. The moving standard deviation can effectively quantify these fluctuations and help detect abnormal change amplitudes. The moving standard deviation provides a dynamic measure of fluctuations, which can reflect the degree of change in network performance within a certain time window. This helps detect abnormal fluctuations in the system, identify potential faults in advance, and enhances the sensitivity and accuracy of the fault detection algorithm. By extracting the rate of change, the problem of being unable to capture rapid changes in network indicators and emergencies is resolved. In wireless communication networks, some faults may manifest as rapid signal changes, and traditional static indicators may not be able to accurately capture this. The extraction of the rate of change can measure the speed of change between a certain time point and the previous time point, helping to identify sudden faults. By calculating the rate of change of the data, the system can detect the speed of data change and emergencies, which is crucial for timely identifying sudden faults in the network. The extraction of the rate of change improves the timeliness of fault detection, enabling the system to respond to rapid changes in the network and quickly locate potential fault sources. By combining the original data with the extracted features, such as the moving average, moving standard deviation, and rate of change, to form an extended feature vector, the problem of insufficient information in a single data point is resolved. A single original data point often cannot provide enough information to accurately judge faults, while combining multiple features to form an extended feature vector can represent the network state more comprehensively. The construction of the extended feature vector provides richer data input for subsequent fault detection. By integrating information from multiple dimensions, the detection system can make more accurate fault judgments based on a more comprehensive network state. The fusion of features enhances the overall robustness of the system and improves the diversity and accuracy of fault detection.

[0166] Based on the extended feature vector, a fault detection strategy is executed, specifically including:

[0167] Set the normal operating range for each data type indicator:

[0168] S6, ;

[0169] S7, ;

[0170] S8, ;

[0171] S9, ;

[0172] Set the fault label:

[0173] ;

[0174] in, is the fault status label at time t;

[0175] when , it indicates a fault;

[0176] when , indicating normal.

[0177] By setting the normal operating range of each data type, such as base station signal strength, network delay, data packet loss rate, user equipment signal quality, etc., the problem of how to distinguish between normal and abnormal states is solved. In wireless communication networks, different performance indicators have their own normal fluctuation ranges. Too broad or inappropriate normal ranges may lead to inaccurate fault identification. By presetting the normal operating range of each indicator, it is possible to more accurately determine whether the data is within the normal range, thereby effectively distinguishing between normal and abnormal states. This operation effectively improves the accuracy of fault detection. By clarifying the normal range, the system can more clearly distinguish between fault states and normal states, avoid misjudgment and missed judgment, and improve the stability and reliability of the fault detection system. By setting a fault label to identify the fault state of each data point, the problem of being unable to clearly identify the fault location in the data is solved. In actual applications, the fault state in the network may be instantaneous, sudden, or continuous. If there is no clear fault label to identify the occurrence of the fault, subsequent positioning and repair work will become difficult. By setting the fault label, the system can clearly identify whether the data point at each moment has a fault, thereby providing a clear mark for fault location and recovery. The setting of fault labels makes fault identification more accurate and clear, providing very important support for subsequent fault location. This labeled fault representation method simplifies the operation process of the fault detection system and improves the speed and accuracy of fault response. By judging the fault label, the system can quickly identify the time when the fault occurs in the real-time data stream, avoiding the defects of traditional methods that require manual judgment or delayed judgment. Such a real-time fault marking and judgment mechanism improves the efficiency of network fault detection and speeds up the speed of fault response, providing network operation and maintenance personnel with instant alarm information, and further improving network stability and user experience.

[0178] The fault detection strategy is executed based on the extended feature vector, specifically including:

[0179] Compute the similarity of the extended feature vectors:

[0180] ;

[0181] in, is the extended feature vector and similarity; , are the i-th and j-th extended feature vectors respectively; is the Gaussian kernel width parameter;

[0182] Initialize cluster centers For randomly selected sample points, iterative update:

[0183] ;

[0184] in, is the k-th class center after the t+1th iteration; is the sample set contained in the kth class; is the number of samples in the kth category.

[0185] By calculating the similarity between extended feature vectors, the problem of how to identify similar fault modes in multidimensional data space is solved. Each extended feature vector contains multiple indicators, such as signal strength, delay, packet loss rate, signal quality, etc. Through similarity calculation, samples with similar features can be effectively identified, thereby helping the system understand which data points belong to the same type of fault mode. The introduction of Gaussian kernel function can make the similarity measurement smoother and can handle the complex relationship between different feature dimensions. By calculating similarity by Gaussian kernel, it is possible to effectively focus on fault modes with similar performance, thereby improving the accuracy of fault detection. Compared with traditional Euclidean distance or other simple similarity calculation methods, Gaussian kernel function can better capture the nonlinear relationship and potential complex patterns of data, effectively improving the accuracy of fault mode recognition. By initializing the cluster center as a randomly selected sample point, the problem of how to efficiently start the clustering process is solved. In cluster analysis, the selection of the initial cluster center has an important influence on the final clustering result. By randomly selecting some sample points from the data set as the initial center, it is possible to avoid over-reliance on certain prior assumptions and increase the flexibility and adaptability of the clustering process. This step enables the clustering process to start from a variety of initial conditions, enhances the diversity and exploratory nature of the clustering results, and avoids the problem of local optimal solutions caused by improper selection of initial points. Randomly initializing the cluster center helps to improve the effectiveness of the clustering algorithm and prevent clustering from falling into the local optimal solution. By iteratively updating the cluster center, the problem of how to continuously optimize the results in the clustering process is solved. By continuously adjusting the cluster center, the system can better adapt to data changes and improve the accuracy of clustering. In each iteration, the cluster center is updated according to the sample set of the current class, so as to gradually converge to the real fault mode. Iterative updates enable the clustering results to be optimized as the data is continuously adjusted, avoiding errors caused by inaccurate initial cluster centers. Iterative updates of cluster centers greatly enhance the flexibility and accuracy of clustering algorithms. Each update enables the class center to more accurately reflect the distribution of data, thereby improving the accuracy of fault mode recognition. In wireless communication network fault detection, real-time data and network conditions may change at any time. Iterative updates enable clustering algorithms to dynamically adapt to these changes and maintain efficient fault identification capabilities. Through multiple iterative updates of cluster centers, the problem of how to ensure that the clustering algorithm eventually converges to the correct fault mode is solved. In multiple iterations, the cluster center will gradually approach the true distribution of the samples to which it belongs, making the clustering results more and more accurate. The convergence of clustering ensures that the final fault detection results can reflect the true structure of the data and avoid errors caused by insufficient iterations or early stopping. The process of cluster center convergence ensures the stability and accuracy of the clustering results, so that fault mode identification can accurately reflect the actual state of the network and reduce misjudgments caused by clustering failure to converge or unstable results.

[0186] Based on the extended feature vectors, execute a fault detection strategy, which specifically includes:

[0187] Randomly assign all the extended feature vectors to K initial clusters, where the cluster centers are randomly selected sample points:

[0188] ;

[0189] wherein, is the center of the initial cluster k; p is a random index;

[0190] Calculate the similarity between each sample and all cluster centers, and assign the sample to the nearest cluster:

[0191] ;

[0192] wherein, is the set of samples included in the k-th cluster; is the sample to the cluster center similarity;

[0193] Update the cluster centers according to the sample sets of each cluster:

[0194] ;

[0195] wherein, is the center of the k-th class after the (t + 1)-th iteration; is the set of samples of the k-th class; is the number of samples in the k-th class;

[0196] Set the clustering convergence threshold to ;

[0197] Check whether the change in the cluster centers is lower than the threshold :

[0198] ;

[0199] If the changes in all cluster centers are lower than the threshold, the clustering process terminates;

[0200] Classify the faults.

[0201] The problem of how to reasonably initialize the clustering centers is solved by randomly assigning all extended feature vectors to K initial clusters and initializing the cluster centers as randomly selected sample points. In traditional clustering algorithms, the selection of the initial cluster centers is crucial for the clustering effect. Randomly selecting the cluster centers can increase the diversity of clustering and avoid over-reliance on certain prior knowledge or assumptions. Randomly initializing the cluster centers makes the clustering process more flexible, capable of exploring more potential fault patterns and avoiding biases caused by over-reliance on preliminary data analysis. Through this initialization method, the algorithm can better adapt to the complex and variable fault patterns in the network, improving the exploratory and accuracy of clustering. The problem of how to reasonably classify samples according to their characteristics is solved by calculating the similarity between each sample and all cluster centers and assigning the sample to the closest cluster. By calculating the similarity between the sample and the cluster centers, the sample can be classified into its most likely fault pattern, ensuring the accurate attribution of the fault pattern. This process can efficiently group similar fault samples together, avoiding misclassification and false judgment. In a wireless communication network, different fault patterns may have similar manifestations. Clustering similarity calculation can help the algorithm quickly identify and accurately group them, thereby improving the accuracy of fault detection. The problem of how to iteratively optimize the clustering results is solved by updating the cluster centers according to the sample sets of each cluster. By continuously updating the cluster centers, the clustering can gradually converge to a more accurate state, thus better reflecting the actual distribution of the samples. This step improves the adaptability and accuracy of the clustering algorithm, enabling the clustering to be optimized as the data changes. By updating the cluster centers, the system can continuously optimize the clustering results, making the fault pattern represented by each cluster more precise. The updated cluster centers can better capture the changes in the data, increasing the stability of the clustering and ensuring the accurate identification of real-time fault patterns in the wireless communication network. The problem of how to determine whether the clustering algorithm has converged is solved by setting a clustering convergence threshold and checking whether the change in the cluster centers is lower than the threshold. By setting the threshold, it is possible to automatically determine whether the clustering has stabilized at the end of the clustering process, avoiding unnecessary calculations and iterations and improving the efficiency of the algorithm. Setting the clustering convergence threshold enables the clustering process to terminate in a timely manner after reaching a stable state, avoiding computational waste caused by over-iteration and ensuring the accuracy and efficiency of the clustering results. In wireless communication network fault detection, quickly terminating the clustering process helps improve the real-time performance and response speed of the system. The problem of how to extract specific fault patterns from the clustered samples is solved by classifying the faults after the clustering process is completed. By classifying the samples within each cluster, each fault pattern can be accurately located and marked, providing clear guidance for subsequent fault troubleshooting and handling.

[0202] The classification of the faults specifically includes:

[0203] After clustering is completed, each cluster is mapped to a fault mode or normal state, specifically:

[0204] Predefined fault modes ;

[0205] Use the majority voting method to assign a fault label to each cluster :

[0206] ;

[0207] where is the label of cluster k; is the indicator function, which takes the value of 1 if otherwise it takes the value of 0;

[0208] When there is a cluster that fails to match an existing fault mode, it is marked as an unknown fault , and its characteristics are recorded;

[0209] Locate the detected faults.

[0210] The problem of how to convert the clustering results into specific fault mode labels is solved. After the clustering step is completed, by corresponding each cluster to a predefined fault mode or normal state, the clustering results can be converted into specific fault types, providing clear guidance for subsequent processing. This step ensures that each clustering cluster can be accurately labeled as a specific fault mode, avoiding ambiguity or misunderstanding of the clustering results. This step enables the clustering results to be docked with the actual fault modes, enhancing the practical application value of fault detection. By corresponding to the actual fault modes, each network problem can be accurately labeled, enhancing the interpretability and operability of the detection and providing clearer fault information for network maintenance personnel. The problem of how to handle the fault label assignment for each cluster is solved. In practical applications, some clusters may contain multiple different fault modes. Using the majority voting method can effectively avoid the influence of human intervention on the labels, making the label assignment more objective and fair. The majority voting method selects the label with the highest frequency as the final cluster label by counting the labels of each sample in the cluster, thus ensuring the reliability of the cluster label. This step effectively reduces the randomness and error in label assignment, improving the classification accuracy. Especially when facing diverse fault samples, the majority voting method can ensure that the main fault mode within the cluster is accurately reflected, enhancing the robustness of the system. The problem of how to handle unknown faults is solved. In a wireless communication network, due to the diverse and rapidly changing fault types, there may be some unknown fault modes. By identifying clusters that do not match the predefined fault modes during the classification process and labeling them as "unknown faults", new fault types that are not foreseen by the database can be discovered in a timely manner. This step provides the necessary flexibility for the comprehensive monitoring of network faults. This step effectively improves the adaptability of fault detection, ensuring that the system can handle and record unknown faults and avoid missing potential problems. Over time and with the discovery of new faults, the characteristics of these unknown fault clusters can be further accumulated, ultimately providing data support for building a more comprehensive fault mode library, thereby continuously improving the accuracy and comprehensiveness of the system.

[0211] The fault location for the detected faults specifically includes:

[0212] Set the set of identified fault modes, denoted as set Y;

[0213] Obtain all the identified fault modes and add them to set Y;

[0214] For each fault mode , calculate the similarity with the collected data :

[0215] ;

[0216] Where, and respectively represent the value of the j-th feature in the failure mode and data ; and are respectively the means of the failure mode and data ; is the weight of the j-th feature;

[0217] By calculating the similarity of each failure mode, calculate the priority of each failure mode:

[0218] ;

[0219] wherein, represents the priority of the failure mode ; d is the total number of failure modes; is the similarity between the failure mode and the current data ;

[0220] Sort all the failure modes according to the priority from high to low;

[0221] Output the failure mode with the highest priority and the data with the highest similarity to the failure mode with the highest priority.

[0222] By setting the set Y of identified fault modes and obtaining all fault modes, the problem of how to systematically manage and organize the identified fault modes is solved. In a wireless communication network, the number of fault modes is large and may increase over time. Therefore, storing all identified fault modes centrally in the set Y can help the system effectively track historical fault modes and provide a complete fault mode reference for subsequent fault location. The maintenance and update of this set ensure that the system can continuously optimize and enrich the fault mode library as the network state changes. By constructing a complete set Y of fault modes, the system can flexibly handle new fault types and modes, and at the same time use the existing fault modes for rapid comparative analysis, improving the accuracy and efficiency of fault location. The management of historical fault modes helps reduce the detection and resolution time of future similar problems. By calculating the similarity between each fault mode and the collected data, the problem of how to compare fault modes with real-time collected data to accurately determine the fault type is solved. In an actual network, fault modes are diverse, and the characteristics of each fault mode may vary at different time points or environmental conditions. By calculating the similarity between each fault mode and the collected data, the fault mode most relevant to the current data can be efficiently identified, avoiding simple rule matching or overly rough judgment methods. This step makes the matching between fault modes and real-time data more accurate by quantifying the similarity index. The similarity score between each fault mode and the data provides a clear basis for subsequent priority calculation, further improving the accuracy of fault detection. This method can reduce the false alarm rate and missed alarm rate, and improve the intelligent level of fault location. By calculating the priority of each fault mode, the problem of how to evaluate the severity and urgency of fault modes based on similarity is solved. In a wireless communication network, different fault modes may have different impacts on the network. Some faults may cause major system crashes, while others may be small-scale problems. By calculating the priority of each fault mode, it can be ensured that the fault location system can give priority to faults with greater impact on the network, thus improving the response speed and processing efficiency of network operation and maintenance. This step makes the system able to automatically judge and focus on the most critical faults by quantifying the priority of each fault mode, avoiding waste of resources and redundancy in processing. Through priority sorting, network operation and maintenance personnel can quickly handle high-priority faults, reduce the network interruption time, and improve the availability and stability of the network. By sorting all fault modes from high to low according to priority, the problem of how to sort multiple fault modes and select the most urgent fault for processing is solved. Since multiple fault modes may be detected simultaneously during the fault detection process, how to sort these fault modes and determine the processing order directly affects the efficiency and effect of fault handling. By sorting the fault modes according to priority, the system can ensure that the most critical problems are solved first, thus reducing the impact on network services.The sorting mechanism enables network operation and maintenance to quickly identify the most important faults and respond within the shortest time. In this way, the system can allocate resources more efficiently in the face of multiple faults, improve the overall fault repair efficiency, reduce the network fault recovery time, and optimize the network service quality. By outputting the fault mode with the highest priority and the data with the highest similarity, the problem of how to associate the detected faults with specific data features and further analyze and locate them is solved. After sorting the fault modes, outputting the fault mode with the highest priority and the collected data most similar to it helps to accurately lock the moment and specific problems when the fault occurs, providing a more intuitive basis for fault troubleshooting. This step helps the operation and maintenance personnel quickly locate the cause and background of the fault by outputting the fault mode with the highest priority and the most similar data. Through the precise docking with the data, the operation and maintenance personnel can quickly make fault repair decisions based on specific historical data and similar network conditions, improving the efficiency and accuracy of fault diagnosis.

[0223] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.

[0224] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A fault detection method for a wireless communication network, characterized in that: include: Collecting data from wireless communication networks; The collected data includes: Base station signal strength, denoted by S t , represents the signal strength from the base station to the user equipment at time t, in dBm; Network delay, denoted as D t , represents the communication delay between the user equipment and the base station at time t, in ms; Data packet loss rate, denoted as P t , represents the packet loss ratio of data transmission at time t, expressed as a percentage; The signal quality of the user equipment is denoted by Q t , represents the quality index of the signal received by the user terminal equipment at time t, in dBm; Set the sampling period to T1, the total observation time to T2, and calculate the number of sampling points: Where N is the total number of sampling points; Each sampling generates a data vector, which is recorded as: X t =S t ,D t ,P t ,Q t ,t=1,2,...,N; Among them, X t is the comprehensive data vector at time t, including signal strength, network delay, data packet loss rate and signal quality; Store the sampled data in the database in chronological order; Preprocess the collected data; The preprocessing of the collected data specifically includes: The data in the database are classified and extracted according to base station signal strength, network delay, data packet loss rate and user equipment signal quality, the data of the same type are grouped into a data set, and the classified and extracted data are sorted in the order of storage time; For any kind of dataset, perform the following steps: S1. Use the triple standard deviation method to detect outliers, specifically: Calculate the mean of all values ​​in a dataset: Among them, μ X is the mean of the data set; M is the total number of elements in the data set; X i is the value of the i-th data point in the data set; Calculate the standard deviation: Among them, σ X is the standard deviation of the data set; When the data point X i Satisfy the conditions: |X i -μ X |>3·σ X When the data point X i Mark as outlier; Among them, |X i -μ X | is the data point X i With mean μ X The absolute value of the difference; S2. Use linear interpolation to fill in missing values: When the data point X i When data is missing, linear interpolation is used to fill in the missing values, specifically: Among them, X ’ i is the data value after filling; X i-1 is the data point before the missing value; X i+1 is the data point after the missing value; Perform feature extraction on the preprocessed dataset; The feature extraction in the preprocessed data set specifically includes: Set the sliding window size to k and extract the following features at each time point: S3. Extract sliding mean: Among them, μ t,k is the sliding mean at time t; k is the sliding window size; X r is the rth data point in the window; S4. Extract sliding standard deviation: Among them, σ t,k is the sliding standard deviation at time t; S5. Extract change rate: Among them, Δ t,k is the rate of change at time t; X t is the data point at time t; X t-k is the data point at the starting time of the window; Combine the original data and features to form an extended feature vector: F t =S t ,D t ,P t ,Q t ,m t,k ,s t,k ,D t,k ; Among them, F t is the extended feature vector at time t; Based on the extended feature vector, a fault detection strategy is executed; The fault detection strategy is executed based on the extended feature vector, specifically including: Set the normal operating range for each data type metric: S6、S t ∈-90,-40dBm; S7、D t <100ms; S8、P t <5%; S9、Q t >-70dBm; Set the fault label: Among them, L t is the fault status label at time t; When L t =1, indicating a fault; When L t =0, indicating normal; Compute the similarity of the extended feature vectors: Among them, SimF i ,F j is the extended eigenvector F i With F j Similarity of F i 、F j are the i-th and j-th extended eigenvectors respectively; σ is the Gaussian kernel width parameter; Initialize cluster centers For randomly selected sample points, iterative update: in, is the k-th class center after the t+1th iteration; is the sample set contained in the kth class; is the number of samples in the kth category; All extended feature vectors Randomly assign to K initial clusters, with the cluster center being a randomly selected sample point: in, is the center of the initial cluster k; p is a random index; Calculate the similarity between each sample and all cluster centers and assign the sample to the nearest cluster: in, is the sample set contained in the kth cluster; SimF t ,C k For sample F t To cluster center C k similarity; Update the cluster centers based on the sample set of each cluster: in, is the k-th class center after the t+1th iteration; is the k-th class sample set; is the number of samples in the kth category; Set the clustering convergence threshold to ∈; Check if the cluster center change is below a threshold ∈: If the changes in all cluster centers are below the threshold, the clustering process terminates; The fault classification includes: After clustering is completed, each cluster is mapped to a failure mode or normal state, specifically: Predefined M failure modes Use majority voting to select each cluster Assign fault labels: Among them, L k is the label of cluster k; ⅡL t =m is the indicator function, if L t =m, the value is 1, otherwise the value is 0; When there is a cluster that does not match the existing fault pattern, it is marked as an unknown fault F. u , and record its characteristics; Fault location for detected faults; The fault locating of the detected fault specifically includes: Set the set of identified fault modes, denoted as set Y; Get all identified failure modes and add them to set Y; For each failure mode Y i , calculated and collected data X t The similarity between: Among them, Y ij and X tj Represents failure mode Y i and data X t The value of the jth feature in ; and Failure mode Y i and data X t The mean value of j is the weight of the jth feature; By calculating the similarity of each failure mode, the priority of each failure mode is calculated: Among them, P i Indicates failure mode Y i priority; d is the total number of failure modes; Sim'Y j ,X t Is the failure mode Y j With the current data X t similarity; All failure modes are classified according to priority P i Sort from high to low; Output the highest priority failure mode and the data with the highest similarity to the highest priority failure mode.