Method and device for pre-judging home and guest network faults in advance

By adopting a large-model fault diagnosis method based on the Transformer layer in the fault diagnosis of Jiake network, the problems of low fault diagnosis efficiency and insufficient accuracy in the existing technology are solved, and the accurate positioning and cause analysis of Jiake network faults are realized, which significantly improves network stability.

CN120200899APending Publication Date: 2025-06-24INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510518797.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing fault diagnosis technology of Jiake network is inefficient and it is difficult to comprehensively analyze multi-source heterogeneous data, resulting in low accuracy of fault prediction in advance, which cannot meet the growing demand for stability of Jiake network.

Method used

The fault location and cause analysis of Jiake network is achieved through model construction, data collection and organization, data preprocessing, large-model inference and fault determination and verification processes.

Benefits of technology

It realizes accurate positioning and cause analysis of Jiake network failures, improves the accuracy of fault prediction, significantly improves the stability of Jiake network, and reduces the impact of faults on users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120200899A_ABST
    Figure CN120200899A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of communication networks, and particularly provides a method and device for pre-judging home and guest network faults in advance, and the method comprises the following steps: S1, constructing a model; s2, data collection and arrangement; s3, data preprocessing; s4, performing large model reasoning and fault judgment; and S5, a verification process. Compared with the prior art, the method has the advantages that fault positioning and reason analysis can be accurately carried out, operation and maintenance personnel can quickly take targeted measures to repair faults, the stability of home and guest networks is remarkably improved, and the influence duration of the faults on users is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of communication networks, and specifically provides a method and device for pre-determining home customer network failures in advance. Background Art

[0002] Traditional home customer network fault determination mainly relies on manual experience and simple testing tools. Network maintenance personnel use the ping command to check connectivity and the tracert command to trace the packet path. However, this method is inefficient. In the face of large-scale network failures, manual troubleshooting is time-consuming and laborious. Moreover, simple tools cannot comprehensively integrate complex information such as the performance of each device layer in the access network, broadband bandwidth, and topological structure, resulting in inaccurate pre-judgment and positioning of faults. For example, when the broadband bandwidth is insufficient, it is difficult to quickly determine whether it is caused by network congestion, line aging, or device configuration problems; ordinary network speed test software can only provide surface network speed data and cannot deeply analyze the internal relationship between network performance, devices, and bandwidth.

[0003] Existing AI technologies in network fault diagnosis are mostly based on simple machine learning algorithms such as decision trees and support vector machines. Home customer network data types are diverse, including device performance data, broadband bandwidth information, network speed and performance data, etc., and the data structures and characteristics vary greatly. Simple machine learning algorithms are difficult to effectively integrate and analyze these multi-source heterogeneous data, and cannot fully explore the complex relationships between data, resulting in a low accuracy rate of pre-judging faults in advance and being unable to meet the growing demand for stability of home customer networks. For example, in the face of complex fault scenarios where device performance anomalies and bandwidth fluctuations occur simultaneously, simple algorithms are difficult to accurately determine the cause and location of the fault.

[0004] In recent years, mobile home broadband services have shown a booming development trend, and the number of users has increased exponentially. However, once the broadband network fails, the user experience will be severely affected and the satisfaction will drop significantly.

[0005] In this context, how to ensure the stable operation of the home broadband network and achieve rapid fault diagnosis to guarantee user satisfaction has become a key problem to be solved urgently. Summary of the Invention

[0006] The present invention aims at the above-mentioned deficiencies of the prior art and provides a method for pre-determining home customer network failures in advance with strong practicability.

[0007] The further technical task of the present invention is to provide a device for pre-determining home customer network failures in advance with reasonable design, safety and applicability.

[0008] The technical solution adopted by the present invention to solve its technical problems is:

[0009] A method for pre-determining home customer network failures in advance has the following steps:

[0010] S1, Model construction;

[0011] S2, Data collection and collation;

[0012] S3, Data preprocessing;

[0013] S4, Large model inference and fault determination;

[0014] S5, Verification process.

[0015] Furthermore, in step S1, it includes:

[0016] S1-1, Input layer;

[0017] Organize the access network device performance data into sequences according to device hierarchy and metric type, organize the broadband bandwidth and basic information into sequences, and form corresponding sequences for broadband network speed and performance data;

[0018] S1-2, Transformer layer;

[0019] The input sequence enters the embedding layer, which maps each data element to a high-dimensional vector space;

[0020] S1-3, Output layer;

[0021] The features output by the Transformer layer are mapped to the fault category space through the fully connected layer.

[0022] Furthermore, in step S1-1, the access network TOPO structure information is transformed into the form of an adjacency matrix. Assuming there are N device nodes in the network, for the element A_{ij} of the adjacency matrix A, when device i is connected to device j, A_{ij} = 1, otherwise A_{ij} = 0;

[0023] Expand the adjacency matrix row by row into a one-dimensional sequence, concatenate it with other data sequences as the input of the model. At the same time, encode the topological structure information and integrate the device hierarchy information into the topological structure sequence.

[0024] Furthermore, in step S1-2, the numerical data and the encoded topological structure information are mapped to the same-dimensional vector space through linear transformation, and the word vectors of the text data directly use the pre-trained results;

[0025] The multi-head attention mechanism calculates the attention scores for the input sequence in multiple subspaces in parallel. Assuming there are h attention heads, each head focuses on different parts of the input sequence. For each attention head, calculate the query vector (Query, Q), key vector (Key, K), and value vector (Value, V). The attention score Attention(Q, K, V) calculates the similarity between Q and K through the dot product and weights and sums V. The formula is:

[0026] Attention(Q, K, V) = softmax(\frac{QK^T}{\sqrt{d_k}})V;

[0027] where \(d_k\) is the dimension of the key vector. The multi-head attention mechanism concatenates the results of multiple attention heads and then performs a linear transformation to obtain a fused feature representation, enabling the model to capture data relationships from multiple perspectives;

[0028] The data processed by the multi-head attention mechanism enters the feed-forward neural network FFN. FFN consists of two linear transformations and a ReLU activation function, further extracting and transforming data features. The linear transformation performs a linear combination of the output features of the attention mechanism, and the ReLU activation function increases the non-linear expression ability of the model.

[0029] Furthermore, in step S1-3, assuming there are \(m\) types of fault types in total, the fully connected layer outputs an \(m\)-dimensional vector. Each element of the vector represents the probability score corresponding to the fault type. The weight matrix of the fully connected layer is obtained through training and learning, effectively associating the output features of the Transformer layer with the fault types;

[0030] The probability scores are converted into a probability distribution through the Softmax function to obtain the final fault prediction result;

[0031] The formula of the Softmax function is:

[0032] P(y_i|x) = \frac{e^{z_i}}{\sum_{j = 1}^{m}e^{z_j}};

[0033] where \(z_i\) is the \(i\)-th element of the output vector of the fully connected layer, and P(y_i|x) is the probability that the input data \(x\) belongs to the fault type \(y_i\). The Softmax function converts the output values of the fully connected layer into a probability distribution, and the model outputs the probability of each fault type occurring, which is used for fault location and determination.

[0034] Furthermore, in step S2, various data of the home broadband network are collected in real time, and the collected data are sorted and associated according to the user and device levels to form a complete network data record.

[0035] Furthermore, in step S3, it includes:

[0036] S3-1. Clean the collected data to remove noise and outliers;

[0037] S3-2. Normalize the numerical data. The numerical data is normalized using the min-max normalization method to map the data to the [0, 1] interval, eliminating the dimension difference. The formula is:

[0038] $x_{norm}=\frac{x - x_{min}}{x_{max} - x_{min}}$

[0039] Where $x$ is the original data, $x_{min}$ and $x_{max}$ are the minimum and maximum values in the dataset respectively, and $x_{norm}$ is the normalized data;

[0040] S3-3. Vectorize text data. For text data, use natural language processing techniques for word segmentation and word vector conversion to convert text information into a vector form that the model can understand and process;

[0041] S3-4. Process TOPO structure information. Encode and process the TOPO structure information of the access network and convert it into a format suitable for model input.

[0042] Further, in step S4, the preprocessed data is integrated into an input vector and input into the trained large AI model. The input vector is input into the trained large model, and the model performs inference through a multi-head attention mechanism and a multi-layer neural network to output the probability distribution of different fault types;

[0043] According to a preset threshold, determine whether there is a potential fault risk. If so, output the fault location information and cause analysis, and issue warnings at different levels according to the fault probability value.

[0044] Further, in step S5, it includes:

[0045] S5-1. In model verification, use the trained large model to process the simulated network data, check whether the model can accurately locate the fault location, verify whether the analysis of the fault cause by the model is accurate, and for each injected fault scenario, compare whether the fault cause analysis output by the model is consistent with the actually set fault cause;

[0046] S5-2. In system performance verification, measure the response time of the system from receiving network data to outputting the fault prediction result. Through multiple tests, statistically calculate the average response time and the maximum response time to evaluate whether the system can meet the real-time requirement;

[0047] Calculate the fault prediction accuracy rate of the model in the simulated network environment.

[0048] An apparatus for early prediction and determination of home customer network faults includes: at least one memory and at least one processor;

[0049] The at least one memory is used to store machine-readable programs;

[0050] The at least one processor is used to call the machine-readable program to execute a method for early prediction and determination of home customer network faults.

[0051] Compared with the prior art, a method and device for early pre - determination of home customer network faults according to the present invention have the following outstanding beneficial effects:

[0052] The present invention can accurately locate faults and analyze the causes, enabling operation and maintenance personnel to quickly take targeted measures to repair faults, significantly improving the stability of the home customer network, and reducing the duration of the impact of faults on users. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0054] Attached Figure 1 is a schematic flowchart of a method for early pre - determination of home customer network faults. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] In order to enable those skilled in the art of the present technology to better understand the solutions of the present invention, the following will further elaborate on the present invention in conjunction with specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0056] The following gives a preferred embodiment:

[0057] As Figure 1 shown, a method for early pre - determination of home customer network faults in this embodiment has the following steps:

[0058] S1. Model construction;

[0059] Including:

[0060] S1 - 1. Input layer;

[0061] Organize the access network device performance data into sequences according to the device level and index type, such as [aggregation switch CPU usage rate, aggregation switch memory utilization rate, aggregation switch port traffic, corridor switch port status, corridor switch forwarding error rate, home gateway signal strength, home gateway packet loss rate,...].

[0062] Organize the broadband bandwidth and basic information into sequences, including subscribed bandwidth, actual available bandwidth, user account code, line type code, etc., such as [subscribed bandwidth, actual available bandwidth, user account code, line type code,...].

[0063] The broadband network speed and performance data form corresponding sequences, such as [download speed, upload speed, network latency, jitter, ……].

[0064] The access network TOPO structure information is transformed into the form of an adjacency matrix. Assuming there are N device nodes in the network, for the element A_{ij} of the adjacency matrix A, when device i is connected to device j, A_{ij} = 1, otherwise A_{ij} = 0. The adjacency matrix is expanded into a one-dimensional sequence by rows and concatenated with other data sequences as the input of the model. At the same time, to better reflect the hierarchical relationship of the network topology structure, the topology structure information is encoded, and the device level information (aggregation switch, corridor switch, home gateway) is incorporated into the topology structure sequence. For example, each element in the topology structure sequence is weighted according to the device level to help the model identify the connection relationships of devices at different levels.

[0065] S1-2, Transformer layer;

[0066] The input sequence enters the embedding layer, which maps each data element into a high-dimensional vector space. Numerical data and encoded topology structure information are mapped into the same-dimensional vector space through linear transformation, and the word vectors of text data directly use the pre-trained results. For example, the device performance data is linearly transformed into a 128-dimensional vector space, which is consistent with the dimension of the word vectors of text data for subsequent processing.

[0067] The multi-head attention mechanism calculates the attention scores for the input sequence in multiple subspaces in parallel. Suppose there are h attention heads, and each head focuses on different parts of the input sequence. For each attention head, the query vector (Query, Q), key vector (Key, K), and value vector (Value, V) are calculated. Taking the device performance data sequence as an example, the query vector Q_i is based on the device performance metric vector at the current position, the key vector K_j is the device performance metric vector at other positions, and the value vector V_j is also the device performance metric vector at other positions. The attention score Attention(Q, K, V) calculates the similarity between Q and K through the dot product and weights and sums V. The formula is:

[0068] Attention(Q,K,V) = softmax(\frac{QK^T}{\sqrt{d_k}})V

[0069] where d_k is the dimension of the key vector. The multi-head attention mechanism concatenates the results of multiple attention heads and then performs a linear transformation to obtain a fused feature representation, enabling the model to capture data relationships from multiple perspectives and improving the ability to identify complex fault patterns.

[0070] The data processed by the multi-head attention mechanism enters the feed-forward neural network (FFN). The FFN consists of two linear transformations and a ReLU activation function, which further extracts and transforms the data features. The linear transformation performs a linear combination of the features output by the attention mechanism, and the ReLU activation function increases the non-linear expression ability of the model, enabling the model to learn more complex fault features and providing a more effective feature representation for subsequent fault location and determination.

[0071] S1-3. Output layer;

[0072] The features output by the Transformer layer are mapped to the fault category space through the fully connected layer. Assuming there are m types of fault types in total, the fully connected layer outputs an m-dimensional vector, and each element of the vector represents the probability score corresponding to the fault type. The weight matrix of the fully connected layer is obtained through training to effectively associate the features output by the Transformer layer with the fault types.

[0073] The probability scores are transformed into a probability distribution through the Softmax function to obtain the final fault prediction result. The formula of the Softmax function is:

[0074] P(y_i|x) = \frac{e^{z_i}}{\sum_{j=1}^{m}e^{z_j}}

[0075] where z_i is the i-th element of the output vector of the fully connected layer, and P(y_i|x) is the probability that the input data x belongs to the fault type y_i. The Softmax function transforms the output values of the fully connected layer into a probability distribution, and the model outputs the probability of each fault type occurring, which is used for fault location and determination.

[0076] S2. Data collection and collation;

[0077] Real-time monitoring data collection: Real-time collect various data of the home broadband network through the network monitoring system, including the performance data of access network devices, broadband bandwidth information, broadband network speed and performance data, etc. Continuously update the access network TOPO structure information to ensure the real-time nature of the data.

[0078] Data collation and association: Collate and associate the collected data according to users, device levels, etc. to form a complete network data record. For example, integrate the performance data of the same user's device, broadband bandwidth information, and network speed and performance data for subsequent analysis.

[0079] S3. Data preprocessing;

[0080] Including:

[0081] S3-1. Data cleaning: Clean the collected data to remove noise and outliers. For example, verify and correct data that significantly exceeds the reasonable range in device performance data (such as CPU usage showing 200%) or extremely low-probability extreme values in network speed test data (such as download speed being 0 or far exceeding the subscribed bandwidth). If verification is not possible, mark or delete it to ensure data accuracy and reliability.

[0082] S3-2. Normalization of numerical data: Perform normalization on numerical data using the min-max normalization method to map the data to the [0,1] interval and eliminate the difference in dimensions. The formula is:

[0083] x_{norm}=\frac{x - x_{min}}{x_{max}-x_{min}}

[0084] Where x is the original data, x_{min} and x_{max} are the minimum and maximum values in the dataset respectively, and x_{norm} is the normalized data.

[0085] S3-3. Vectorization of text data: For text data such as user feedback on fault descriptions and device names, use natural language processing techniques for word segmentation and word vector conversion to convert text information into a vector form that the model can understand and process.

[0086] S3-4. Processing of TOPO structure information: Code and process the TOPO structure information of the access network to convert it into a format suitable for model input. For example, represent the topological structure as an adjacency matrix form and incorporate device hierarchy information to enhance the model's understanding of the network topological structure.

[0087] S4. Inference of the large model and fault determination;

[0088] Integrate the preprocessed data into an input vector and input it into the trained AI large model. For example, concatenate the normalized device performance data, encoded topological structure information, and text data word vectors in a specific order into a long vector as the model input.

[0089] Input the input vector into the trained large model. The model performs inference through the multi-head attention mechanism and multi-layer neural network, and outputs the probability distribution of different fault types.

[0090] According to the preset threshold, determine whether there is a potential fault risk. If so, output the fault location information and cause analysis, and issue warnings at different levels according to the fault probability value.

[0091] S5. Verification process;

[0092] Includes:

[0093] S5-1. Model verification;

[0094] Use the trained large model to process the simulated network data and check whether the model can accurately locate the fault location. For example, when injecting fault data into the corridor switch port, observe whether the model can correctly identify the specific corridor switch and port where the fault occurs.

[0095] Verify whether the model's analysis of the fault cause is accurate. For each injected fault scenario, compare whether the fault cause analysis output by the model is consistent with the actually set fault cause. For example, in the scenario of simulating insufficient broadband bandwidth due to network congestion, check whether the model can correctly analyze that network congestion is the cause of insufficient bandwidth.

[0096] S5-2. Measure the response time of the measurement system from receiving network data to outputting the fault prediction result. Through multiple tests, count the average response time and the maximum response time to evaluate whether the system can meet the real-time requirements. For example, the system is required to complete the prediction of common fault scenarios within 1 minute.

[0097] Calculate the fault prediction accuracy rate of the model in the simulated network environment.

[0098] In the data collection and integration link, comprehensive and accurate data collection lays a solid foundation for stability guarantee. By using the SNMP protocol to collect the performance data of each device layer in the access network in real time, abnormal device operations can be detected in a timely manner. For example, when the CPU usage rate of the aggregation switch exceeds 80%, the risk of excessive network load can be pre-warned in advance, and the operation and maintenance personnel can adjust the resource allocation in advance based on this to prevent network failures caused by device overload and ensure the stable operation of the network. The acquisition of broadband bandwidth and basic information enables the operator to handle problems such as network congestion, line attenuation, or improper device configuration in a timely manner based on the difference between the subscribed and actual available bandwidth, maintain the stability of the network bandwidth, and ensure that users enjoy the subscribed bandwidth service. The collection of broadband network speed and performance data, by means of monitoring the network delay change curve, etc., can detect potential congestion or device failures in advance, take measures before the faults occur, avoid network lags, and improve network stability.

[0099] The application of the large model fault location algorithm and inference technology based on the Transformer architecture has greatly enhanced the fault location and cause analysis capabilities. When the large model inference technology analyzes the device performance data, it accurately focuses on the relationships between key indicators such as CPU usage rate and memory utilization rate and faults. When analyzing broadband bandwidth information, it pays attention to the impact of the difference between the subscribed and actual available bandwidths on network performance. When it detects that the forwarding error rate of the corridor switch port increases and the user network speed in the area decreases, combined with the topology structure information, it can quickly determine the link problem between this switch and the aggregation switch, and further analyze factors such as link bandwidth usage and signal strength to comprehensively judge the cause of the fault. This accurate fault location and cause analysis enable the operation and maintenance personnel to quickly take targeted measures to repair the fault, significantly improving the stability of the home customer network and reducing the impact duration of the fault on users.

[0100] When collecting and integrating data, with the help of the SNMP protocol, the performance data of each device layer in the access network is collected in real time. The aggregation switch focuses on collecting the CPU usage rate, memory utilization rate, port traffic, port error rate, etc. These data reflect the operating load of the switch and the accuracy of data transmission. If the CPU usage rate continues to be higher than 80%, it may indicate that the network load is too high. The corridor switch collects the port status, forwarding error rate, and port bandwidth utilization rate to evaluate its working status in the corridor network. The home gateway collects the signal strength, the number of data packets, the packet loss rate, the network connection mode, etc. to understand the user's home network access situation, providing rich device operating status information for early fault prediction.

[0101] Connect to the operator's network management system to obtain the user's broadband bandwidth information, including the subscribed bandwidth, the real-time available bandwidth, as well as basic information such as the user account information, package type, line type, and access mode. By comparing the subscribed and actual available bandwidths, if the actual available bandwidth is much lower than the subscribed bandwidth, it may be caused by network congestion, line attenuation, or device configuration problems, providing multi-dimensional basic data support for early fault prediction.

[0102] Deploy network speed test probes at the user's home terminal or network edge to regularly test the network speed at preset intervals and obtain performance index data such as download speed, upload speed, network latency, and jitter. Combine professional network performance monitoring tools to collect the historical network performance data during the fault occurrence period and analyze the network performance change trend. If the network latency gradually increases, it may indicate network congestion or device failure, and potential network performance problems can be discovered in advance.

[0103] Collect various types of data and converge them to the data convergence server. Conduct preliminary sorting and association according to the user, and sort them in chronological order. Clean the data to remove noise and outliers, ensuring data accuracy and reliability. Perform normalization on numerical data, unifying the value range to the interval [0, 1], eliminating the dimension difference, facilitating subsequent data analysis and model processing, and making the performance data of different devices comparable.

[0104] When constructing the access network TOPO, use a network topology discovery tool based on the ICMP protocol to send ICMP packets to network devices, identify devices and connection relationships according to device responses, and generate a preliminary topology structure. However, in a special configuration network environment, some links may not be able to be normally detected through ICMP packets.

[0105] The network administrator conducts manual inspection and improvement on the automatically generated topology structure according to the actual network deployment. Manually add parts that cannot be automatically recognized, such as special configuration links and cross-region connections, and detail the physical location of devices, port connections, etc. Finally, present the improved topology structure in a graphical way, clearly showing the hierarchical connection relationships of each device, providing an intuitive and accurate network structure reference for early fault prediction, such as marking the specific user home gateway connected to the corridor switch and the link bandwidth between the aggregation switch and the corridor switch.

[0106] Establish a dynamic update mechanism for the topology structure. Regularly (such as at midnight every day) conduct a topology scan of the network, compare the current topology structure with that in the database, and update the topology information in the database in a timely manner when there are changes. When a new user home gateway accesses the network, the topology structure can be updated in a timely manner, accurately determining its location and connection relationship in the network, and providing the latest network structure information for early fault determination.

[0107] When selecting and training the AI large model, select a large language model similar to the GPT-4 architecture, which has strong capabilities in natural language processing and complex data relationship mining, can effectively process and analyze the multi-source heterogeneous data of the home broadband network, and learn complex fault patterns and data association relationships from the vast amount of network data.

[0108] Further preprocess the collected data and perform feature engineering. In addition to normalization for numerical data, extract statistical features such as mean, variance, maximum value, and minimum value. For example, calculate the mean and variance of the device CPU usage rate. A large variance indicates that the CPU usage rate fluctuates violently, which may pose a potential fault risk. Perform word segmentation and word vector conversion on text data. Represent the access network TOPO structure information in the form of an adjacency matrix and encode it, integrating device hierarchical information to enhance the model's understanding of the network topology structure. For example, weight the elements of the adjacency matrix according to the device hierarchy to help the model better identify the impact of the connection relationships of different hierarchical devices on fault propagation.

[0109] Construct a rich training dataset that includes a large number of normal network state samples and various types of fault samples to simulate complex fault scenarios. There are not only single-device fault samples, but also complex samples of multiple devices failing simultaneously and the interaction between device faults and bandwidth problems, enabling the model to learn various fault combination patterns and accurately identify the complex correlation relationships between different network states and fault types.

[0110] Use the training dataset to perform supervised training on the large model, and adopt the cross-entropy loss function to measure the difference between the model prediction and the true fault label. Assume that the model predicts the fault probability distribution as P(y|x), where x is the input network data and y is the true fault label. The calculation formula for the cross-entropy loss function L is:

[0111] L = -\sum_{i = 1}^{n}y_{i}\log(P(y_{i}|x))

[0112] Adjust the model parameters, such as weights and biases, through the backpropagation algorithm to minimize the value of the loss function and improve the accuracy of fault prediction. Adopt an adaptive learning rate algorithm (such as Adagrad, Adadelta) to accelerate the convergence speed and avoid local optimal solutions. During the training process, regularly evaluate the model performance on the validation set, and evaluate metrics such as accuracy and recall every 10 epochs. Stop training when the evaluation metrics do not improve for 5 consecutive times to obtain an optimized model.

[0113] When performing the fault early prediction inference mechanism, when new network data is input, it is processed in the same preprocessing manner as during training. Numerical data is normalized and statistical features are extracted, text data is tokenized and converted into word vectors, and the access network TOPO structure information is encoded. The processed data is integrated into an input vector and input into the trained AI large model. For example, the normalized device performance data, encoded topology structure information, and word vectors of text data are concatenated into a long vector in a specific order as the input.

[0114] The large model uses its powerful inference ability to deeply analyze the input data. The multi-head attention mechanism simultaneously focuses on data features at different positions, automatically learns the correlation weights between data, and captures complex fault patterns. When analyzing device performance data, it focuses on the relationship between key indicators such as CPU usage and memory utilization and network faults; when analyzing broadband bandwidth information, it pays attention to the impact of the difference between the subscribed and actual available bandwidth on network performance. Through a multi-layer neural network, it comprehensively analyzes information such as device performance anomalies, bandwidth changes, and network speed performance metrics in the input data, and combines the access network TOPO structure information to infer the cause and location of the fault. For example, if it is detected that the port forwarding error rate of the corridor switch increases and the network speed of users in the area decreases, combined with the topology structure, it is judged that there may be a problem with the link between this switch and the aggregation switch, and further analyze factors such as link bandwidth usage and signal strength to comprehensively judge the cause of the fault.

[0115] The model outputs the probability distributions of different fault types. A preset threshold (such as 0.8) is set. When the probability value of a certain fault type exceeds the threshold, it is determined as a potential fault risk, and the fault location information and cause analysis are output. For example, it outputs that "the probability of a fault in the corridor switch port of Building XX in XX Community is 0.9, and the possible reason is that the port aging leads to an increase in the forwarding error rate". Multiple thresholds are set to improve the reliability of pre-judgment. A yellow warning is issued when the fault probability is between 0.8 and 0.9, and a red warning is issued when it is greater than 0.9.

[0116] The large model fault location algorithm is based on the Transformer architecture and has obvious advantages in processing sequence data. Network data (such as device performance data, bandwidth information, network speed performance data, etc.) and topology structure information are converted into a sequence form and input into the model. The multi-head attention mechanism in the Transformer architecture automatically learns the correlation weights between data, and at the same time pays attention to the data features at different positions, effectively capturing complex fault patterns. For example, when analyzing the device performance data sequence, it can pay attention to the relationship between multiple device performance indicators and faults at the same time, rather than being limited to a single indicator. During the model training process, the model parameters are adjusted by minimizing the cross-entropy loss function, so that the model prediction results are close to the true fault labels, improving the model's ability to identify and locate faults.

[0117] Based on the above method, a device for early pre-judgment of home customer network faults in this embodiment includes: at least one memory and at least one processor;

[0118] The at least one memory is used to store machine-readable programs;

[0119] The at least one processor is used to call the machine-readable program and execute a method for early pre-judgment of home customer network faults.

[0120] The above specific implementation manners are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific implementation manners. Any technical solution that conforms to the technical solutions recorded in the above specific implementation manners of the present invention and any appropriate changes or substitutions made by those of ordinary skill in the art shall fall within the patent protection scope of the present invention.

[0121] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for early prediction of home-guest network failures, characterized in that: The steps are as follows: S1, model construction; S2, data collection and compilation; S3, data preprocessing; S4, large model reasoning and fault determination; S5. Verification process.

2. According to claim 1, a method for early prediction of home-guest network failures is characterized in that: In step S1, it includes: S1-1, input layer; The access network equipment performance data is sorted into a sequence according to the equipment level and indicator type, the broadband bandwidth and basic information are sorted into a sequence, and the broadband network speed and performance data form a corresponding sequence; S1-2, Transformer layer; The input sequence enters the embedding layer, which maps each data element into a high-dimensional vector space; S1-3, output layer; The output features of the Transformer layer are mapped to the fault category space through the fully connected layer.

3. According to claim 2, a method for early prediction of home-guest network failures is characterized in that: In step S1-1, the access network TOPO structure information is converted into an adjacency matrix form. Assuming that there are N device nodes in the network, the element A_{ij} of the adjacency matrix A, when device i is connected to device j, A_{ij}=1, otherwise A_{ij}=0; The adjacency matrix is ​​expanded into a one-dimensional sequence by row, and then concatenated with other data sequences as the input of the model. At the same time, the topological structure information is encoded and the device level information is integrated into the topological structure sequence.

4. According to claim 3, a method for early prediction of home-guest network failures is characterized in that: In step S1-2, the numerical data and the encoded topological structure information are mapped to the same dimensional vector space through linear transformation, and the word vector of the text data directly uses the pre-training result; The multi-head attention mechanism calculates the attention score of the input sequence in multiple subspaces in parallel. Assume that there are h attention heads, each of which focuses on a different part of the input sequence. For each attention head, the query vector (Query, Q), key vector (Key, K) and value vector (Value, V) are calculated. The attention score Attention(Q, K, V) calculates the similarity between Q and K through the dot product and weighted sums V. The formula is: Attention(Q,K,V)=softmax(\frac{QK^T}{\sqrt{d_k}})V; Where d_k is the key vector dimension. The multi-head attention mechanism concatenates the results of multiple attention heads and transforms them linearly to obtain a fused feature representation, allowing the model to capture data relationships from multiple angles. The data processed by the multi-head attention mechanism enters the feed-forward neural network FFN, which consists of two linear transformations and a ReLU activation function to further extract and transform data features. The linear transformation linearly combines the output features of the attention mechanism, and the ReLU activation function increases the nonlinear expression ability of the model.

5. A method for predicting home-guest network failure in advance according to claim 4, characterized in that: In step S1-3, assuming that there are m types of faults, the fully connected layer outputs an m-dimensional vector, each element of which represents the probability score of the corresponding fault type. The weight matrix of the fully connected layer is obtained through training and learning, which effectively associates the output features of the Transformer layer with the fault type. The probability score is converted into probability distribution through the Softmax function to obtain the final fault prediction result; The Softmax function formula is: P(y_i|x)=\frac{e^{z_i}}{\sum_{j=1}^{m}e^{z_j}}; Where z_i is the i-th element of the fully connected layer output vector, P(y_i|x) is the probability that the input data x belongs to the fault type y_i, and the Softmax function is used to convert the output value of the fully connected layer into a probability distribution. The model outputs the probability of each fault type for fault location and judgment.

6. A method for predicting home-guest network failure in advance according to claim 5, characterized in that: In step S2, various data of the home guest network are collected in real time, and the collected data are sorted and associated according to the user and device levels to form a complete network data record.

7. A method for predicting home-guest network failure in advance according to claim 6, characterized in that: In step S3, it includes: S3-1, clean the collected data to remove noise and outliers; S3-2, normalization of numerical data, normalization of numerical data, using the minimum-maximum normalization method, mapping the data to the [0,1] interval to eliminate dimensional differences, the formula is: x_{norm}=\frac{x-x_{min}}{x_{max}-x_{min}} Where x is the original data, x_{min} and x_{max} are the minimum and maximum values ​​in the data set, respectively, and x_{norm} is the normalized data; S3-3, vectorization of text data: for text data, natural language processing technology is used to segment words and convert word vectors to convert text information into a vector form that the model can understand and process; S3-4, TOPO structure information processing, encoding and processing of the access network TOPO structure information, and converting it into a format suitable for model input.

8. A method for predicting home-guest network failure in advance according to claim 7, characterized in that: In step S4, the preprocessed data is integrated into an input vector and input into the trained AI big model. The input vector is input into the trained big model, and the model performs reasoning through a multi-head attention mechanism and a multi-layer neural network to output the probability distribution of different fault types; Based on the preset threshold, determine whether there is a potential failure risk. If so, output the fault location information and cause analysis, and issue different levels of warnings based on the failure probability value.

9. A method for early prediction of home-guest network failure according to claim 8, characterized in that: In step S5, it includes: S5-1. In model verification, the trained large model is used to process the simulated network data to check whether the model can accurately locate the fault location and verify whether the model's analysis of the fault cause is accurate. For each injected fault scenario, the fault cause analysis output by the model is compared with the actual set fault cause to see if they are consistent; S5-2, in the system performance verification, measure the system response time from receiving network data to outputting fault prediction results. Through multiple tests, calculate the average response time and maximum response time to evaluate whether the system can meet the real-time requirements; Calculate the fault prediction accuracy of the model in a simulated network environment.

10. A device for predicting home-guest network failures in advance, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is configured to call the machine-readable program to execute the method according to any one of claims 1 to 9.