Cloud-based collaborative cpe device remote operation and fault prediction method
By constructing a cloud-based collaborative heterogeneous correlation graph and performing root cause analysis, the problem of difficulty in identifying the source of CPE equipment failure was solved, achieving high-precision fault prediction and automated recovery, and reducing the false positive rate and repair time.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- N RADIO TECH CO LTD
- Filing Date
- 2026-06-11
- Publication Date
- 2026-07-10
AI Technical Summary
Existing technologies cannot effectively distinguish whether a CPE device failure originates from local hardware anomalies, base station sector beam offset, or transmission link capacity degradation. This results in low fault prediction accuracy, high false alarm and missed alarm rates, and the automatic recovery phase may incorrectly determine the root cause, thus prolonging the repair time.
Construct a heterogeneous correlation graph based on cloud collaboration, identify device clusters through multi-dimensional operational indicator data, perform root cause analysis by combining network topology data, generate a pre-diagnostic instruction set, and realize automated recovery operations.
Accurately identify faulty equipment clusters, improve fault prediction accuracy, reduce false positive rate, realize end-to-end automated recovery closed loop, and shorten fault repair time.
Smart Images

Figure CN122372447A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of fault operation and maintenance prediction technology, and in particular to a method for remote operation and maintenance and fault prediction of CPE devices based on cloud collaboration. Background Technology
[0002] With the large-scale deployment of 5G networks and the rapid growth in demand for gigabit broadband access, fixed wireless access solutions based on 5G CPE terminal devices have become an important alternative to broadband services for homes and small and medium-sized enterprises. However, CPE devices are usually deployed on the user side, lacking a professional operation and maintenance environment, and their operating status is highly dependent on multi-dimensional dynamic factors such as wireless channel quality, base station resource scheduling, and backhaul link stability, resulting in complex causes of failures and difficulties in locating them.
[0003] In existing remote operation and maintenance systems, traditional methods mainly rely on single-device threshold alarms or simple rule-based correlation analysis. For example, when a CPE reports a signal strength (RSRP) below -110 dBm or a throughput drop of more than 50%, the system triggers a weak coverage or performance degradation alarm. While these methods are simple to implement, they have significant drawbacks: on the one hand, they cannot distinguish whether the fault originates from a CPE's local hardware anomaly, base station sector beam offset, or transmission link capacity degradation; on the other hand, since multiple CPEs share the same sector or backhaul link in a 5G network, individual alarms often have strong correlations. Without aggregation and root cause analysis, alarm storms can easily occur, leading to a waste of operation and maintenance resources.
[0004] A deeper problem lies in the fact that existing methods fail to construct quantifiable and propagable degradation flow signals to characterize the propagation path of anomalies in the topology. In real networks, anomalies exhibit cascading propagation characteristics: for example, anomalies on the wireless side initially lead to performance degradation of the downstream CPEs, and these degradation indicators further attempt to propagate towards the core network through the backhaul link; however, if a backhaul link itself has a high bit error rate, congestion, or bandwidth bottleneck, it may block the continued propagation of the degradation signal, causing the anomaly energy to accumulate at the link entry point, forming a local convergence point, which is very likely the real source of the fault. However, current technology has neither defined a unified degradation flow metric nor established a modulation mechanism for anomaly propagation based on link state, resulting in the system's inability to effectively distinguish between source faults and passively affected devices, thus misjudging many problems that should be attributed to upstream infrastructure as terminal-side problems.
[0005] In summary, during the fault prediction phase, it is difficult to accurately identify device clusters with common degradation patterns and the actual faulty nodes behind them from a massive number of CPEs, resulting in low prediction accuracy and high false negative and false positive rates. During the automated recovery phase, due to the coarse determination of root cause types, the system may issue incorrect pre-diagnostic instructions, which not only fail to solve the problem but may also interrupt user services, or even mask the true fault phenomenon, thus prolonging the mean time to repair (MTBL). Summary of the Invention
[0006] This application aims to at least partially address one of the technical problems in the related art.
[0007] To achieve the above objectives, this application proposes a cloud-based collaborative method for remote operation and maintenance and fault prediction of CPE devices, including the following steps:
[0008] Step 1: Collect operational indicator data from multiple target devices. The operational indicator data includes signal quality indicators, throughput indicators, resource utilization indicators, and corresponding sampling timestamps. Determine a time window based on the sampling timestamps.
[0009] Step 2: Obtain network topology data related to the target device from the network management system. The network topology data includes the base station identifiers accessed by each target device and the transmission link identifiers on which the base stations depend.
[0010] Step 3: Based on the target device, operational indicator data, base station identifier, and transmission link identifier as nodes, construct association edges between the nodes, and construct a heterogeneous association graph based on the nodes and association edges;
[0011] Step 4: Construct multiple degradation indicators based on the operational indicator data within a continuous time window; identify target base station nodes as predicted fault points in the heterogeneous association graph based on the degradation indicators; and identify target device nodes directly connected to the predicted fault points as device clusters.
[0012] Step 5: Perform root cause analysis on the device cluster to obtain the root cause ranking result; perform composite judgment logic on the top n target devices in the root cause ranking result to determine the fault type, generate a pre-diagnosis instruction set, and perform automated recovery operation.
[0013] Based on the target device, operational indicator data, base station identifier, and transmission link identifier as nodes, association edges are constructed between the nodes, and a heterogeneous association graph is constructed based on the nodes and association edges, including the following steps:
[0014] Step 31: Define each target device as a target device node, define each type of operational indicator data as an indicator data node, define each base station identifier as a base station node, and define each transmission link identifier as a transmission link node.
[0015] Step 32: Construct a first association edge based on the target device node and the base station node; construct a second association edge based on the base station node and the transmission link node;
[0016] Step 33: Construct a third association edge based on the target device node and the indicator data node; construct a fourth association edge based on the base station node and the indicator data node; and construct a fifth association edge based on the relationship between the transmission link node and the indicator data node.
[0017] Step 34: Based on the target device node, indicator data node, base station node, transmission link node, and the first, second, third, fourth, and fifth associated edges connecting the nodes, a heterogeneous association graph is constructed.
[0018] Multiple degradation indicators are constructed based on the operational indicator data within a continuous time window. Target base station nodes are identified as predicted fault points in the heterogeneous correlation graph based on these degradation indicators. Target device nodes directly connected to the predicted fault points are designated as a device cluster. The process includes the following steps:
[0019] Step 41: Based on each target device node, obtain the operation index data within a continuous time window, calculate the first-order time difference for each type of operation index data as the degradation rate of each operation index data; multiply the degradation rate of all operation index data of the current target device node by a preset weight to obtain the degradation rate vector of the current target device node.
[0020] Step 42: Calculate the mean of all elements in the degradation rate vector and record it as the scalar degradation intensity;
[0021] Step 43: Based on the target device nodes connected to each base station node, sector groups are formed in conjunction with antenna sector configuration information. A weighted average is then performed based on the degradation rate vector of the target device nodes within each sector group to obtain a sector-level degradation representative vector. The sector-level degradation representative vectors of all sector groups are then concatenated in order to form a structured degradation state vector.
[0022] Step 44: Perform sparse principal component analysis based on the structured deterioration state vector to obtain the principal mode direction vector;
[0023] Step 45: Obtain multiple state parameter time series of each base station node within multiple consecutive time windows and filter the state master parameters; calculate the first-order time difference based on the state master parameter time series corresponding to the state master parameters and record it as the base station degradation rate; construct a multivariate time series matrix based on the base station degradation rate and the scalar degradation intensity of each target device node, wherein each row of the multivariate time series matrix corresponds to a time point, the first column of the multivariate time series matrix is the base station degradation rate at the current time point, and the remaining columns in the multivariate time series matrix are the scalar degradation intensity of each target device node at the current time point;
[0024] Step 46: Calculate the Granger causality strength and reverse causality strength based on the base station degradation rate and each scalar degradation strength of the multivariate time series matrix; if the average Granger causality strength of the base station degradation rate for at least two target device nodes is greater than the reverse causality strength, then calculate the inner product of the main mode direction vector and the base station structured degradation state vector under the current time window to obtain the degradation main mode activation strength; use the degradation main mode activation strength as a degradation flow signal and send it from the base station node to each connected transmission link node;
[0025] Step 47: Construct a link state feature vector based on the transmission link node; when the current degraded flow signal arrives at the transmission link node, concatenate the current degraded main mode activation intensity with the current link state feature vector to form a current joint feature vector; calculate the Mahalanobis distance between the current joint feature vector and each historical event in the preset memory bank to obtain multiple Mahalanobis distance values; if the minimum Mahalanobis distance exceeds the dynamic tolerance, determine that the current transmission link node is in an abnormal conduction state, block the continued downstream propagation of the degraded flow signal, and generate a propagation blocking binary marker;
[0026] Step 48: For each base station node, calculate the proportion of transmission link nodes whose propagation of degraded flow signals is blocked to the total number of transmission link nodes; if the proportion exceeds a preset proportion threshold, it is determined that the degraded energy has locally converged at the current base station; and calculate the cosine distance between the main mode direction vector of the current base station and the main mode direction vector of the adjacent base stations; if the cosine distance is greater than the isolation threshold, it is determined that the degraded mode of the base station has topological isolation; mark the base station nodes that simultaneously satisfy the local convergence condition and the topological isolation condition as predicted fault points, and aggregate all directly connected target device nodes into a cluster of devices to be dealt with.
[0027] The link state feature vector includes real-time bit error rate, real-time throughput, real-time transmission latency, and real-time load rate.
[0028] The dynamic tolerance is obtained based on the real-time bit error rate, real-time load rate, and preset basic tolerance: if the real-time load rate is less than 0.3 and the real-time bit error rate is less than 0.02, the dynamic tolerance is defined as the preset basic tolerance; if the real-time load rate is greater than 0.7 or the real-time bit error rate is greater than 0.1, the dynamic tolerance is defined as x times the preset basic tolerance, where x is between 0.1 and 0.3; otherwise, the dynamic tolerance is calculated based on the real-time bit error rate, real-time load rate, and preset basic tolerance.
[0029] Root cause analysis is performed on the device cluster to obtain root cause ranking results; based on the top n target devices in the root cause ranking results, composite judgment logic is executed to determine the fault type, and a pre-diagnostic instruction set is generated to execute automated recovery operations, including the following steps:
[0030] Step 51: Mark each target device node in the device cluster as a candidate device node, mark the base station node connected to the candidate device node as a candidate base station node, and mark the propagation link node connected to the candidate base station node as a candidate link node; construct base station sector nodes based on the candidate device nodes connected to the candidate base station nodes; form a first type of edge based on the candidate device node pointing to its corresponding base station sector node; form a second type of edge based on the candidate device node pointing to the candidate link node it uses; form a third type of edge based on the candidate link node pointing to its corresponding base station sector node; construct a heterogeneous operation and maintenance graph based on the candidate device nodes, base station sector nodes, candidate link nodes, first type of edge, second type of edge, and third type of edge.
[0031] Step 52: Define a first initial anomaly score for each candidate device node based on the heterogeneous operation and maintenance diagram, and define a second initial anomaly score for each base station sector node and candidate link node, wherein the second initial anomaly score is 0;
[0032] Step 53: Iteratively perform deterministic anomaly propagation on the heterogeneous operation and maintenance graph. In each round, update the results based on the first initial anomaly score of each candidate device node to obtain the root cause score of each candidate device node, its associated base station sector node, and candidate link node.
[0033] Step 54: Sort all candidate device nodes in descending order of root cause score to generate root cause ranking result; if the root cause score of the base station sector node to which the first n candidate device nodes belong in the root cause ranking result is greater than the first sector threshold, and the proportion of the number of target devices whose propagation link nodes within the base station sector node to the total number of target devices in the sector group is greater than the preset density threshold, then the fault type is determined to be antenna beamforming offset fault type.
[0034] Step 55: If the root cause score of the transmission link node used by the top n candidate device nodes in the root cause sorting result is greater than the first link threshold, and the dynamic tolerance of the current transmission link node shows a monotonically decreasing trend in the most recent P consecutive time windows, then the fault type is determined to be the link carrying capacity degradation fault type.
[0035] Step 56: If the root cause scores of the top n candidate device nodes in the root cause sorting results are greater than the preset root cause significance threshold, the root cause scores of the base station sector nodes to which they belong are less than or equal to the first sector threshold, and the root cause scores of the transmission link nodes used are less than or equal to the first link threshold, then the fault type is determined to be the target device local hardware abnormality fault type.
[0036] Step 57, otherwise, determine the fault type as a composite abnormal fault type;
[0037] Step 58: Match the predefined fault type template according to the determined fault type, generate the corresponding pre-diagnostic instruction set and execute it to perform automated recovery operation on the target device.
[0038] Iteratively perform deterministic anomaly propagation on the heterogeneous operation and maintenance graph. In each round, update the results based on the first initial anomaly score of each candidate device node to obtain the root cause score of each candidate device node, as well as its associated base station sector node and transmission link node. This includes the following steps:
[0039] Step 531: In each round, the initial anomaly scores of the candidate device nodes, base station sector nodes, and candidate link nodes in the heterogeneous operation and maintenance graph are iteratively updated:
[0040] a. Based on the base station sector node, obtain the candidate device node and candidate link node of the incoming edge, denoted as the first incoming edge and the second incoming edge, respectively; determine the sum of the initial anomaly scores of all the first incoming edges of the base station sector node and multiply them by the first propagation coefficient, denoted as the first anomaly score; determine the sum of the initial anomaly scores of all the second incoming edges of the base station sector node and multiply them by the second propagation coefficient, denoted as the second anomaly score; calculate the sum of the first anomaly score and the second anomaly score as the base station sector anomaly score;
[0041] b. Based on the candidate link node, obtain the candidate device node of the incoming edge, which is denoted as the third incoming edge. Calculate the sum of the initial anomaly scores of all third incoming edges and multiply it by the third propagation coefficient to obtain the third anomaly score, which is used as the anomaly score of the candidate link node.
[0042] Step 532: If the number of propagation rounds already executed reaches the preset maximum number of rounds, and the maximum absolute change in the abnormal scores of all updated base station sector nodes and candidate link nodes in two consecutive propagation rounds is less than the preset convergence threshold, then stop the iterative propagation and proceed to the next step:
[0043] Step 533: Calculate the root cause score for each target device based on the updated anomaly scores of the base station sector nodes, candidate link nodes, and the first initial anomaly score.
[0044] The first initial score is obtained by weighted summation of the scalar degradation intensity and the binary label of the current candidate device node.
[0045] The first propagation coefficient, the second propagation coefficient, and the third propagation coefficient satisfy the following condition: 0 < second propagation coefficient < first propagation coefficient < third propagation coefficient < 1.
[0046] Compared with existing technologies, the cloud-based collaborative remote operation and maintenance and fault prediction method for CPE devices provided in this application constructs a heterogeneous association graph based on three types of entities: the target CPE device, its serving base station, and the transmission link it depends on. This graph structure can explicitly identify whether multiple CPEs share the same base station or transmission link, thereby capturing base station-level or link-level degradation patterns. When multiple CPEs experience performance degradation simultaneously, if they converge at the same upstream node in the graph, it indicates a local abnormal energy convergence phenomenon, and this convergence point is the potential source of the fault. Based on this, the device clusters affected by the same infrastructure disturbance are accurately selected as the analysis objects for subsequent fault prediction, effectively avoiding misattributing common-cause degradation to multiple isolated terminal failures or vaguely judging it as an overall base station anomaly. It shifts from phenomenon aggregation to structural tracing, significantly narrowing down the root cause candidate set and avoiding misjudging common-cause degradation as multiple independent terminal failures or vaguely attributing it to an overall base station anomaly, fundamentally eliminating root cause ambiguity.
[0047] Subsequently, a heterogeneous operation and maintenance graph was constructed for the device cluster, and each node in the graph was assigned an initial score. An iterative propagation and update mechanism was implemented: the node score was transmitted to neighboring nodes along the graph edges according to preset rules, and in each iteration, the signals received from the neighbors were merged to continuously correct its own score. After several rounds of convergence, each node obtained a stable root cause score, and the upstream node with the most significant convergence of abnormal energy would show the highest root cause score, thereby improving the accuracy of locating the real fault point.
[0048] Based on this, a composite judgment logic is executed: if the highest root cause score is located at a CPE node and there is no commonality, it is judged as a local hardware anomaly; if it is located at a base station node and its downstream CPEs show common degradation, it is judged as a wireless-side anomaly; if it is located at a transmission link node and the link status deteriorates synchronously, it is judged as a transmission link anomaly. Finally, according to the determined fault type, the corresponding recovery command is automatically generated (such as restarting the CPE, switching the wireless frequency band, triggering link rerouting, etc.), and distributed to edge devices or network controllers through the cloud management platform to achieve end-to-end automated recovery closed loop. By deeply integrating the three information of topology location, root cause score distribution, and real-time link status, fine-grained differentiation of fault types is achieved, and the false positive rate is significantly reduced. Attached Figure Description
[0049] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0050] Figure 1 A flowchart illustrating the cloud-based collaborative remote operation and maintenance and fault prediction method for CPE devices provided in this application embodiment;
[0051] Figure 2 This is a structural block diagram of the cloud-based collaborative remote operation and maintenance and fault prediction system for CPE devices provided in the embodiments of this application;
[0052] Figure 3 This is a block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0053] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0054] The following describes a cloud-based collaborative method for remote operation and maintenance and fault prediction of CPE devices according to embodiments of this application, with reference to the accompanying drawings.
[0055] like Figure 1 As shown, the cloud-based collaborative remote operation and maintenance and fault prediction method for CPE devices includes the following steps:
[0056] Step 1: Collect operational indicator data based on multiple target devices. The operational indicator data includes signal quality indicators, throughput indicators, resource utilization indicators, and corresponding sampling timestamps. Determine the time window based on the sampling timestamps.
[0057] This embodiment targets multiple 5G CPE target devices deployed at the edge. A lightweight local proxy module collects local operational metrics data at preset intervals (e.g., every 30 seconds), including signal quality metrics (e.g., RSRP, SINR, BLER), throughput metrics (e.g., downlink / uplink throughput, TCP retransmission rate), and resource utilization metrics (e.g., CPU utilization, memory usage, device temperature). A local high-precision sampling timestamp is appended to each record. Subsequently, this raw metric data is uploaded to the cloud-based operations and maintenance platform in real time via a secure channel (e.g., TLS encryption), achieving synergy between edge awareness and cloud aggregation. Upon receiving data reported from multiple devices, the cloud platform calibrates and aligns the sampling timestamps of each device based on a unified clock source, eliminating time deviations caused by device local clock drift. Furthermore, the cloud dynamically configures sliding time window parameters (e.g., window length of 15 minutes, sliding step of 5 minutes) according to business needs and anomaly detection sensitivity, and divides the data of all target devices into corresponding time windows based on the calibrated sampling timestamps. Data within each time window is aggregated by device ID, forming a structured, time-aligned multidimensional metric sequence.
[0058] Step 2: Obtain network topology data related to the target device from the network management system. The network topology data includes the base station identifiers accessed by each target device and the transmission link identifiers on which the base stations depend.
[0059] The cloud-based operations and maintenance platform proactively initiates a structured query request to the operator's network management system, using the unique device identifier of each target device as input. Based on its maintained real-time configuration database, the network management system returns the identifier of the base station currently registered or served by each target device, along with the identifier of the backhaul transmission link bound to that base station. This topology information is encapsulated by the network management system according to a standard interface protocol and synchronized to the cloud-based operations and maintenance platform via an authentication and authorization channel. Upon receiving the raw topology data, the cloud side performs format parsing and consistency verification, eliminating transient inconsistencies caused by handover processes and retaining only stable access relationships. Ultimately, each target device is precisely mapped to its currently serving base station, and each base station is associated with one or more primary transmission links it relies on, forming a three-layer topology mapping relationship of "target device—base station—transmission link".
[0060] Step 3: Based on the target device, operational indicator data, base station identifier, and transmission link identifier as nodes, construct association edges between the nodes, and construct a heterogeneous association graph based on the nodes and association edges;
[0061] Step 31: Define each target device as a target device node, define each type of operational indicator data as an indicator data node, define each base station identifier as a base station node, and define each transmission link identifier as a transmission link node.
[0062] The cloud-based operations and maintenance platform performs structured abstraction of all entities and attributes. For each target device with valid operational indicator data within a time window, a uniquely identified target device node is created, whose node attributes include device ID, geographical location, and user information. For each type of operational indicator data, a corresponding indicator data node is created. Each type of indicator data node exists independently and carries the indicator type name, unit, and normal threshold range as metadata. At the same time, based on the topology mapping relationship obtained in step S2, each appearing base station identifier is instantiated as a base station node, whose node attributes include base station ID, frequency band configuration, and coverage area. Similarly, each referenced transmission link identifier is instantiated as a transmission link node, whose attributes include link ID, bandwidth capacity, current bit error rate, and QoS level.
[0063] Step 32: Construct a first association edge based on the target device node and the base station node; construct a second association edge based on the base station node and the transmission link node;
[0064] Step 33: Construct a third association edge based on the target device node and the indicator data node; construct a fourth association edge based on the base station node and the indicator data node; and construct a fifth association edge based on the relationship between the transmission link node and the indicator data node.
[0065] Step 34: Based on the target device node, indicator data node, base station node, transmission link node, and the first, second, third, fourth, and fifth associated edges connecting the nodes, a heterogeneous association graph is constructed.
[0066] For each target device node and its currently connected base station node, a first association edge is created. This edge is labeled as "access" and carries a time window identifier to reflect dynamic association. For each base station node and its dependent transmission link node, a second association edge is created. This edge is labeled as "bearer" and includes uplink / downlink direction attributes to reflect the link usage pattern. Subsequently, combining the operational indicator data aggregated by time window in step S1, the platform further establishes semantic associations between indicators and entities: For each target device node, the three types of operational indicators collected within the corresponding time window are linked to the corresponding indicator data nodes, forming a third association edge, with the edge type being "generation"; for a base station node, if multiple target devices under its service exhibit common signal quality degradation or throughput decline within the same time window, the base station node is connected to the relevant indicator data node through a fourth association edge, with the edge type being "influence," used to characterize the aggregation effect of base station-level performance on indicators; similarly, if a transmission link experiences high bit error rate or congestion within a specific time window, and this occurs synchronously with abnormal resource utilization of multiple base stations or target devices, a fifth association edge is established between the transmission link node and the corresponding indicator data node, with the edge type being "constraint." All nodes and the five types of association edges are ultimately uniformly loaded into the graph database, forming a heterogeneous association graph with multi-hop transmission paths, multi-granularity semantic relationships, and temporal context constraints. This graph not only fully preserves the physical network topology but also integrates the dynamic distribution characteristics of operational status indicators.
[0067] Step 4: Construct multiple degradation indicators based on the operational indicator data within a continuous time window; identify target base station nodes as predicted fault points in the heterogeneous association graph based on the degradation indicators; and identify target device nodes directly connected to the predicted fault points as device clusters.
[0068] Step 41: Based on each target device node, obtain the operation index data within a continuous time window, calculate the first-order time difference for each type of operation index data as the degradation rate of each operation index data; multiply the degradation rate of all operation index data of the current target device node by a preset weight to obtain the degradation rate vector of the current target device node.
[0069] First, for each target device node in the diagram, the cloud-based operations and maintenance platform traces back through multiple consecutive time windows (e.g., the most recent 5 sliding windows, each 5 minutes apart, covering 25 minutes of history) already defined in step S1, extracting the three types of operational indicator data corresponding to that node within each window. For each type of indicator, the platform arranges its numerical sequence in chronological order and performs a first-order time difference operation on the indicator values between adjacent time windows, i.e., subtracting the previous window value from the current window value, and then normalizing the difference to obtain the change in the indicator within that time period; if the difference is negative and exceeds the normal fluctuation threshold (e.g., SINR drops by more than 3dB), it is considered a degradation trend, and the original difference is retained as the degradation rate; if it is positive or in a stable range, the degradation rate is set to zero or compressed according to a decay function to suppress interference in non-degradation directions. Subsequently, business awareness weights are pre-configured for each type of metric (e.g., signal quality weight 0.5, throughput weight 0.3, resource utilization weight 0.2, with a total weight of 1, determined by operation and maintenance strategies or historical fault statistics). The degradation rate of each type of metric is then multiplied by its corresponding weight to form a weighted degradation component. Finally, all weighted components are concatenated in order of metric category to form a fixed-dimensional degradation rate vector.
[0070] Step 42: Calculate the mean of all elements in the degradation rate vector and record it as the scalar degradation intensity.
[0071] The degradation rate vector is a one-dimensional array whose dimension is equal to the number of running metric categories (e.g., three categories of metrics correspond to a three-dimensional vector). Each element of the vector already contains a weighted degradation rate value. The platform iterates through all elements of the vector, sums them algebraically, and divides the sum by the number of vector dimensions. This arithmetic average is used to calculate the overall average level, and the result is defined as the scalar degradation intensity of the target device node in the current time window.
[0072] Step 43: Based on the target device nodes connected to each base station node, sector groups are formed in conjunction with antenna sector configuration information. A weighted average is then performed based on the degradation rate vector of the target device nodes within each sector group to obtain a sector-level degradation representative vector. The sector-level degradation representative vectors of all sector groups are then concatenated in order to form a structured degradation state vector.
[0073] First, for each base station node in the diagram, the cloud-based operations and maintenance platform extracts all target device nodes directly connected to it through the first associated edge. Simultaneously, it retrieves the antenna sector configuration information of the base station from the network configuration database or base station metadata, including the number of sectors (usually three, corresponding to 0°, 120°, and 240° directions), the coverage angle range of each sector, and a unique sector identifier. Based on the geographical location information of each target device node (already recorded in the target device node attributes) and the beam pointing and coverage area of each sector, the platform precisely maps the connected target device nodes to the corresponding sectors, thereby completing sector grouping and ensuring that each target device node belongs to only one sector.
[0074] Next, for the target device node set within each sector group, the platform performs weighted average aggregation: the weights are dynamically set based on the historical service activity or current throughput of each target device node (for example, if there are 3 target device nodes in a sector with downlink throughputs of 20 Mbps, 30 Mbps, and 50 Mbps respectively, the total throughput is 100 Mbps, corresponding to weights of 0.2, 0.3, and 0.5 respectively), to highlight the impact of high-load users on sector performance; if no dynamic weights are available, equal weights are used by default. Specifically, the degradation rate vectors of all target device nodes in the same sector are weighted and averaged dimension by dimension to generate a new vector with the same dimension as the original degradation rate vector, which is the sector-level degradation representative vector for that sector. This vector effectively summarizes the comprehensive degradation trend of the user group under that sector, possessing spatial locality and service representativeness.
[0075] Finally, for all sectors (e.g., 3) under the jurisdiction of the base station node, the platform concatenates the sector-level degradation representative vectors of each sector in a preset sector numbering order to form a one-dimensional vector with a length of sector number × indicator category number, denoted as the structured degradation state vector of the base station node. For example, if the number of sectors is 3 and the number of indicator categories is 3, then the vector length is 9, with a clear structure and preservation of inter-sector differences.
[0076] Step 44: Perform sparse principal component analysis based on the structured deterioration state vector to obtain the principal mode direction vector;
[0077] The cloud-based operation and maintenance platform organizes the structured degradation state vectors corresponding to all base station nodes in the current network into a matrix, where each row is the structured degradation state vector of a base station, and the number of columns is equal to the product of the number of sectors and the number of operational indicator categories (for example, if there are 3 sectors and 3 types of indicators, then the length of each vector is 9).
[0078] Subsequently, the platform performs sparse principal component analysis on the matrix. This process employs standard numerical computation methods, retaining the ability of traditional principal component analysis to extract the direction of maximum variance while introducing sparsity constraints. This ensures that the resulting principal component direction vectors retain only the dimensions that significantly contribute to the global degradation pattern, while the coefficients of other dimensions are compressed to zero or near zero. In practice, a mature numerical computation library can be called, the number of principal components can be set to 1, and appropriate sparsity control parameters can be configured (e.g., setting the L1 regularization strength based on experience with the vector dimensions) to solve for the direction vector corresponding to the first principal component.
[0079] The obtained first principal component direction vector is the main mode direction vector. This vector has the same dimension as the original structured degradation state vector, and its non-zero elements clearly indicate the specific sectors and index combinations that commonly exhibit significant degradation across the network, thus revealing the main common patterns of network degradation in the current period.
[0080] Step 45: Obtain multiple state parameter time series of each base station node within multiple consecutive time windows and filter the state master parameters; calculate the first-order time difference based on the state master parameter time series corresponding to the state master parameters and record it as the base station degradation rate; construct a multivariate time series matrix based on the base station degradation rate and the scalar degradation intensity of each target device node, wherein each row of the multivariate time series matrix corresponds to a time point, the first column of the multivariate time series matrix is the base station degradation rate at the current time point, and the remaining columns in the multivariate time series matrix are the scalar degradation intensity of each target device node at the current time point;
[0081] First, for each base station node, the cloud-based operations and maintenance platform extracts time series of multiple raw state parameters from the performance monitoring database within the most recent T consecutive time windows (e.g., 5 minutes per window, 12 windows in total, covering 1 hour). These state parameters include, but are not limited to: average user throughput, wireless call drop rate, PRB utilization, handover failure rate, uplink / downlink block error rate, etc., totaling P items (e.g., P=8). Each parameter is arranged in chronological order, forming a time series of length T.
[0082] Subsequently, the platform performs state master parameter screening on the aforementioned P time series. Specifically, it calculates the coefficient of variation (COP) of each parameter's time series within the current time period, i.e., the ratio of the standard deviation to the mean, to measure the significance of its dynamic changes. The parameter with the largest COP is selected as the state master parameter for that base station. For example, if the COP of a base station's PRB utilization rate in the past hour is 0.45, which is higher than other parameters (such as throughput CV=0.32, call drop rate CV=0.28), then the PRB utilization rate is determined as the state master parameter for that base station. After obtaining the state master parameters, normalization processing is performed to obtain normalized state master parameters.
[0083] Next, the platform extracts the time series corresponding to the main state parameter and calculates its first-order time difference, reflecting the changing trend of the base station's main state parameter over time. If the main state parameter is a resource utilization indicator (an increasing value indicates performance degradation), a difference greater than 0 indicates accelerated degradation; if it is a throughput indicator (a decreasing value indicates degradation), the original series is pre-negated before differencing to unify the degradation direction. Finally, the difference is defined as the base station degradation rate at time point t.
[0084] Finally, the platform constructs a multivariate time series matrix. Assuming the base station currently connects to N target device nodes (N > 2), the matrix has N+1 columns. For each valid time point t = 2, 3, ..., T, the (t-1)th row of the matrix is filled as follows: the first column represents the base station degradation rate at that time point; the second to (N+1)th columns represent the scalar degradation intensity of each target device node at the same time point t (arranged in a fixed order by device ID). The resulting matrix has a dimension of (T-1) × (N+1), with each row corresponding to a time point, fully characterizing the temporal correlation between the base station degradation rate and the degradation intensity of its target devices.
[0085] Step 46: Calculate the Granger causality strength and reverse causality strength based on the base station degradation rate and each scalar degradation strength of the multivariate time series matrix; if the average Granger causality strength of the base station degradation rate for at least two target device nodes is greater than the reverse causality strength, calculate the inner product of the main mode direction vector and the base station structured degradation state vector under the current time window to obtain the degradation main mode activation strength; send the degradation main mode activation strength as a degradation flow signal from the base station node to each connected transmission link node.
[0086] Extract the current base station degradation rate sequence and the target device node scalar degradation intensity sequence from the multivariate time series matrix. Define a prediction historical time step number p, where p=2, indicating that the data from the previous two time points are used to predict the current value. The degradation intensity of the target device is the predicted variable, and its lag term and the lag term of the base station degradation rate are used together as explanatory variables. First, construct a constrained model:
[0087] s t The actual degradation intensity of the target equipment at time t, where α0 is a constant term, α0-α p These are the regression coefficients, s t-1 -s t-p denoted as scalar degradation intensity of the target device at the first p time points. Let be the predicted residuals of the model at time t. After fitting the model using the least squares method, calculate the sum of squared first residuals for all valid time points (from t=p+1 to t=T).
[0088] Similarly, an unrestricted model is constructed, with the same framework as above. The difference is that multiple terms multiplying the base station degradation rate with the regression coefficient are added to the above model. After fitting the model using the least squares method, the second residual sum of squares is calculated for all effective time points (from t=p+1 to t=T).
[0089] Perform an F-test and calculate the Granger causality strength: [(first residual sum of squares - second residual sum of squares) / p] / [second residual sum of squares / (T - 2p - 1)].
[0090] When calculating the reverse causality strength, the base station degradation rate is used as the predicted variable, and the equipment degradation strength is used as the potential explanatory variable. A constrained model and an unconstrained model are constructed to calculate the residual sum of squares and the F-test to obtain the reverse causality strength. The calculation process is the same as described above, and will not be repeated in detail in this embodiment.
[0091] The platform repeats the two rounds of F-tests for all target device nodes to obtain each pair of (Granger causality strength, reverse causality strength). If the number of target device nodes satisfying the condition that the Granger causality strength is greater than the reverse causality strength is ≥ 2, then the base station is determined to be the dominant degradation source. The inner product of the dominant mode direction vector and the base station's structured degradation state vector under the current time window is calculated to obtain the degradation dominant mode activation strength. Finally, the degradation dominant mode activation strength is used as a scalar degradation flow signal and broadcast by the base station node to all its connected transmission link nodes for cross-domain propagation analysis.
[0092] If the average Granger causality of the base station degradation rate with respect to at least two target device nodes is greater than the reverse causality, the core purpose is to identify whether the base station is in a dominant downlink degradation source state, thereby avoiding misjudging local terminal anomalies as network-side problems and improving the accuracy of fault location and the effectiveness of alarms.
[0093] In wireless communication networks, base stations, as central nodes, experience performance degradation (such as RF unit aging, clock drift, and decreased power amplifier efficiency). This degradation affects multiple associated terminals or edge devices simultaneously through factors like wireless signal quality, scheduling resource allocation, and synchronization mechanisms. This manifests as a deterioration in scalar degradation metrics such as increased bit error rate, decreased throughput, or worsened connection stability. This impact is unidirectional, concurrent, and shared-causal: degradation is initiated by the base station and propagates to multiple downstream devices, rather than multiple independent devices simultaneously causing base station anomalies. Therefore, if the base station degradation rate significantly improves the predictive ability for the degradation intensity of two or more devices over time (i.e., high Granger causality), and this forward causal relationship is stronger than the reverse impact of devices on the base station (i.e., lower reverse causality), it is reasonable to infer that the current base station itself is undergoing a systemic, radial performance degradation and has become the source of degradation propagation.
[0094] Step 47: Construct a link state feature vector based on the transmission link node; when the current degraded flow signal arrives at the transmission link node, concatenate the current degraded dominant mode activation intensity with the current link state feature vector to form a current joint feature vector; calculate the Mahalanobis distance between the current joint feature vector and each historical event in the preset memory bank to obtain multiple Mahalanobis distance values; if the minimum Mahalanobis distance exceeds the dynamic tolerance, determine that the current transmission link node is in an abnormal conduction state, block the continued downstream propagation of the degraded flow signal, and generate a propagation blocking binary marker. The link state feature vector includes real-time bit error rate, real-time throughput, real-time transmission delay, and real-time load rate.
[0095] When a degraded signal is transmitted from a base station node to a transmission link node, the link node first collects and normalizes four key indicators of its current operating status: real-time bit error rate, real-time throughput, real-time transmission latency, and real-time load rate. These four indicators together constitute a four-dimensional link state feature vector. Subsequently, the received degraded main mode activation intensity is concatenated with this four-dimensional vector to form a five-dimensional current joint feature vector. This joint feature vector simultaneously contains the intensity information of the upstream base station's degraded mode and the current link's own real-time operating status, providing a complete basis for subsequent judgment on whether the link can faithfully transmit the degraded signal.
[0096] The platform pre-maintains a memory that stores several joint feature vector samples from historically confirmed normal transmission events. These samples all originate from scenarios where the base station deteriorated, but the transmission link functioned normally, and the deteriorated signal was accurately transmitted. The memory also synchronously stores the overall statistical distribution characteristics of these historical samples, including the mean of each dimension and the covariance relationship between each dimension. When a new joint feature vector arrives, its Mahalanobis distance to the overall distribution in the memory is calculated. Mahalanobis distance is a distance metric that considers the correlation and dimensional differences between feature dimensions. For example, high load is often accompanied by low throughput and high latency; this coupling relationship is captured by the covariance structure, thus the distance calculation can more accurately identify abnormal deviations.
[0097] If the calculated minimum Mahalanobis distance exceeds the currently set dynamic tolerance, the transmission link node is determined to be in an abnormal conduction state. At this time, two operations are immediately performed: first, the degraded flow signal is blocked from continuing to propagate to downstream nodes to prevent distorted or amplified degraded information from misleading subsequent analysis; second, a propagation interruption binary flag (with a value of 1) is generated to record this propagation interruption event for use by the subsequent root cause localization module.
[0098] The dynamic tolerance is obtained based on the real-time bit error rate, real-time load rate, and preset basic tolerance: if the real-time load rate is less than 0.3 and the real-time bit error rate is less than 0.02, the dynamic tolerance is defined as the preset basic tolerance; if the real-time load rate is greater than 0.7 or the real-time bit error rate is greater than 0.1, the dynamic tolerance is defined as x times the preset basic tolerance, where x is between 0.1 and 0.3; otherwise, the dynamic tolerance is calculated based on the real-time bit error rate, real-time load rate, and preset basic tolerance.
[0099] When the real-time load rate is less than 30% and the real-time bit error rate is less than 2%, the link is in a light-load, high-quality state, with high system stability and sensitivity to external disturbances. At this time, the dynamic tolerance is set to a preset base tolerance value, and stricter criteria are used to ensure that even minor anomalies can be effectively identified.
[0100] When the real-time load rate exceeds 70% or the real-time bit error rate exceeds 10%, the link is already in a high-load or high-error-rate state, and the network's inherent fluctuations increase significantly. If strict tolerances are still used, normal fluctuations are easily misjudged as abnormal propagation. Therefore, the dynamic tolerance is reduced to between 10% and 30% of the basic tolerance, that is, the judgment criteria are significantly relaxed to adapt to high-noise operating conditions.
[0101] In other intermediate states (i.e., real-time load rate between 30% and 70%, and real-time bit error rate between 2% and 10%), the dynamic tolerance is calculated using linear interpolation based on the real-time load rate, real-time bit error rate, and a preset basic tolerance. The specific implementation process is as follows:
[0102] First, the impact weights of load rate and bit error rate on dynamic tolerance are calculated separately. For the real-time load rate, the ratio of the portion exceeding 30% to the 40% interval length (i.e., 70% minus 30%) is used to obtain the load impact factor; for the real-time bit error rate, the ratio of the portion exceeding 2% to the 8% interval length (i.e., 10% minus 2%) is used to obtain the bit error rate impact factor. Both impact factors are limited to between zero and one.
[0103] Subsequently, the system assigns preset adjustment coefficients to the load impact factor and the bit error rate impact factor. Both of these adjustment coefficients are constants between zero and one, used to reflect the relative importance of load and bit error rate on link stability in network operation and maintenance experience. For example, in a network primarily used for data services, the load rate may have a greater impact on transmission fidelity. In this case, the load adjustment coefficient can be set to 0.6, and the bit error rate adjustment coefficient can be set to 0.4.
[0104] Next, the two weighted impact factors are added together to obtain the overall adjustment ratio. This ratio is also limited to between zero and one. Finally, the overall adjustment ratio is subtracted from one and then multiplied by the preset basic tolerance to obtain the dynamic tolerance under the current operating conditions.
[0105] For example: If the current real-time load rate is 50%, the real-time bit error rate is 5%, the preset basic tolerance is 2.5, the load adjustment coefficient is 0.6, and the bit error adjustment coefficient is 0.4, then the load impact factor is (50%−30%)÷40% =0.5, and the bit error impact factor is (5%−2%)÷8% ≈ 0.375; after weighting, they are 0.3 and 0.15 respectively, and the comprehensive adjustment ratio is 0.45; the dynamic tolerance is (1−0.45)×2.5 = 1.375.
[0106] Because the effective propagation of degraded flow signals depends on the health of the transmission link itself. If the link itself is congested, interfered with, or has hardware anomalies, then its transmission of upstream degraded signals is no longer a faithful process, but may be superimposed with its own noise, or even generate false degraded characteristics. If it continues to propagate without discrimination, it will cause downstream nodes to make incorrect decisions based on distorted information, such as misjudging core network failures or triggering unnecessary protection switching. Therefore, the reliability of the propagation must be verified at each hop link. This embodiment can accurately block abnormal propagation paths, blocking the signal only when the relay link itself is abnormal, avoiding a one-size-fits-all approach and ensuring the end-to-end transmission of normal degraded flows; it has adaptive anti-interference capabilities, and the dynamic tolerance mechanism enables the system to remain stable under complex conditions such as high load and high bit error rate, significantly reducing the false alarm rate; it supports efficient root cause localization, and the binary marker for blocking propagation directly indicates where the degraded propagation is interrupted, providing key clues for operation and maintenance personnel, suggesting that the problem may lie in the link itself, rather than the upstream base station or downstream equipment.
[0107] Step 48: For each base station node, calculate the proportion of transmission link nodes whose propagation of degraded flow signals is blocked to the total number of transmission link nodes; if the proportion exceeds a preset proportion threshold, it is determined that the degraded energy has locally converged at the current base station; and calculate the cosine distance between the main mode direction vector of the current base station and the main mode direction vector of the adjacent base stations; if the cosine distance is greater than the isolation threshold, it is determined that the degraded mode of the base station has topological isolation; mark the base station nodes that simultaneously satisfy the local convergence condition and the topological isolation condition as predicted fault points, and aggregate all directly connected target device nodes into a cluster of devices to be dealt with.
[0108] For each base station node, the system first obtains the total number of all downlink transmission link nodes connected to it (i.e., the total number of relay or backhaul links served by the base station), denoted as the total number of links. Then, the system counts how many of these links actively blocked the continued propagation of the degraded flow signal due to the abnormal conduction determination in step S47 during the current degradation event period, denoted as the blocked link number. The proportion of blocked links to the total number of links is calculated: Blocking ratio = Blocked link number / Total number of links. If this ratio exceeds a preset threshold (e.g., set to 60%), it is determined that the degradation energy has locally converged at the current base station.
[0109] The preset threshold ratio is typically set between 50% and 70%, preferably 60%. This value is derived from statistical analysis of numerous historical fault cases: when the base station itself experiences hardware degradation (such as RF unit aging, clock jitter, power amplifier nonlinear distortion, etc.), the resulting degraded signals propagate outward through multiple links; however, since the degradation source is concentrated in the base station itself, multiple downstream links often trigger blockages due to abnormal conditions or inability to maintain fidelity during transmission. Actual measurement data shows that, on average, over 60% of downlinks exhibit conduction anomalies during the precursor stage of real base station-level faults. If the threshold is too low (e.g., below 40%), temporary disturbances on individual links may be misjudged as base station aggregation; if it is too high (e.g., exceeding 80%), early degradation may be missed. Therefore, 60% is an empirical balance point that considers both sensitivity and specificity.
[0110] Extract the dominant mode direction vector of the current base station and calculate the cosine distance between it and the dominant mode direction vectors of all geographically or topologically adjacent base stations. The cosine distance is defined as 1 minus the cosine of the angle between the two vectors, ranging from 0 to 2. The larger the cosine distance, the more inconsistent the evolution directions of the two degradation modes in the feature space. If the cosine distance between the current base station and all adjacent base stations (topological adjacency data of the current base station retrieved from the communication network management system, including other base stations that share the same aggregation node, are connected to the same transmission ring network, or are interconnected through the same core router. These base stations constitute direct adjacency relationships physically or logically due to sharing some transmission resources) is greater than a preset isolation threshold (e.g., set to 0.7), then the degradation mode of the base station is determined to have topological isolation.
[0111] The isolation threshold is typically set between 0.6 and 0.8, preferably 0.7. This threshold is based on a comparative analysis of normal cooperative degradation scenarios (such as regional interference, power fluctuations, and weather effects) and single-point hardware failure scenarios: In regional events, the main degradation modes of adjacent base stations are highly similar, and the cosine distance is generally less than 0.5; while in single-point hardware failures, the degradation mode is determined by the characteristics of local devices and is unrelated to neighboring stations, and the cosine distance is often greater than 0.7. Through ROC curve analysis, 0.7 is the optimal segmentation point to distinguish between the two types of scenarios, which can effectively suppress false alarms of regional events while maintaining a high detection rate for isolated fault points.
[0112] A base station is marked as a predicted fault point only if it meets both of the following conditions:
[0113] The proportion of degraded flow blocking exceeds the preset proportion threshold (local convergence is established).
[0114] The cosine distance between the main mode direction vectors of the neighboring base stations and all base stations in the set is greater than the isolation threshold (topological isolation holds).
[0115] Once a device is marked as a predicted fault point, all target device nodes directly served by it (such as user terminals, IoT sensors, micro base stations, etc.) are automatically aggregated into a cluster of devices to be handled, which can be used for subsequent preventive maintenance, resource rescheduling, alarm push or automatic isolation operations.
[0116] True base station-level hardware degradation typically exhibits a dual characteristic of energy convergence and mode isolation. On one hand, the degradation source is concentrated within the base station itself, causing the degraded signal output to be obstructed on multiple downlinks (manifested as a high blocking ratio). On the other hand, this type of degradation is caused by local component aging or configuration errors, and has no common cause with surrounding base stations. Therefore, its degradation evolution direction differs significantly from other base stations in the feature space. Judging solely by a high blocking ratio may misclassify regional interference (such as thunderstorms causing simultaneous degradation at multiple stations) as a single point of failure; judging solely by directional isolation may mislabel individual link anomalies (failure to converge) as failures. Only by combining both approaches can we accurately identify potential base station-level hazards that are about to occur but have not yet completely failed.
[0117] By employing dual-condition filtering, regional events and localized link noise are effectively eliminated, focusing on truly high-risk single-point degraded base stations and significantly improving prediction accuracy. The generation of clusters of devices to be addressed allows for targeted maintenance actions (such as handover or rate limiting only for users under that base station), avoiding network-wide disruption.
[0118] Step 5: Perform root cause analysis on the device cluster to obtain the root cause ranking result; perform composite judgment logic on the top n target devices in the root cause ranking result to determine the fault type, generate a pre-diagnosis instruction set, and perform automated recovery operation.
[0119] Step 51: Mark each target device node in the device cluster as a candidate device node, mark the base station node connected to the candidate device node as a candidate base station node, and mark the propagation link node connected to the candidate base station node as a candidate link node; construct base station sector nodes based on the candidate device nodes connected to the candidate base station nodes; form a first type of edge based on the candidate device node pointing to its corresponding base station sector node; form a second type of edge based on the candidate device node pointing to the candidate link node it uses; form a third type of edge based on the candidate link node pointing to its corresponding base station sector node; construct a heterogeneous operation and maintenance graph based on the candidate device nodes, base station sector nodes, candidate link nodes, first type of edge, second type of edge, and third type of edge.
[0120] Based on the cluster of devices to be processed output in step S48, each target device node in the cluster is uniformly marked as a candidate device node. These devices are the final receivers or service objects of the degraded signals, and their abnormal behavior is a direct manifestation of the fault's impact.
[0121] For each candidate device node, query the identifier of the base station node that it is currently registered or connected to, and mark all the referenced base station nodes as candidate base station nodes after deduplication.
[0122] For each candidate base station node, retrieve all downlink transmission link nodes involved in step S47 and mark these link nodes as candidate link nodes.
[0123] This embodiment reuses the sector grouping result: for each candidate base station node, iterate through the sector group set generated in the previous steps; for each non-empty sector group, create a unique base station sector node.
[0124] After the node hierarchy is established, the system constructs semantic edges according to the following rules:
[0125] Type 1 edge (device → sector)
[0126] For each candidate device node, a directed edge is created from the device node to the base station sector node to which it belongs, based on its affiliation in the preceding sector group. This edge directly reflects the preceding grouping result, ensuring that the graph structure is consistent with the actual coverage relationship.
[0127] Second type of edge (device → link)
[0128] Based on the business flow path record, a directed edge is created for each candidate device node pointing to the candidate link node it uses, representing the transmission path dependency.
[0129] Type III edge (link → sector)
[0130] For each candidate link node, the base station sector node it serves is determined based on the network configuration data, and a directed edge is created from the link node to the sector node, representing the infrastructure support relationship for the wireless sector.
[0131] Finally, the following elements are integrated to construct a heterogeneous operation and maintenance diagram:
[0132] Nodes: Candidate device nodes, base station sector nodes, candidate link nodes;
[0133] Edges: Type 1 edge (belonging), Type 2 edge (use), Type 3 edge (support).
[0134] The graph structure strictly inherits the semantics of the preceding sector grouping, ensuring that each base station sector node represents a subgroup of devices that has been verified to have degraded behavior, thereby making the entire graph focus on local network areas where abnormal propagation paths actually exist.
[0135] Step 52: Define a first initial anomaly score for each candidate device node based on the heterogeneous operation and maintenance diagram, and define a second initial anomaly score for each base station sector node and candidate link node, where the second initial anomaly score is 0; obtain the first initial score by weighted summation based on the scalar degradation intensity and the blocking propagation binary label of the current candidate device node. The weighting coefficient for the scalar degradation intensity is 0.7, and the weighting coefficient for the blocking propagation binary label is 0.3.
[0136] Step 53: Iteratively perform deterministic anomaly propagation on the heterogeneous operation and maintenance graph. In each round, update the results based on the first initial anomaly score of each candidate device node to obtain the root cause score of each candidate device node, its associated base station sector node, and candidate link node.
[0137] Step 531: In each round, the initial anomaly scores of the candidate device nodes, base station sector nodes, and candidate link nodes in the heterogeneous operation and maintenance graph are iteratively updated:
[0138] a. Based on the base station sector node, obtain the candidate device node and candidate link node of the incoming edge, denoted as the first incoming edge and the second incoming edge, respectively; determine the sum of the initial anomaly scores of all the first incoming edges of the base station sector node and multiply them by the first propagation coefficient, denoted as the first anomaly score; determine the sum of the initial anomaly scores of all the second incoming edges of the base station sector node and multiply them by the second propagation coefficient, denoted as the second anomaly score; calculate the sum of the first anomaly score and the second anomaly score as the base station sector anomaly score;
[0139] The first propagation coefficient, the second propagation coefficient, and the third propagation coefficient satisfy the following condition: 0 < second propagation coefficient < first propagation coefficient < third propagation coefficient < 1;
[0140] b. Based on the candidate link node, obtain the candidate device node of the incoming edge, which is denoted as the third incoming edge. Calculate the sum of the initial anomaly scores of all third incoming edges and multiply it by the third propagation coefficient to obtain the third anomaly score, which is used as the anomaly score of the candidate link node.
[0141] First, for each base station sector node, the following operations are performed in each iteration: All first-type edges pointing to the base station sector node are identified, with their origins being candidate device nodes belonging to that sector; these are collectively referred to as first incoming edges. Simultaneously, all third-type edges pointing to the base station sector node are identified, with their origins being candidate link nodes supporting the sector's transmission function; these are collectively referred to as second incoming edges. Then, the anomaly scores of all candidate device nodes corresponding to the first incoming edges in the current iteration are accumulated, and the sum is multiplied by a preset first propagation coefficient (preferably 0.6) to obtain the first anomaly score. Similarly, the anomaly scores of all candidate link nodes corresponding to the second incoming edges in the current iteration are accumulated, and the sum is multiplied by a preset second propagation coefficient (preferably 0.3) to obtain the second anomaly score. Finally, the first anomaly score and the second anomaly score are added together to obtain the anomaly score of the base station sector node in the next iteration. This design fully reflects that the abnormal state of a base station sector is influenced by both the behavior of the terminal group it serves and the health status of its underlying transmission links; neither is dispensable.
[0142] Secondly, for each candidate link node, the system performs the following operation in each iteration: It identifies all second-type edges pointing to the candidate link node. These edges originate from candidate device nodes using the link for data transmission and are collectively referred to as third-incoming edges. The anomaly scores of all candidate device nodes corresponding to these third-incoming edges in the current iteration are accumulated, and the sum is multiplied by a preset third propagation coefficient (preferably 0.85). The result is used as the anomaly score of the candidate link node in the next iteration. This mechanism is based on a core operational understanding: anomalies in the link itself cannot be directly observed and can only be indirectly inferred through anomalies in the quality of multiple terminal services it carries. Therefore, only when multiple devices via the same link simultaneously exhibit high anomaly scores will the system consider the link potentially degraded and assign it a higher anomaly score.
[0143] The reason why the above update rules must be designed this way is that the direction and semantics of various edges in the heterogeneous operation and maintenance graph have clear causal orientation. Device nodes are the source of abnormal behavior, and their performance degradation is an observable fact; while base station sector nodes and candidate link nodes belong to the infrastructure layer, and their abnormal states need to be deduced from the behavior of upper-layer devices. Therefore, information can only be propagated unidirectionally from device nodes to sector or link nodes, and cannot be assigned in reverse. If the propagation direction is reversed, it will lead to confusion in causal logic, making fault location lose its physical basis. In addition, if any incoming edge type is ignored (for example, only considering the impact of devices on sectors and ignoring links), the system will not be able to accurately identify the real root cause when facing pure transmission failures or mixed failure scenarios, which will seriously weaken the robustness and applicability of the solution.
[0144] Step 532: If the number of propagation rounds already executed reaches the preset maximum number of rounds, and the maximum absolute change in the abnormal scores of all updated base station sector nodes and candidate link nodes in two consecutive rounds of propagation is less than the preset convergence threshold, then stop the iterative propagation and execute the next step.
[0145] In this embodiment, after each round of abnormal score update, a termination judgment procedure is immediately initiated to check whether the following two conditions are met simultaneously:
[0146] First condition: The number of propagation rounds has reached the maximum.
[0147] This embodiment maintains a round counter, initially set to 0, which increments by 1 after each complete update of S531. The first condition is considered satisfied when the current value of this counter equals the preset maximum number of propagation rounds (e.g., 5 rounds). This constraint aims to prevent infinite loops in the iteration process due to complex graph structures or unusual distributions, ensuring the algorithm's time controllability in the worst-case scenario.
[0148] The second condition is that abnormal score changes tend to stabilize.
[0149] Iterate through all base station sector nodes and candidate link nodes in the current heterogeneous operation and maintenance graph. For each such node, calculate the absolute difference between its abnormal score after the current update and its abnormal score after the previous update. Then, take the maximum value among all these differences and record it as the maximum absolute change. If the maximum absolute change is less than a preset convergence threshold (e.g., set to 0.01), it is considered to meet the second condition.
[0150] Only when both of the above conditions are met is it determined that the abnormal propagation process has fully converged, the subsequent iterations are immediately terminated, and the process proceeds to the next step; otherwise, the next round of S531 update continues.
[0151] Step 533: Calculate the root cause score for each target device based on the updated anomaly scores of the base station sector nodes, candidate link nodes, and the first initial anomaly score.
[0152] First, an initial anomaly score is obtained as the basic anomaly strength. Then, for the associated base station sector node, the final anomaly score obtained after the termination of anomaly propagation iterations is extracted and divided by a preset sector anomaly reference threshold (e.g., 0.8, derived from historical fault data statistical analysis: over 95% of non-serious sector anomalies have scores below this value, while truly alarm-inducing sector anomalies usually exceed this threshold), thus obtaining a sector anomaly standardization ratio. This ratio is then multiplied by a preset first environmental commonality suppression coefficient (value range greater than 0 and less than 1), and the product is subtracted from 1. If the result is less than 0, it is truncated to 0, ultimately obtaining the sector health confidence. Similarly, for the transmission link node used by the device, its final anomaly score is divided by a preset link anomaly reference threshold (e.g., 0.75, determined based on the anomaly score distribution of historical congestion and bit error events in the backhaul link), modulated by a second environmental commonality suppression coefficient (value range greater than 0 and less than 1), and the result is subtracted from 1 (again, truncated to 0 at the lower limit) to generate the link health confidence. Ultimately, the root cause score of the device is calculated by multiplying the first initial anomaly score by the sector health confidence score by the link health confidence score.
[0153] Sector and link are the two major upstream shared sources of influence for device anomalies. Their health status jointly determines whether the device anomaly is individual-specific. Only when both are in a relatively healthy state should a high initial anomaly score be highly trusted as the true root cause. If either upstream resource is severely abnormal, the corresponding health confidence level drops significantly, thereby effectively suppressing the root cause score of that device and avoiding misjudging common environmental degradation as terminal failure. Using a product-based fusion method can accurately model the joint necessary condition of sector health and link health.
[0154] Step 54: Sort all candidate device nodes in descending order of root cause score to generate root cause ranking result; if the root cause score of the base station sector node to which the first n candidate device nodes belong in the root cause ranking result is greater than the first sector threshold, and the proportion of the number of target devices whose propagation link nodes within the base station sector node to the total number of target devices in the sector group is greater than the preset density threshold, then the fault type is determined to be antenna beamforming offset fault type.
[0155] In the specific implementation of this application, all candidate device nodes are sorted in descending order of root cause score to generate a root cause ranking result. If the root cause score of the base station sector node to which the top n (n=5) candidate device nodes belong in the root cause ranking result is greater than the first sector threshold, and the proportion of the number of target devices whose propagation link nodes within the base station sector node are blocked to the total number of target devices in the sector group is greater than a preset density threshold, then the fault type is determined to be an antenna beamforming offset fault type. The fault type determination is based on a high-confidence fault identification logic constructed from the physical characteristics of the beamforming mechanism and the abnormal propagation topology features in 5G millimeter wave / massive MIMO networks. Its core idea is that antenna beamforming offset is a typical sector-level spatial directional fault that does not cause a full sector service interruption, but it will cause multiple terminals in a specific spatial area to degrade synchronously. These terminals exhibit a special pattern in the abnormal propagation map where the link is not broken, but the signal quality drops sharply. Therefore, only when both the high sector root cause score and the high spatial anomaly density are met simultaneously can the fault type be uniquely identified.
[0156] First, antenna beamforming relies on the base station antenna array to accurately align the user equipment's spatial location with the beam. If the calibration parameters are incorrect, the mechanical tilt angle is offset, or the digital beam weight configuration is abnormal, the main lobe direction will deviate from the expected coverage area, causing multiple CPE terminals that should be in the strong signal area to fall into the side lobe or zero depression area at the same time, resulting in a sharp drop in SINR and a plunge in throughput, but the physical link (such as fiber backhaul, IP connectivity) is not interrupted. In such scenarios: (1) the node in this sector itself has accumulated a high abnormal score due to the deterioration of the KPI of a large number of terminals; (2) multiple terminals connected to this sector, although the link is normal (the propagation path is not blocked), are marked as abnormal receivers because they cannot effectively receive beam energy; (3) these terminals are often geographically clustered in a certain azimuth angle or distance range, forming a high-density abnormal cluster. Therefore, by checking whether the top n high root cause devices are concentrated in the same sector (sector root cause score > first sector threshold) and verifying whether the proportion of devices in the sector with unbroken links but abnormal exceeds the density threshold, the typical fingerprint of beam offset can be effectively captured. This judgment logic directly reflects the two essential characteristics of beamforming faults: spatial correlation and non-link interruption.
[0157] First sector threshold (e.g., set to 0.65): This value is derived from the statistical distribution of anomaly scores for various sector-level events in the historical fault database. Preset density threshold (e.g., set to 40%): This value originates from spatial simulation and experimental data of typical 5G CPE deployment scenarios. In the 3.5GHz or 26GHz bands, a sector typically covers 8–20 CPEs, and the area affected by beam offset generally accounts for 30%–60% of the sector's coverage area. Experiments show that when the main lobe of the beam offsets by more than 10°, the proportion of affected CPEs typically exceeds 40%. Therefore, setting the density threshold to 40% ensures that a judgment is triggered only when abnormal devices form a significant spatial cluster within the sector, avoiding false alarms caused by a small number of randomly distributed abnormal terminals.
[0158] Secondly, this dual-condition design possesses strong fault specificity and anti-interference capabilities. On one hand, a sector root cause score greater than the first sector threshold excludes isolated anomalies caused by individual terminal hardware failures or user-side issues (such as router failures), ensuring that the problem has a sector-level impact range. On the other hand, a link not broken but anomaly device density greater than a preset density threshold effectively distinguishes beam offset from other common faults. For example, if it is a fiber break or a base station main control board failure, it usually leads to a complete blockage of link propagation, and the anomaly device appears as a link node failure in the propagation diagram, rather than a normal link but abnormal reception. If it is external interference (such as radar or illegal emission sources), the anomaly device may be distributed across sectors, making it difficult to form a high-density cluster within a single sector. Therefore, only when the antenna beam itself experiences spatial pointing deviation will both the significant sector anomaly and the high-density non-link-interruption type terminal anomaly be simultaneously satisfied. This combined criterion significantly reduces the false positive rate, enabling the system to accurately pinpoint beamforming offset from dozens of potential faults and automatically trigger antenna calibration or beam re-optimization processes, significantly shortening MTTR (Mean Time To Repair).
[0159] Step 55: If the root cause score of the transmission link node used by the top n candidate device nodes in the root cause sorting result is greater than the first link threshold, and the dynamic tolerance of the current transmission link node shows a monotonically decreasing trend in the most recent P consecutive time windows, then the fault type is determined to be the link carrying capacity degradation fault type.
[0160] If the root cause score of the transmission link node used by the top n candidate device nodes in the root cause ranking results is greater than the first link threshold, and the dynamic tolerance of the current transmission link node shows a monotonically decreasing trend in the most recent P consecutive time windows, then the fault type is determined to be a "link carrying capacity degradation fault type". This is not a collection of empirical rules, but a high-confidence diagnostic logic built on the progressive characteristics of transmission network performance degradation and the behavior patterns of link nodes in the abnormal propagation graph. Its core lies in identifying a typical latent, slowly deteriorating fault, that is, the physical link is not interrupted, but its effective bandwidth, latency stability or bit error rate and other key indicators continue to deteriorate, causing the quality of multiple high-priority CPE services carried by it to decline simultaneously. Such faults (such as optical module aging, microbending loss accumulation, QoS policy configuration drift, etc.) often cannot trigger traditional link interruption alarms, but will significantly affect user experience. Therefore, early identification must be carried out through the fusion of multi-dimensional dynamic indicators.
[0161] First link threshold (e.g., set to 0.7): This value is determined based on the statistical distribution of root cause scores for historical link failure events. Analysis of the final root cause scores of link nodes identified as cases of optical module aging, microbending loss, QoS policy errors, etc., revealed that the scores for actual capacity degradation events were generally higher than 0.65, while the scores for transient congestion or configuration errors were mostly between 0.4 and 0.6. Setting the threshold to 0.7 effectively filters non-degradation anomalies while maintaining high accuracy (>85%). P-value (e.g., P=5): Represents the number of continuous time windows to be examined. Each window defaults to 5 minutes, covering the trend of the most recent 25 minutes. A P-value that is too small (e.g., P=2) is susceptible to short-term noise interference, while a P-value that is too large (e.g., P=10) results in a delayed response. Backtesting on 3 months of data from the live network showed that P=5 can stably capture over 90% of progressive degradation events with a false alarm rate of less than 3%. A strict definition of monotonically decreasing: small fluctuations (e.g., ±2%) are allowed between adjacent windows, but the overall trend slope must be negative, and at least (P−1) windows must satisfy the month-on-month decrease. This design balances mathematical rigor with engineering robustness, avoiding missed detections due to measurement jitter.
[0162] Link capacity degradation has two essential characteristics: (1) The impact is concentrated on a specific transmission path. All terminal devices sharing the link will experience synchronous phenomena such as decreased throughput and increased jitter, thus causing the link node to accumulate a high root cause score in the abnormal propagation graph; (2) The performance degradation is temporally continuous and monotonic. Unlike sudden congestion or instantaneous interference, the capacity decay caused by hardware aging or configuration drift is usually manifested as a gradual decrease in service capacity (i.e., dynamic tolerance) within multiple consecutive time windows, rather than random fluctuations. Therefore, only when the link node on which the high root cause device depends simultaneously satisfies both significant abnormality (root cause score > first link threshold) and continuous service capacity decay (monotonically decreasing dynamic tolerance) can it uniquely point to the root cause of capacity degradation. This judgment logic accurately captures the topological clustering and temporal evolution regularity of such faults.
[0163] This dual-criteria design possesses strong fault differentiation capabilities and early warning value. On one hand, a link root cause score greater than the first link threshold ensures that the problem has a common link-level nature, excluding interference from individual terminal application layer anomalies or wireless-side interference. On the other hand, a monotonically decreasing dynamic tolerance within P windows effectively distinguishes between capacity degradation and other transient link problems. For example, in the case of a sudden DDoS attack or a temporary large volume of uploads, the dynamic tolerance may drop sharply but not monotonically, and recover quickly after the attack ends; in the case of routing oscillations, abnormal devices may be scattered across multiple links, making it difficult to form a high root cause cluster on a single link. True hardware aging or configuration drift inevitably leads to irreversible and gradual degradation of service capabilities. Therefore, this mechanism can not only accurately identify degradation events that have already caused significant business impact, but also issue early warnings before performance falls below the SLA threshold, enabling predictive maintenance.
[0164] Step 56: If the root cause scores of the top n candidate device nodes in the root cause ranking result are greater than the preset root cause significance threshold, the root cause scores of the base station sector nodes to which they belong are less than or equal to the first sector threshold, and the root cause scores of the transmission link nodes used are less than or equal to the first link threshold, then the fault type is determined to be the target device local hardware abnormality fault type.
[0165] First, the top n devices with the root cause ranking results are extracted from the candidate device set that has completed root cause scoring. Each device is then verified to ensure that it meets all three conditions: (1) the device's own root cause score is higher than the preset root cause significance threshold; (2) the root cause score of the base station sector node to which it belongs does not exceed the first sector threshold; and (3) the root cause score of the transmission link node to which it depends does not exceed the first link threshold. Only when all three conditions are met is the device marked as a local hardware abnormality fault type of the target device, and a corresponding maintenance work order is generated, recommending on-site device replacement or return to the factory for testing.
[0166] The engineering rationale behind this judgment logic stems from a deep understanding of terminal failure modes in 5G Fixed Wireless Access (FWA) networks. Real-world CPE local hardware failures (such as RF front-end damage, baseband processor crashes, and power module aging) are typically highly isolated and non-propagating, meaning they only affect a single device and do not cause synchronous degradation of other terminals or services on shared links within the same sector. Therefore, in the anomaly propagation map, such devices exhibit extremely high individual anomaly intensity (high root cause score), while their upstream sectors and link nodes remain in a low-anomaly state because they are unaffected. By setting threshold boundaries, it is possible to effectively distinguish between genuine hardware failures and pseudo-anomalies (such as temporary obstruction, user-side router problems, or weak coverage edge effects). For example, if a CPE experiences a brief signal interruption due to obstruction from a construction crane outside its window, its root cause score may increase, but it usually won't exceed the significance threshold (because its degradation is limited and recoverable). Even if it does exceed it, if other devices in the same sector experience similar fluctuations, the sector's root cause score will rise and trigger a common fault determination (such as beam shift), thus excluding it from the S56 logic. Only those devices with exceptionally prominent and clean contexts will be classified as local hardware faults.
[0167] Preset root cause significance threshold (typical value: 0.85): This threshold was derived by modeling the root cause score distribution of over 2,300 cases confirmed as CPE hardware failures in the past 12 months. Statistics show that 92% of genuine hardware failure devices had a final root cause score ≥ 0.83, while high-abnormality devices caused by non-hardware reasons (such as user configuration errors or temporary interference) had scores concentrated in the 0.60–0.78 range. To balance recall and precision, a threshold of 0.85 was set, ensuring a fault detection rate of over 88% while keeping the false alarm rate below 5%.
[0168] Step 57, otherwise, determine the fault type as a composite abnormal fault type.
[0169] In the specific implementation process of this application, when the criteria conditions of steps S54, S55 and S56 are not met, the fault type of the current abnormal event is determined to be a composite abnormal fault type, based on the high confidence composite causal inference obtained after cross-analysis of multi-dimensional indicators in the abnormal propagation spectrum. The specific implementation process is as follows: traverse the first n candidate device nodes in the root cause sorting results, and verify in turn whether they trigger all the necessary conditions of any fault type in S54, S55 or S56; if all the condition combinations are not met, start the composite abnormal assessment process, and further check whether there is a phenomenon of multiple weak abnormal sources coexisting, for example: (1) the root cause score of a certain sector is slightly lower than the threshold of the first sector (such as 0.62 < 0.65), but some of the devices under its jurisdiction also experience slight link degradation (the link root cause score is close to but does not exceed the first link threshold); (2) multiple low-intensity interference sources (such as neighboring PCI conflict and Wi-Fi co-frequency interference) superimposed cause the terminal status parameters to fluctuate continuously; (3) CPE firmware version defects and base station scheduling strategies are incompatible, causing intermittent connection failures. In such scenarios, a single fault model cannot fully explain the observed anomaly propagation topology. Therefore, a composite anomaly category is introduced to accurately reflect the complexity of real-world operations and maintenance.
[0170] Step 58: Match the predefined fault type template according to the determined fault type, generate the corresponding pre-diagnostic instruction set and execute it to perform automated recovery operation on the target device.
[0171] In the specific implementation of this application, taking the link carrying capacity degradation fault as an example: when it is determined that the bit error rate of the transmission link used by a certain CPE device continues to increase due to the aging of the optical module (link root cause score > first link threshold), the corresponding recovery template is immediately retrieved from the template library and the automatic recovery engine is started. The engine first sends a Telemetry subscription command to the access gateway through the NetConf protocol to collect the optical power, FEC error correction count and queue packet loss rate of the link in real time; if it is confirmed within 5 seconds that the received optical power is lower than -28dBm and the proportion of uncorrectable FEC frames exceeds 0.1%, the link redundancy switching process is automatically triggered: (1) query the backup link resource pool bound to the CPE (such as another PON port or microwave backup link); (2) send a port redirection command to the OLT to map the CPE's GEM Port to the backup port; (3) synchronously update the N6 interface routing table of the UPF to ensure seamless migration of the user plane path. The entire switching process is completed within 30 seconds. During this period, the system continuously monitors the throughput and latency indicators of the CPE. If the KPI recovers to the baseline level within 10 seconds (such as the downlink rate recovering to more than 90% of the contracted bandwidth), the recovery is considered successful, the alarm is automatically turned off, and the operation log is archived. If the KPI does not improve or a new anomaly occurs, a safety rollback is immediately performed: the original port configuration is restored, the fault level is upgraded, and the case is handed over to manual handling.
[0172] For local hardware failures of the target device, automated recovery focuses on ensuring business continuity rather than repairing the device itself. After confirming that the device has a high root cause score and the sector / link is normal, a remote hard reboot is first attempted (sending a reboot command via the TR-069 protocol); if the RSRP is still below -110dBm and the connection success rate is <50% within 3 minutes after reboot, the service drift mechanism is initiated: (1) query whether there are other online CPEs or smart terminals supporting Wi-Fi 6 under the same home gateway; (2) issue a network selection command through the ANDSF policy to guide user traffic to switch to the backup access point; (3) push the device replacement work order to the CRM system at the same time, and coordinate with the warehouse system to pre-allocate a new machine of the same model. During this process, the system continuously verifies whether the user's business has been restored. Once the video stuttering rate of the alternative path is detected to drop below 1%, it is considered that the automated recovery is complete.
[0173] like Figure 2 As shown, this embodiment also discloses a cloud-based collaborative remote operation and maintenance and fault prediction system for CPE devices, including the following modules:
[0174] The first acquisition module is used to collect operational indicator data based on multiple target devices;
[0175] The second acquisition module is used to acquire network topology data related to the target device according to the network management system;
[0176] The graph construction module is used to construct association edges between the target device, operational indicator data, base station identifier, and transmission link identifier as nodes, and to construct a heterogeneous association graph based on the nodes and association edges.
[0177] The prediction module is used to construct multiple degradation indicators based on the operational indicator data within a continuous time window, identify target base station nodes as predicted fault points in the heterogeneous correlation graph based on the degradation indicators, and identify target device nodes directly connected to the predicted fault points as device clusters.
[0178] The judgment and analysis module is used to perform root cause analysis on the device cluster to obtain the root cause ranking result; perform composite judgment logic on the top n target devices in the root cause ranking result to determine the fault type, generate a pre-diagnosis instruction set, and perform automated recovery operation.
[0179] To implement the above embodiments, this application also proposes an electronic device. Please see [link to relevant documentation]. Figure 3 , Figure 3 This is a schematic diagram of the structure of the electronic device provided in an embodiment of this application. For example... Figure 3As shown, the electronic device 500 includes: a processor 501 and a memory 502 communicatively connected to the processor 501; the memory 502 stores computer-executable instructions; the processor 501 executes the computer-executable instructions stored in the memory to implement the method provided in the foregoing embodiments.
[0180] To implement the above embodiments, this application also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided in the foregoing embodiments.
[0181] To implement the above embodiments, this application also proposes a computer program product, including a computer program that, when executed by a processor, implements the methods provided in the foregoing embodiments.
[0182] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A cloud-based collaborative method for remote operation and maintenance and fault prediction of CPE devices, characterized in that: Includes the following steps: Step 1: Collect operational indicator data based on multiple target devices. The operational indicator data includes signal quality indicators, throughput indicators, resource utilization indicators, and corresponding sampling timestamps. The time window is determined based on the sampling timestamp; Step 2: Obtain network topology data related to the target device from the network management system. The network topology data includes the base station identifiers accessed by each target device and the transmission link identifiers on which the base stations depend. Step 3: Based on the target device, operational indicator data, base station identifier, and transmission link identifier as nodes, construct association edges between the nodes, and construct a heterogeneous association graph based on the nodes and association edges; Step 4: Construct multiple degradation indicators based on the operational indicator data within a continuous time window; identify target base station nodes as predicted fault points in the heterogeneous association graph based on the degradation indicators; and identify target device nodes directly connected to the predicted fault points as device clusters. Step 5: Perform root cause analysis on the device cluster to obtain the root cause ranking result; perform composite judgment logic on the top n target devices in the root cause ranking result to determine the fault type, generate a pre-diagnosis instruction set, and perform automated recovery operation.
2. The method for remote operation and maintenance and fault prediction of CPE devices based on cloud collaboration according to claim 1, characterized in that, Based on the target device, operational indicator data, base station identifier, and transmission link identifier as nodes, association edges are constructed between the nodes, and a heterogeneous association graph is constructed based on the nodes and association edges, including the following steps: Step 31: Define each target device as a target device node, define each type of operational indicator data as an indicator data node, define each base station identifier as a base station node, and define each transmission link identifier as a transmission link node. Step 32: Construct a first association edge based on the target device node and the base station node; construct a second association edge based on the base station node and the transmission link node; Step 33: Construct a third association edge based on the target device node and the indicator data node; construct a fourth association edge based on the base station node and the indicator data node; and construct a fifth association edge based on the relationship between the transmission link node and the indicator data node. Step 34: Based on the target device node, indicator data node, base station node, transmission link node, and the first, second, third, fourth, and fifth associated edges connecting the nodes, a heterogeneous association graph is constructed.
3. The method for remote operation and maintenance and fault prediction of CPE devices based on cloud collaboration according to claim 2, characterized in that, Multiple degradation indicators are constructed based on the operational indicator data within a continuous time window. Target base station nodes are identified as predicted fault points in the heterogeneous correlation graph based on these degradation indicators. Target device nodes directly connected to the predicted fault points are designated as a device cluster. The process includes the following steps: Step 41: Based on each target device node, obtain the operation index data within a continuous time window, calculate the first-order time difference for each type of operation index data as the degradation rate of each operation index data; multiply the degradation rate of all operation index data of the current target device node by a preset weight to obtain the degradation rate vector of the current target device node. Step 42: Calculate the mean of all elements in the degradation rate vector and record it as the scalar degradation intensity; Step 43: Based on the target device nodes connected to each base station node, sector groups are formed in conjunction with antenna sector configuration information. A weighted average is then performed based on the degradation rate vector of the target device nodes within each sector group to obtain a sector-level degradation representative vector. The sector-level degradation representative vectors of all sector groups are then concatenated in order to form a structured degradation state vector. Step 44: Perform sparse principal component analysis based on the structured deterioration state vector to obtain the principal mode direction vector; Step 45: Obtain multiple state parameter time series of each base station node within multiple consecutive time windows and filter the state master parameters; calculate the first-order time difference based on the state master parameter time series corresponding to the state master parameters and record it as the base station degradation rate; construct a multivariate time series matrix based on the base station degradation rate and the scalar degradation intensity of each target device node, wherein each row of the multivariate time series matrix corresponds to a time point, the first column of the multivariate time series matrix is the base station degradation rate at the current time point, and the remaining columns in the multivariate time series matrix are the scalar degradation intensity of each target device node at the current time point; Step 46: Calculate the Granger causality strength and reverse causality strength based on the base station degradation rate and each scalar degradation strength of the multivariate time series matrix; if the average Granger causality strength of the base station degradation rate for at least two target device nodes is greater than the reverse causality strength, then calculate the inner product of the main mode direction vector and the base station structured degradation state vector under the current time window to obtain the degradation main mode activation strength; use the degradation main mode activation strength as a degradation flow signal and send it from the base station node to each connected transmission link node; Step 47: Construct a link state feature vector based on the transmission link node; when the current degraded flow signal arrives at the transmission link node, concatenate the current degraded main mode activation intensity with the current link state feature vector to form a current joint feature vector; calculate the Mahalanobis distance between the current joint feature vector and each historical event in the preset memory bank to obtain multiple Mahalanobis distance values; if the minimum Mahalanobis distance exceeds the dynamic tolerance, determine that the current transmission link node is in an abnormal conduction state, block the continued downstream propagation of the degraded flow signal, and generate a propagation blocking binary marker; Step 48: For each base station node, calculate the proportion of transmission link nodes whose propagation of degraded flow signals is blocked to the total number of transmission link nodes; if the proportion exceeds a preset proportion threshold, it is determined that the degraded energy has locally converged at the current base station; and calculate the cosine distance between the main mode direction vector of the current base station and the main mode direction vector of the adjacent base stations; if the cosine distance is greater than the isolation threshold, it is determined that the degraded mode of the base station has topological isolation; mark the base station nodes that simultaneously satisfy the local convergence condition and the topological isolation condition as predicted fault points, and aggregate all directly connected target device nodes into a cluster of devices to be dealt with.
4. The method for remote operation and maintenance and fault prediction of CPE devices based on cloud collaboration according to claim 3, characterized in that, The link state feature vector includes real-time bit error rate, real-time throughput, real-time transmission latency, and real-time load rate.
5. The cloud-based collaborative remote operation and maintenance and fault prediction method for CPE devices according to claim 4, characterized in that, The dynamic tolerance is obtained based on the real-time bit error rate, real-time load rate, and preset basic tolerance: if the real-time load rate is less than 0.3 and the real-time bit error rate is less than 0.02, the dynamic tolerance is defined as the preset basic tolerance; if the real-time load rate is greater than 0.7 or the real-time bit error rate is greater than 0.1, the dynamic tolerance is defined as x times the preset basic tolerance, where x is between 0.1 and 0.3; otherwise, the dynamic tolerance is calculated based on the real-time bit error rate, real-time load rate, and preset basic tolerance.
6. The method for remote operation and maintenance and fault prediction of CPE devices based on cloud collaboration according to claim 5, characterized in that, Root cause analysis is performed on the device cluster to obtain root cause ranking results; based on the top n target devices in the root cause ranking results, composite judgment logic is executed to determine the fault type, and a pre-diagnostic instruction set is generated to execute automated recovery operations, including the following steps: Step 51: Mark each target device node in the device cluster as a candidate device node, mark the base station node connected to the candidate device node as a candidate base station node, and mark the propagation link node connected to the candidate base station node as a candidate link node; construct base station sector nodes based on the candidate device nodes connected to the candidate base station nodes; form a first type of edge based on the candidate device node pointing to its corresponding base station sector node; form a second type of edge based on the candidate device node pointing to the candidate link node it uses; form a third type of edge based on the candidate link node pointing to its corresponding base station sector node; construct a heterogeneous operation and maintenance graph based on the candidate device nodes, base station sector nodes, candidate link nodes, first type of edge, second type of edge, and third type of edge. Step 52: Define a first initial anomaly score for each candidate device node based on the heterogeneous operation and maintenance diagram, and define a second initial anomaly score for each base station sector node and candidate link node, wherein the second initial anomaly score is 0; Step 53: Iteratively perform deterministic anomaly propagation on the heterogeneous operation and maintenance graph. In each round, update the results based on the first initial anomaly score of each candidate device node to obtain the root cause score of each candidate device node, its associated base station sector node, and candidate link node. Step 54: Sort all candidate device nodes in descending order of root cause score to generate root cause ranking result; if the root cause score of the base station sector node to which the first n candidate device nodes belong in the root cause ranking result is greater than the first sector threshold, and the proportion of the number of target devices whose propagation link nodes within the base station sector node to the total number of target devices in the sector group is greater than the preset density threshold, then the fault type is determined to be antenna beamforming offset fault type. Step 55: If the root cause score of the transmission link node used by the top n candidate device nodes in the root cause sorting result is greater than the first link threshold, and the dynamic tolerance of the current transmission link node shows a monotonically decreasing trend in the most recent P consecutive time windows, then the fault type is determined to be the link carrying capacity degradation fault type. Step 56: If the root cause scores of the top n candidate device nodes in the root cause sorting results are greater than the preset root cause significance threshold, the root cause scores of the base station sector nodes to which they belong are less than or equal to the first sector threshold, and the root cause scores of the transmission link nodes used are less than or equal to the first link threshold, then the fault type is determined to be the target device local hardware abnormality fault type. Step 57, otherwise, determine the fault type as a composite abnormal fault type; Step 58: Match the predefined fault type template according to the determined fault type, generate the corresponding pre-diagnostic instruction set and execute it to perform automated recovery operation on the target device.
7. The cloud-based collaborative remote operation and maintenance and fault prediction method for CPE devices according to claim 6, characterized in that, Iteratively perform deterministic anomaly propagation on the heterogeneous operation and maintenance graph. In each round, update the results based on the first initial anomaly score of each candidate device node to obtain the root cause score of each candidate device node, as well as its associated base station sector node and transmission link node. This includes the following steps: Step 531: In each round, the initial anomaly scores of the candidate device nodes, base station sector nodes, and candidate link nodes in the heterogeneous operation and maintenance graph are iteratively updated: a. Based on the base station sector node, obtain the candidate device node and candidate link node of the incoming edge, denoted as the first incoming edge and the second incoming edge, respectively; determine the sum of the initial anomaly scores of all the first incoming edges of the base station sector node and multiply them by the first propagation coefficient, denoted as the first anomaly score; determine the sum of the initial anomaly scores of all the second incoming edges of the base station sector node and multiply them by the second propagation coefficient, denoted as the second anomaly score; calculate the sum of the first anomaly score and the second anomaly score as the base station sector anomaly score; b. Based on the candidate link node, obtain the candidate device node of the incoming edge, which is denoted as the third incoming edge. Calculate the sum of the initial anomaly scores of all third incoming edges and multiply it by the third propagation coefficient to obtain the third anomaly score, which is used as the anomaly score of the candidate link node. Step 532: If the number of propagation rounds already executed reaches the preset maximum number of rounds, and the maximum absolute change in the abnormal scores of all updated base station sector nodes and candidate link nodes in two consecutive propagation rounds is less than the preset convergence threshold, then stop the iterative propagation and proceed to the next step: Step 533: Calculate the root cause score for each target device based on the updated anomaly scores of the base station sector nodes, candidate link nodes, and the first initial anomaly score.
8. The method for remote operation and maintenance and fault prediction of CPE devices based on cloud collaboration according to claim 7, characterized in that, The first initial score is obtained by weighted summation of the scalar degradation intensity and the binary label of the current candidate device node.
9. The method for remote operation and maintenance and fault prediction of CPE devices based on cloud collaboration according to claim 7, characterized in that, The first propagation coefficient, the second propagation coefficient, and the third propagation coefficient satisfy the following condition: 0 < second propagation coefficient < first propagation coefficient < third propagation coefficient < 1.