B5G base station multi-dimensional health trend prediction method and system based on improved Transform

By improving the multidimensional health trend prediction method of Transformer, the problems of over-maintenance and under-maintenance in B5G base station operation and maintenance are solved, achieving efficient and accurate prediction of base station health trends, reducing resource consumption and improving fault detection capabilities.

CN121750127APending Publication Date: 2026-03-27BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

The existing B5G base station operation and maintenance model suffers from over-maintenance and under-maintenance, lacks effective online intelligent health trend prediction methods, resulting in resource waste and operational risks. Existing prediction models are difficult to adapt to the limited computing resources at the base station edge and capture early weak fault characteristics.

Method used

A multidimensional health trend prediction method based on an improved Transformer is adopted. By constructing a multidimensional health input sequence, a trend-sensitive sparse attention strategy and a dual-path feature fidelity distillation mechanism are introduced, combined with a cross-level health aggregation strategy, to achieve multidimensional health trend prediction for B5G base stations.

Benefits of technology

It reduces the computational complexity of long sequences, enhances the ability to capture non-stationary oscillation signals, reduces computational resource consumption, improves the sensitivity to faults and the foresight of system-level assessment, reduces the risk of error accumulation, and improves the accuracy and stability of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121750127A_ABST
    Figure CN121750127A_ABST
Patent Text Reader

Abstract

The invention relates to a B5G base station multi-dimensional health trend prediction method and system based on an improved Transform, and the method comprises the following steps: obtaining the historical health degree data of a B5G base station multi-dimensional functional circuit, and constructing a time-aligned multi-dimensional input tensor; a multi-dimensional trend prediction model is constructed, an encoder adopts a trend-sensitive sparse attention strategy, and a local volatility index is introduced to correct sparsity measurement so as to capture early weak oscillation features neglected by a traditional method; two-way feature fidelity distillation is executed between encoder levels, and reference information in long-sequence downsampling is prevented from being lost through a parallel local feature extraction path and a trend keeping path; utilizing a multi-dimensional full-connection decoder to synchronously output future multi-step predicted values; and executing cross-level aggregation based on a dynamic and static weight fusion strategy, and generating a system-level health trend. The method can effectively solve the problem that an existing model neglects high-frequency fault symptoms and error accumulation in long sequence modeling, and improves the accuracy and robustness of edge side prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent operation and maintenance technology for communication network equipment, specifically relating to a multi-dimensional health trend prediction method and system for B5G base stations based on an improved Transformer. Background Technology

[0002] With the development of 5G and B5G (Beyond 5G) communication technologies, the integration and complexity of base station hardware are increasing exponentially. Modern B5G base stations, especially AAU (Active Antenna Unit) and BBU (Baseband Processing Unit), integrate a massive number of radio frequency transceiver channels and high-precision digital signal processing circuits. These devices operate 24 / 7 under high-frequency, high-power conditions, and their health status directly affects the stability and reliability of the communication network.

[0003] However, current operation and maintenance management of B5G base stations mainly relies on reactive maintenance or preventative maintenance based on fixed cycles, lacking effective online intelligent health trend prediction methods. This leads to serious adverse consequences of over-maintenance and under-maintenance in practical engineering applications.

[0004] First, there is the waste of resources caused by over-maintenance: Traditional periodic maintenance often ignores the actual health differences of individual equipment. Even if some functional circuits are still in good condition, they may be forcibly replaced within the same maintenance cycle, resulting in a huge waste of spare parts resources and maintenance manpower.

[0005] Secondly, there are operational risks caused by insufficient maintenance: B5G base station failures are often sudden and hidden. Between two scheduled maintenance periods, if key components experience performance degradation, such as an increase in nonlinear distortion in the radio frequency link, existing monitoring systems based on simple threshold alarms can only detect the problem when the indicators deteriorate completely and cause service interruption. This delayed monitoring makes it impossible to predict and intervene in the failure in advance, which can easily lead to base station outages or even large-scale network paralysis. Summary of the Invention

[0006] To address the technical shortcomings of over-maintenance and under-maintenance in existing B5G base station operation and maintenance models, and the technical problems of existing general prediction models being difficult to adapt to the limited computing resources at the base station edge and difficult to capture early weak fault characteristics, this invention provides a multi-dimensional health trend prediction method for B5G base stations based on an improved Transformer. This prediction method can be executed by one or more processors and includes the following steps:

[0007] S1: The one or more processors construct the multidimensional health input sequence for the B5G base station, including the following steps:

[0008] S11: Obtain historical health monitoring data of multiple functional circuits in the B5G base station;

[0009] S12: Based on the historical health monitoring data of multiple key functional circuits in the B5G base station, and with a preset time granularity as the benchmark, the monitoring data of each functional circuit is timestamped and normalized.

[0010] S13: Abstract the dispersed circuit monitoring data into a multidimensional input tensor;

[0011] S2: The one or more processors construct a multidimensional health trend prediction model for B5G base stations and configure a sparse attention strategy based on trend-sensitive weighting in the Transformer encoder. When calculating the sparsity metric of the query vector and the key vector, a local volatility weighting factor is introduced. Explicit query vectors with significant volatility characteristics are preferentially selected for dot product attention calculation. This reduces the computational complexity of long sequences while enhancing the ability to capture non-stationary oscillation signals of various functional circuits of the B5G base station.

[0012] S3: The one or more processors establish a dual-path feature-fidelity self-attention distillation mechanism between model levels: on the basis of the main path that uses one-dimensional convolution and max pooling to extract significant features, a parallel average pooling residual path is added to fuse significant fault features and baseline aging trend features, and a pyramid-shaped feature extraction structure is constructed to prevent the loss of health baseline information during long-cycle downsampling.

[0013] S4: The one or more processors train and infer the B5G base station multidimensional health trend prediction model; in the B5G base station multidimensional health trend prediction model, the decoder is configured as a multidimensional fully connected output structure, which maps the deep features extracted by the encoder into a multidimensional prediction matrix, and outputs the predicted values ​​of the health of all functional circuits at the next H time points at one time to reduce the cumulative impact of errors caused by stepwise regression.

[0014] S5: The one or more processors execute a cross-level health aggregation strategy, including the following steps:

[0015] S51: Construct a hierarchical association model for B5G base stations, including functional circuits, subsystems, and systems;

[0016] S52: Using dynamic weights and dynamic weight fusion strategies, subsystem-level aggregation and system-level aggregation are performed sequentially: the subsystem-level aggregation aggregates the predicted health values ​​of each functional circuit level into the health trend of the B5G base station subsystem level; the system-level aggregation aggregates the health trends of the B5G base station subsystem level into the health trend of the B5G base station system level, generating a prediction sequence that reflects the overall future evolution trend of the B5G base station.

[0017] Preferably, step S13 includes the following steps:

[0018] S131: Abstract the key functional circuits in the B5G base station into a set of feature dimensions in the input tensor;

[0019] S132: Abstract the historical monitoring time window into a set of time steps in the input tensor;

[0020] S133: Construct a multidimensional input sequence of shape B·W·M based on the relationship between functional circuits, time steps, and position codes, where: M is the number of functional circuits; W is the length of the historical time window; and B is the batch size.

[0021] Furthermore, S13 also includes the following steps:

[0022] S134: Calculate the location code PE of the B5G base station's multidimensional health input sequence using the following formula. (pos,2i) and PE (pos,2i+1) To preserve the temporal dependency properties of the sequence:

[0023]

[0024] In the formula: pos is the position index; i is the dimension index; d model Embed the dimension of the model's features;

[0025] S135: To address the differences in the numerical distribution characteristics of health scores across different circuits, the original health score data x is Z-score standardized using the following formula to obtain standardized data x':

[0026]

[0027] In the formula: μ is the mean of the historical data of this functional circuit, and σ is the standard deviation;

[0028] S136: Combine the standardized data x' with the location code PE (pos,2i) and PE (pos,2i+1) The sum is used as the model input.

[0029] Furthermore, the sparse attention strategy in S2 adopts a trend-sensitive probabilistic sparse mechanism: in the trend-sensitive weighted sparse attention strategy, the i-th query vector q is calculated. i An improved sparsity metric M(q) for the set of key vectors K i The improved sparsity metric M(q), K), i M(q) is based not only on the difference in probability distribution, but also on the local fluctuation characteristics of the query vector, therefore M(q) i The formula for calculating K is derived from...

[0030]

[0031] Improved to:

[0032]

[0033] In the formula: L is the sequence length; d is the feature dimension; k j V is the j-th key vector; i For query vector q i Local volatility index within a time window; λ is the trend sensitivity coefficient; T is the time window;

[0034] By introducing V i This enhances the predictive model's ability to capture the nonlinear fluctuation characteristics of the B5G base station's functional circuits under high load, avoiding the standard sparse attention method from filtering out weak oscillation signals in the early stages of a fault as noise; finally, based on the weighted M(q)... i Sort the query vectors by K and filter the top u explicit query vectors; when the improved sparsity metric is greater than a preset threshold, the query vector is determined to be an explicit query and participates in the calculation.

[0035] Furthermore, S2 also includes the calculation of the sparse attention mechanism, using the selected explicit query matrix Q, key matrix K, and value matrix V to calculate the attention weights.

[0036]

[0037] In the formula: Attn is the attention weight function; Softmax is the normalization exponential function;

[0038] For the decoder output, a multi-dimensional fully connected layer is used to directly map and generate the prediction results for the next H steps.

[0039]

[0040] Where: Where: The matrix represents the health prediction results of each functional circuit in the next H steps; Linear represents the mapping function of the multidimensional fully connected layer; Decoder_Output represents the deep feature sequence output by the decoder; H represents the future prediction step size; and M represents the number of functional circuits. This indicates that the output is a real matrix with shape H rows and M columns.

[0041] Furthermore, the self-attention distillation mechanism in S3 employs a dual-path feature-fidelity distillation structure to address the loss of health baseline information caused by traditional max-pooling operations. The dual-path feature-fidelity distillation structure includes the following steps:

[0042] S31: The output feature sequence X of the j-th layer encoder... j As input, it is processed in two paths:

[0043] Main path: Extract local context through one-dimensional convolution operation, transform it using ELU activation function, and then use max pooling layer MaxPool for downsampling to extract significant fault features;

[0044] Residual path: for X j Perform average pooling (AvgPool) to preserve the health baseline trend information of the functional circuit;

[0045] S32: Input X of the (j+1)th layer j+1 The feature is determined through dual-path feature fusion, as shown in the following formula:

[0046] X j+1 = Norm(MaxPool(ELU(Conv1d(X) j )))+AvgPool(X j ))

[0047] In the formula: X j+1 X is the input feature sequence of the (j+1)th layer encoder; j denoted as the output feature sequence of the j-th layer encoder; Norm is the normalization operation; MaxPool is the max pooling operation used to extract salient features; AvgPool is the average pooling operation used to preserve the average trend information of the sequence; ELU is the exponential linear unit activation function; Conv1d is the one-dimensional convolution operation used to extract local contextual features.

[0048] Furthermore, the cross-level health aggregation strategy in S5 represents the overall operating status of the base station by weighted aggregation of multi-dimensional prediction results. The aggregation strategy includes an aggregation strategy based on static functional weights, a dynamic weight strategy based on the amount of information in the prediction sequence, and a step-by-step aggregation strategy based on multi-dimensional weight fusion.

[0049] The aggregation strategy based on static functional weights includes: constructing a judgment matrix using the Analytic Hierarchy Process (AHP) and calculating the relative importance weight of each functional circuit within its respective subsystem. And the relative importance weight λ of each subsystem in the overall B5G base station. sub The subsystem includes at least an active antenna unit (AAU) and a baseband processing unit (BBU).

[0050] The dynamic weighting strategy based on the information content of the predicted sequence includes: calculating the information entropy of the predicted sequence of each functional circuit based on the inverse entropy weighting method, and then calculating the dynamic information weight. To reflect the impact of predicted sequence fluctuations on the system;

[0051] The hierarchical aggregation strategy based on multi-dimensional weight fusion includes: subsystem-level aggregation and system-level aggregation. First, subsystem-level aggregation is performed: combining the static functional weights... With the dynamic information weight Calculate the overall weight of each functional circuit. Using the aforementioned comprehensive weights The predicted health values ​​of each functional circuit level are aggregated into the health trend of the B5G base station subsystem; then, system-level aggregation is performed: based on the relative importance weight λ of each subsystem in the overall B5G base station. sub The health trends of the B5G base station subsystem are aggregated into the health trends of the B5G base station system.

[0052] Furthermore, the calculation process for cross-level health aggregation includes the following steps:

[0053] S521: For the k-th functional circuit, the information entropy E of the prediction sequence within the future prediction step size H is calculated. k Construct dynamic information weights, where information entropy is defined as:

[0054]

[0055] In the formula: E k p is the information entropy of the predicted sequence of the k-th functional circuit within the future prediction step size H; kt Let be the normalized probability of the predicted value of the k-th circuit at a future time t;

[0056] S522: According to E k Calculate the dynamic information weight of each functional circuit. The dynamic information weight is defined as follows: This increases the weight when the predicted sequence has large fluctuations or low entropy.

[0057]

[0058] In the formula: E represents the dynamic information weight of the k-th functional circuit; k M represents the information entropy of the predicted sequence within the future prediction step size H for the k-th functional circuit; sub The total number of functional circuits contained within the subsystem AAU or BBU to which the functional circuit belongs;

[0059] S523: Combining the aforementioned static function weights With the dynamic information weight Calculate the overall weight of each functional circuit.

[0060]

[0061] In the formula: α is the balance coefficient, which is used to adjust the proportion of static functional weight and dynamic information weight in the comprehensive evaluation. Its value range is 0≤α≤1. The comprehensive weight of the k-th functional circuit; These are static function weights; For dynamic information content weights;

[0062] S524: Perform the subsystem-level aggregation: using the comprehensive weights The health trends of the AAU and BBU subsystems are calculated. AAU (t) and S BBU (t), where Ω AAU and Ω BBU These are the functional circuit sets belonging to AAU and BBU, respectively. k (t) represents the circuit prediction value:

[0063]

[0064] S525: Perform the system-level aggregation: Calculate the base station system-level health trend S based on the subsystem weights λ. total (t):

[0065] S total (t)=λ AAU ·S AAU (t)+λ BBU ·S BBU (t).

[0066] To address the aforementioned technical problems, this invention also provides a B5G base station multidimensional health trend prediction system based on an improved Transformer, used to implement any of the above preferred embodiments of the B5G base station multidimensional health trend prediction method based on an improved Transformer, comprising a multidimensional health input sequence construction module, a trend-sensitive sparse attention encoding module, a dual-path feature fidelity distillation module, a multidimensional synchronous decoding prediction module, and a cross-level health aggregation module.

[0067] The multidimensional health input sequence construction module is configured as follows: first, acquire historical health monitoring data of multiple functional circuits in the B5G base station; second, perform timestamp alignment and normalization processing on the monitoring data of each functional circuit; and finally, construct a multidimensional input tensor.

[0068] The trend-sensitive sparse attention encoding module is configured to perform trend-sensitive weighted sparse attention calculation in the Transformer encoder, wherein the sparsity metric introduces a local volatility weighting factor to filter query vectors participating in the attention calculation.

[0069] The dual-path feature fidelity distillation module is configured as follows: a dual-path distillation structure is established with main path downsampling and residual path downsampling. Convolution and max pooling are performed on the main path, and average pooling is performed on the residual path. The dual-path outputs are then fused to form a pyramid-shaped feature.

[0070] The multidimensional synchronous decoding prediction module is configured to: map the encoded features into a multidimensional prediction matrix through a decoder, and synchronously output the predicted health values ​​of each functional circuit at H future time points;

[0071] The cross-level health aggregation module is configured as follows: based on a hierarchical association model including functional circuit-subsystem-system, the predicted values ​​at the functional circuit level are aggregated at the subsystem level and then at the system level to obtain the base station system-level health trend prediction sequence.

[0072] This invention addresses the gaps in existing intelligent operation and maintenance technologies for B5G base stations. To solve the industry challenges of over-maintenance and under-maintenance, it proposes a multi-dimensional health trend prediction method for B5G base stations that balances low computational power consumption at the edge, sensitive capture of weak fault features, and system-level dynamic evaluation. This prediction method solves the problem of lost key fluctuation features caused by sparse sampling by constructing a trend-sensitive sparse attention mechanism and a dual-path feature-fidelity distillation structure. Simultaneously, it combines a hierarchical association model of B5G base stations, including functional circuits, subsystems, and the system, to dynamically aggregate the prediction results across levels, generating a system-level health trend prediction sequence to assist operation and maintenance personnel in trend assessment and early warning decisions.

[0073] Compared with the prior art, the beneficial effects of the present invention are:

[0074] 1. Reduced overhead for long sequence modeling and improved deployability at the edge: This invention introduces a sparse attention mechanism in the encoder and introduces a sequence downsampling distillation structure based on convolution and pooling between layers to reduce the overhead of attention computation and intermediate feature storage in long sequence scenarios, thereby improving the modeling capability for ultra-long historical monitoring sequences and the feasibility of deployment at the edge computing unit.

[0075] 2. Multi-step, multi-dimensional synchronous output and error accumulation suppression: This invention uses a multi-dimensional fully connected output structure to map decoder features and outputs the multi-functional circuit health prediction results for the next H time points at once, reducing the risk of error accumulation caused by step-by-step iterative prediction and improving the stability of long-term prediction;

[0076] 3. Enhanced ability to utilize information coupled by multiple circuits: This invention uses a multidimensional health sequence as a joint modeling object to simultaneously characterize the change process of the health of multiple functional circuits over time in the same model, enabling the model to learn the correlation features between different circuits, thereby improving the trend representation ability in multi-circuit coupled fault scenarios.

[0077] 4. Dynamic aggregation and system-level trend representation for predicted sequences: This invention constructs dynamic information weights for future predicted sequences and integrates them with static functional importance weights to form a step-by-step aggregation strategy of circuit-subsystem-system. This enables the system-level health trend to adaptively adjust to the fluctuation characteristics of future predicted sequences, thereby improving the sensitivity to trend risks and the foresight of system-level assessment.

[0078] 5. Reduced training and inference resource consumption: This invention combines sparse attention filtering with hierarchical downsampling to reduce the number of queries involved in attention calculation and reduce the length of inter-layer sequences, thereby reducing the computational load and intermediate feature storage overhead during the training and inference stages and improving the model's operating efficiency in resource-constrained environments.

[0079] The breakthroughs and comparisons of the B5G base station trend prediction modeling method based on multifunctional circuit health sequences, which is involved in this invention, with existing technologies are shown in Table 1:

[0080] Table 1:

[0081]

[0082] Attached Figure Description

[0083] Figure 1 This is a schematic diagram of the overall process of a preferred embodiment of the B5G base station multidimensional health trend prediction method based on the improved Transformer of the present invention.

[0084] Figure 2 This is a schematic diagram of the alignment and normalization of historical health monitoring data of a multi-functional circuit under a unified timestamp in a preferred embodiment of the multi-dimensional health trend prediction method for B5G base stations based on the improved Transformer of the present invention.

[0085] Figure 3 This is a schematic diagram of the overall architecture of the improved Transformer multidimensional trend prediction model in a preferred embodiment of the improved Transformer-based B5G base station multidimensional health trend prediction method of the present invention.

[0086] Figure 4 This is a flowchart illustrating the dual-path feature-fidelity self-attention distillation mechanism in a preferred embodiment of the improved Transformer-based B5G base station multidimensional health trend prediction method of the present invention.

[0087] Figure 5 This is a schematic diagram of the future health trend prediction results of some functional circuits in a preferred embodiment of the B5G base station multidimensional health trend prediction method based on the improved Transformer of the present invention.

[0088] Figure 6 This is a schematic diagram of the health trend prediction curve of the AAU subsystem in a preferred embodiment of the B5G base station multidimensional health trend prediction method based on the improved Transformer of the present invention.

[0089] Figure 7 This is a schematic diagram of the health trend prediction curve of the BBU subsystem in a preferred embodiment of the multidimensional health trend prediction method for B5G base stations based on the improved Transformer of the present invention.

[0090] Figure 8 This is a schematic diagram of the overall health trend prediction curve of the B5G base station system level obtained by weighted aggregation of the prediction results of each subsystem in a preferred embodiment of the multidimensional health trend prediction method for B5G base stations based on the improved Transformer of the present invention.

[0091] Figure 9 This is a schematic diagram comparing the error indices of the model of the present invention and the baseline model on a long-sequence trend prediction task in a preferred embodiment of the multi-dimensional health trend prediction method for B5G base stations based on the improved Transformer of the present invention. Detailed Implementation

[0092] The present invention will be further described in detail below with reference to the accompanying drawings. This detailed description is based on exemplary embodiments of the invention, including various details that aid in understanding the invention, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the invention.

[0093] The proposed method and system for predicting the multidimensional health trend of B5G base stations based on the improved Transformer aims to solve the memory bottleneck and computing power barrier problems faced by existing technologies when processing long sequences and multidimensional coupled data, and to build a comprehensive dynamic perception capability from micro-circuits to macro-systems.

[0094] like Figure 1 As shown, the preferred embodiment of this method for predicting the multidimensional health trend of B5G base stations based on an improved Transformer includes the following steps:

[0095] Step 1: Constructing the multidimensional health input sequence for B5G base stations: This step aims to transform the scattered and heterogeneous raw monitoring data within B5G base stations into standard tensors suitable for deep learning model processing;

[0096] Multi-source data acquisition and time-series alignment: First, key functional circuits with strong physical correlation within the B5G base station are selected as monitoring objects. In this embodiment, there are a total of 12 types of monitoring objects, specifically including:

[0097] AAU side: power amplifier, filter, T / R module, optical module, low noise amplifier, AAU power supply;

[0098] BBU side: Digital SOC, FLASH memory, phase-locked loop, BBU power supply, relay, crystal oscillator;

[0099] Since the data sampling frequencies of different sensors are inconsistent, this embodiment uses a preset time granularity as a benchmark to perform spline interpolation on the historical health sequence of each circuit to ensure that it is strictly aligned on the time axis.

[0100] Figure 2 This demonstrates the changes in the health sequence of each functional circuit on a unified time axis after timing alignment. Figure 2 It can be seen that the originally sparsely sampled AAU radio frequency data and the densely sampled BBU digital signal data have been uniformly mapped to the same time step, thus constructing the data foundation for multidimensional coupling analysis.

[0101] Data standardization: To eliminate differences in the distribution characteristics of health values ​​among different circuits, such as different baseline values ​​and different fluctuation amplitudes, the health value x of the m-th circuit at time t is standardized. t,mZ-score normalization is performed to obtain x′ t,m :

[0102]

[0103] In the formula, μ m and σ m These are the mean and standard deviation of the circuit's historical data, respectively. This step unifies all input data into a standard distribution with a mean of 0 and a variance of 1, allowing the model to focus on the relative trend of health status.

[0104] Dataset partitioning and input tensor construction: In order to verify the generalization ability of the model and prevent data leakage, this embodiment strictly follows the chronological order of processing. In this embodiment, the original sequence length is 500, and a sliding window mechanism without mixed cross-boundary windows is adopted; the historical observation window length W = 30 and the prediction step size H = 7 are set.

[0105] The dataset is strictly divided according to chronological order, with the ratio of training, validation, and test sets set to 7:1:2 (70% training, 10% validation, and 20% test). After sliding window processing, the final input tensor shape is as follows: Where B is the batch size, preferably B=32, and feature dimension M=12;

[0106] Positional encoding: To preserve the temporal dependency information of the sequence, a positional encoding vector (PE) is introduced and superimposed on the input data. The positional encoding calculation formula is as follows:

[0107]

[0108] In the formula, pos is the position index; i is the dimension index; d model Embed the dimension of the model's features;

[0109] Step 2: Construct an improved Transformer prediction model:

[0110] like Figure 3 As shown, the prediction model constructed in this invention adopts an encoder-decoder architecture. In response to the dual characteristics of B5G base station data, which are characterized by early faults exhibiting high-frequency weak oscillations and aging exhibiting long-term slow drift, this embodiment has made targeted modifications to the standard Transformer.

[0111] Configure trend-sensitive sparse self-attention: to address the standard Transformer's computational cost of O(K). 2 To address the problem of sparse attention and avoid the shortcomings of traditional sparse attention which ignores weak fluctuation characteristics based solely on probability distribution, this embodiment introduces a local volatility weighting factor.

[0112] Calculate query vector qi The sparsity metric M(q) of the key vector set K i When K), the following improved formula is used:

[0113]

[0114] Among them, V i For query vector q i The local volatility index within the time window, calculated as q in this embodiment, is... i The variance within the neighborhood; λ is the trend sensitivity coefficient, which is 0.5 in this embodiment;

[0115] Physical meaning: When a functional circuit, such as an RF power amplifier, exhibits early nonlinear distortion, its health curve often shows high-frequency oscillations without a significant change in the mean. Traditional KL divergence tends to treat this as noise and ignore it. However, when (1+λ·V) is introduced... i After that, the model will improve the sparsity score of such vectors, making them selected as explicit query vectors;

[0116] Based on this metric, only the top u explicit query vectors are selected to participate in the dot product attention operation, while the attention outputs of the remaining lazy query vectors are directly filled into the mean. This mechanism not only ensures low complexity of O(LlogL) but also significantly improves the sensitivity to fault symptoms.

[0117] Step 3: Establish a self-attention distillation mechanism:

[0118] In order to extract long-term health evolution features and prevent the loss of reference information reflecting the aging trend of the equipment during downsampling, this embodiment constructs a dual-path feature fidelity distillation structure between encoder levels.

[0119] like Figure 4 As shown, this mechanism involves two parallel paths:

[0120] Main path (feature extraction): Includes one-dimensional convolutional layers (Conv1d), ELU activation function, and max pooling layer (MaxPool) to extract significant mutation features;

[0121] Residual path (trend preservation): includes an average pooling layer (AvgPool) to preserve the average trend information of the sequence;

[0122] The input X of the (j+1)th layer j+1 Output X from layer j j The result obtained through dual-path fusion is:

[0123] X j+1 = Norm(MaxPool(ELU(Conv1d(X) j)))+AvgPool(X j ))

[0124] Implementation Results: In this embodiment, the sequence length is halved by using a pooling operation with a step size of 2. Compared to the traditional distillation that only uses MaxPool (which is prone to losing the mean information representing the health baseline), the dual-path structure ensures that the model in the deep network can still remember the overall aging level of the circuit, thereby constructing a pyramid-shaped robust feature extraction structure.

[0125] Step 4: Model Training and Multidimensional Simultaneous Prediction

[0126] Multidimensional fully connected output: In this embodiment, the decoder adopts a generative structure, removing the stepwise regression output method of the standard Transformer and replacing it with a multidimensional fully connected layer;

[0127] This fully connected layer directly maps the deep features extracted by the decoder into a prediction matrix of dimension H×M, and synchronously outputs the predicted health values ​​of the next H time steps at once. In this embodiment, H=7 and M=12. The output tensor shape is... Figure 5 The model's predictions of the future health of some key functional circuits (such as power amplifiers and optical modules) are shown to fit the actual values. This non-autoregressive output method cuts off the accumulation link of single-step prediction errors and improves the accuracy of long-step prediction.

[0128] Step 5: Implement cross-level health aggregation strategy:

[0129] In order to achieve the perception of macroscopic system status from micro-circuit prediction, this embodiment constructs a cross-level aggregation model of functional circuit → subsystem → base station system, and adopts a weighting strategy that combines dynamic and static elements.

[0130] Fusion of dynamic and static weights:

[0131] Static function weight w func Based on the physical structure of the base station and expert experience, the inherent importance of each circuit and subsystem in the system is calculated using the Analytic Hierarchy Process (AHP).

[0132] Dynamic information content weight w info : Analyzing the volatility of future prediction sequences using the inverse entropy weight method;

[0133] Calculate the information entropy E of the predicted sequence of the k-th circuit in the next H steps. k This leads to the information content weight:

[0134]

[0135] This mechanism ensures that when the prediction curve of a circuit fluctuates drastically (indicating potential failure risk and a decrease in entropy value), its weight in the system evaluation will automatically increase.

[0136] Overall Weighting: The final weight is calculated by combining static and dynamic weights.

[0137]

[0138] Hierarchical aggregation calculation and result analysis:

[0139] First-level aggregation (subsystem-level aggregation): Calculate the health trends of the AAU and BBU subsystems separately using comprehensive weights:

[0140]

[0141] Figure 6 and Figure 7 The health trend prediction curves of the AAU subsystem and the BBU subsystem are shown respectively. As can be seen from the figure, the AAU subsystem shows a relatively obvious downward trend due to the aging characteristics of the radio frequency devices, while the BBU subsystem remains relatively stable. This proves that hierarchical aggregation can accurately locate the specific subsystem source that causes the decline in system health.

[0142] Second-level aggregation (system-level aggregation): Calculates the overall health trend of the base station based on the subsystem weight λ.

[0143] S total (t)=λ AAU S AAU (t)+λ BBU S BBU (t)

[0144] Figure 8 The system displays a trend prediction chart of the overall health of the base station system after weighted aggregation. This curve comprehensively reflects the overall operational risk of the base station, showing the main impact of AAU-side decay on the system, and avoiding interference from single circuit noise through weighted fusion, providing operation and maintenance personnel with an intuitive macro decision-making basis.

[0145] Performance Verification: To verify the effectiveness of the method of the present invention, the model of the present invention was compared with the existing mainstream time series prediction models on the same test set;

[0146] Specifically, such as Figure 9As shown, in order to verify the effectiveness of the improved method proposed in this invention, a number of mainstream time series prediction models, including Informer, LSTM, standard Transformer, GCN-LSTM and TCN, were selected as baselines and compared with the model (Ours) proposed in the embodiments of this invention on the same test set. The bar chart in the figure comprehensively shows the performance of the model on three key performance indicators: mean squared error (MSE), mean absolute error (MAE) and symmetric mean absolute percentage error (sMAPE, corresponding to the right axis).

[0147] The model of this invention performs optimally across all error metrics, such as... Figure 9 As shown, the model of this invention (Ours) is at the lowest level in all three indices: MSE, MAE, and sMAPE. The sMAPE index is also significantly lower than that of the comparative model, which indicates that the prediction accuracy and stability of the model of this invention have significant advantages.

[0148] The improvement effect of this invention overcoming the shortcomings of existing technologies is verified by comparing it with the Informer model. Although the Informer model uses a probabilistic sparse attention mechanism to adapt to long sequence prediction, Figure 9 The results show that its error is still significantly higher than that of the model in this invention, which corroborates the analysis in the background of this invention, namely that Informer's sparsity mechanism based on KL divergence easily filters out key high-frequency weak fluctuation features in B5G base station data, thus leading to prediction bias.

[0149] To address the aforementioned technical problems, this invention also provides a B5G base station multidimensional health trend prediction system based on an improved Transformer, used to implement any of the above preferred embodiments of the B5G base station multidimensional health trend prediction method based on an improved Transformer, comprising a multidimensional health input sequence construction module, a trend-sensitive sparse attention encoding module, a dual-path feature fidelity distillation module, a multidimensional synchronous decoding prediction module, and a cross-level health aggregation module.

[0150] The multidimensional health input sequence construction module is configured as follows: first, acquire historical health monitoring data of multiple functional circuits in the B5G base station; second, perform timestamp alignment and normalization processing on the monitoring data of each functional circuit; and finally, construct a multidimensional input tensor.

[0151] The trend-sensitive sparse attention encoding module is configured to perform trend-sensitive weighted sparse attention calculation in the Transformer encoder, wherein the sparsity metric introduces a local volatility weighting factor to filter query vectors participating in the attention calculation.

[0152] The dual-path feature fidelity distillation module is configured as follows: a dual-path distillation structure is established with main path downsampling and residual path downsampling. Convolution and max pooling are performed on the main path, and average pooling is performed on the residual path. The dual-path outputs are then fused to form a pyramid-shaped feature.

[0153] The multidimensional synchronous decoding prediction module is configured to: map the encoded features into a multidimensional prediction matrix through a decoder, and synchronously output the predicted health values ​​of each functional circuit at H future time points;

[0154] The cross-level health aggregation module is configured as follows: based on a hierarchical association model including functional circuit-subsystem-system, the predicted values ​​at the functional circuit level are aggregated at the subsystem level and then at the system level to obtain the base station system-level health trend prediction sequence.

[0155] This invention can effectively solve the problem of existing models ignoring high-frequency fault signs and error accumulation in long sequence modeling, and improve the accuracy and robustness of edge prediction.

[0156] Contents not described in detail in this specification are existing technologies known to those skilled in the art. The above descriptions are merely preferred embodiments of the present invention, and the scope of protection of the present invention is not limited to the above embodiments. Equivalent modifications or variations made by those skilled in the art based on the disclosure of the present invention should be included within the scope of protection of the claims.

Claims

1. A multidimensional health trend prediction method for B5G base stations based on an improved Transformer, characterized in that, The prediction method can be executed by one or more processors and includes the following steps: S1: The one or more processors construct the multidimensional health input sequence for the B5G base station, including the following steps: S11: Obtain historical health monitoring data of multiple functional circuits in the B5G base station; S12: Based on the historical health monitoring data of multiple key functional circuits in the B5G base station, and with a preset time granularity as the benchmark, the monitoring data of each functional circuit is timestamped and normalized. S13: Abstract the dispersed circuit monitoring data into a multidimensional input tensor; S2: The one or more processors construct a multidimensional health trend prediction model for B5G base stations and configure a sparse attention strategy based on trend-sensitive weighting in the Transformer encoder. When calculating the sparsity metric of the query vector and the key vector, a local volatility weighting factor is introduced. Explicit query vectors with significant volatility characteristics are preferentially selected for dot product attention calculation. This reduces the computational complexity of long sequences while enhancing the ability to capture non-stationary oscillation signals of various functional circuits of the B5G base station. S3: The one or more processors establish a dual-path feature-fidelity self-attention distillation mechanism between model levels: on the basis of the main path that uses one-dimensional convolution and max pooling to extract significant features, a parallel average pooling residual path is added to fuse significant fault features and baseline aging trend features, and a pyramid-shaped feature extraction structure is constructed to prevent the loss of health baseline information during long-cycle downsampling. S4: The one or more processors train and infer the B5G base station multidimensional health trend prediction model; in the B5G base station multidimensional health trend prediction model, the decoder is configured as a multidimensional fully connected output structure, which maps the deep features extracted by the encoder into a multidimensional prediction matrix, and outputs the predicted values ​​of the health of all functional circuits at the next H time points at one time to reduce the cumulative impact of errors caused by stepwise regression. S5: The one or more processors execute a cross-level health aggregation strategy, including the following steps: S51: Construct a hierarchical association model for B5G base stations, including functional circuits, subsystems, and systems; S52: Using dynamic weights and dynamic weight fusion strategies, subsystem-level aggregation and system-level aggregation are performed sequentially: the subsystem-level aggregation aggregates the predicted health values ​​of each functional circuit level into the health trend of the B5G base station subsystem level; the system-level aggregation aggregates the health trends of the B5G base station subsystem level into the health trend of the B5G base station system level, generating a prediction sequence that reflects the overall future evolution trend of the B5G base station.

2. The method for predicting the multidimensional health trend of B5G base stations based on the improved Transformer as described in claim 1, characterized in that, S13 includes the following steps: S131: Abstract the key functional circuits in the B5G base station into a set of feature dimensions in the input tensor; S132: Abstract the historical monitoring time window into a set of time steps in the input tensor; S133: Construct a multidimensional input sequence of shape B·W·M based on the relationship between functional circuits, time steps, and position codes, where: M is the number of functional circuits; W is the length of the historical time window; and B is the batch size.

3. The method for predicting the multidimensional health trend of B5G base stations based on the improved Transformer as described in claim 2, characterized in that, S13 further includes the following steps: S134: Calculate the location code PE of the B5G base station's multidimensional health input sequence using the following formula. (pos,2i) and PE (pos,2i+1) To preserve the temporal dependency properties of the sequence: In the formula: pos is the position index; i is the dimension index; d model Embed the dimension of the model's features; S135: To address the differences in the numerical distribution characteristics of health scores across different circuits, the original health score data x is Z-score standardized using the following formula to obtain standardized data x': In the formula: μ is the mean of the historical data of this functional circuit, and σ is the standard deviation; S136: Combine the standardized data x' with the location code PE (pos,2i) and PE (pos,2i+1) The sum is used as the model input.

4. The method for predicting the multidimensional health trend of B5G base stations based on the improved Transformer as described in claim 1, characterized in that, The sparse attention strategy in S2 adopts a trend-sensitive probabilistic sparse mechanism: in the trend-sensitive weighted sparse attention strategy, the i-th query vector q is calculated. i An improved sparsity metric M(q) for the set of key vectors K i The improved sparsity metric M(q), K), i M(q) is based not only on the difference in probability distribution, but also on the local fluctuation characteristics of the query vector, therefore M(q) i The formula for calculating K is derived from... Improved to: In the formula: L is the sequence length; d is the feature dimension; k j V is the j-th key vector; i For query vector q i Local volatility index within a time window; λ is the trend sensitivity coefficient; T represents the time window; By introducing V i This enhances the predictive model's ability to capture the nonlinear fluctuation characteristics of the B5G base station's functional circuits under high load, avoiding the standard sparse attention method from filtering out weak oscillation signals in the early stages of a fault as noise; finally, based on the weighted M(q)... i Sort the query vectors by K and filter the top u explicit query vectors; when the improved sparsity metric is greater than a preset threshold, the query vector is determined to be an explicit query and participates in the calculation.

5. The method for predicting the multidimensional health trend of B5G base stations based on the improved Transformer as described in claim 1, characterized in that, S2 also includes computation for the sparse attention mechanism, utilizing the selected explicit query matrix. The attention weights are calculated using the key matrix K and the value matrix V. In the formula: Attn is the attention weight function; Softmax is the normalization exponential function; For the decoder output, a multi-dimensional fully connected layer is used to directly map and generate the prediction results for the next H steps. In the formula: The matrix represents the health prediction results of each functional circuit in the next H steps; Linear represents the mapping function of the multidimensional fully connected layer; Decoder_Output represents the deep feature sequence output by the decoder; H represents the future prediction step size; and M represents the number of functional circuits. This indicates that the output is a real matrix with shape H rows and M columns.

6. The method for predicting the multidimensional health trend of B5G base stations based on the improved Transformer as described in claim 1, characterized in that, The self-attention distillation mechanism in S3 employs a dual-path feature-fidelity distillation structure to address the loss of health baseline information caused by traditional max pooling operations. The dual-path feature-fidelity distillation structure includes the following steps: S31: The output feature sequence X of the j-th layer encoder... j As input, it is processed in two paths: Main path: Extract local context through one-dimensional convolution operation, transform it using ELU activation function, and then use max pooling layer MaxPool for downsampling to extract significant fault features; Residual path: for X j Perform average pooling (AvgPool) to preserve the health baseline trend information of the functional circuit; S32: Input X of the (j+1)th layer j+1 The feature is determined through dual-path feature fusion, as shown in the following formula: X j+1 =Norm(MaxPool(ELU(Conv1d(X j )))+AvgPool(X j )) In the formula: X j+1 X is the input feature sequence of the (j+1)th layer encoder; j denoted as the output feature sequence of the j-th layer encoder; Norm is the normalization operation; MaxPool is the max pooling operation used to extract salient features; AvgPool is the average pooling operation used to preserve the average trend information of the sequence; ELU is the exponential linear unit activation function; Conv1d is the one-dimensional convolution operation used to extract local contextual features.

7. The method for predicting the multidimensional health trend of B5G base stations based on the improved Transformer as described in claim 1, characterized in that, The cross-level health aggregation strategy in S5 represents the overall operating status of the base station by weighted aggregation of multi-dimensional prediction results. The aggregation strategy includes an aggregation strategy based on static functional weights, a dynamic weight strategy based on the amount of information in the prediction sequence, and a step-by-step aggregation strategy based on multi-dimensional weight fusion. The aggregation strategy based on static functional weights includes: constructing a judgment matrix using the Analytic Hierarchy Process (AHP) and calculating the relative importance weight of each functional circuit within its respective subsystem. And the relative importance weight λ of each subsystem in the overall B5G base station. sub The subsystem includes at least an active antenna unit (AAU) and a baseband processing unit (BBU). The dynamic weighting strategy based on the information content of the predicted sequence includes: calculating the information entropy of the predicted sequence of each functional circuit based on the inverse entropy weighting method, and then calculating the dynamic information weight. To reflect the impact of predicted sequence fluctuations on the system; The hierarchical aggregation strategy based on multi-dimensional weight fusion includes: subsystem-level aggregation and system-level aggregation. First, subsystem-level aggregation is performed: combining the static functional weights... With the dynamic information weight Calculate the overall weight of each functional circuit. Using the aforementioned comprehensive weights The predicted health values ​​of each functional circuit level are aggregated into the health trend of the B5G base station subsystem; then, system-level aggregation is performed: based on the relative importance weight λ of each subsystem in the overall B5G base station. sub The health trends of the B5G base station subsystem are aggregated into the health trends of the B5G base station system.

8. The method for predicting the multidimensional health trend of B5G base stations based on the improved Transformer as described in claim 7, characterized in that, The calculation process for cross-level health aggregation includes the following steps: S521: For the k-th functional circuit, the information entropy E of the prediction sequence within the future prediction step size H is calculated. k Construct dynamic information weights, where information entropy is defined as: In the formula: E k p is the information entropy of the predicted sequence of the k-th functional circuit within the future prediction step size H; kt Let be the normalized probability of the predicted value of the k-th circuit at a future time t; S522: According to E k Calculate the dynamic information weight of each functional circuit. The dynamic information weight is defined as follows: This increases the weight when the predicted sequence has large fluctuations or low entropy. In the formula: E represents the dynamic information weight of the k-th functional circuit; k M represents the information entropy of the predicted sequence within the future prediction step size H for the k-th functional circuit; sub The total number of functional circuits contained within the subsystem AAU or BBU to which the functional circuit belongs; S523: Combining the aforementioned static function weights With the dynamic information weight Calculate the overall weight of each functional circuit. In the formula: α is the balance coefficient, which is used to adjust the proportion of static functional weight and dynamic information weight in the comprehensive evaluation. Its value range is 0≤α≤1. The comprehensive weight of the k-th functional circuit; These are static function weights; For dynamic information content weights; S524: Perform the subsystem-level aggregation: using the comprehensive weights The health trends of the AAU and BBU subsystems are calculated. AAU (t) and S BBU (t), where Ω AAU and Ω BBU These are the functional circuit sets belonging to AAU and BBU, respectively. k (t) represents the circuit prediction value: S525: Perform the system-level aggregation: Calculate the base station system-level health trend S based on the subsystem weights λ. total (t): S total (t)=λ AAU ·S AAU (t)+λ BBU ·S BBU (t)。 9. A B5G base station multidimensional health trend prediction system based on an improved Transformer, used to implement the B5G base station multidimensional health trend prediction method based on an improved Transformer as described in any one of claims 1-8, characterized in that, It includes a multidimensional health input sequence construction module, a trend-sensitive sparse attention encoding module, a dual-path feature fidelity distillation module, a multidimensional synchronous decoding and prediction module, and a cross-level health aggregation module; The multidimensional health input sequence construction module is configured as follows: first, acquire historical health monitoring data of multiple functional circuits in the B5G base station; second, perform timestamp alignment and normalization processing on the monitoring data of each functional circuit; and finally, construct a multidimensional input tensor. The trend-sensitive sparse attention encoding module is configured to perform trend-sensitive weighted sparse attention calculation in the Transformer encoder, wherein the sparsity metric introduces a local volatility weighting factor to filter query vectors participating in the attention calculation. The dual-path feature fidelity distillation module is configured as follows: a dual-path distillation structure is established with main path downsampling and residual path downsampling. Convolution and max pooling are performed on the main path, and average pooling is performed on the residual path. The dual-path outputs are then fused to form a pyramid-shaped feature. The multidimensional synchronous decoding prediction module is configured to: map the encoded features into a multidimensional prediction matrix through a decoder, and synchronously output the predicted health values ​​of each functional circuit at H future time points; The cross-level health aggregation module is configured as follows: based on a hierarchical association model including functional circuit-subsystem-system, the predicted values ​​at the functional circuit level are aggregated at the subsystem level and then at the system level to obtain the base station system-level health trend prediction sequence.