A communication anomaly detection method based on multi-modal learning

By constructing an anomaly diagnosis mechanism based on the complex field Hermitian matrix and fractional Fourier transform gradient, extracting multimodal joint features, and utilizing an improved DVAE learning model, the accuracy and adaptability issues of communication anomaly detection in existing technologies are solved, achieving efficient communication anomaly localization and accurate detection.

CN122339987APending Publication Date: 2026-07-03JIANGSU SHIFANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU SHIFANG TECHNOLOGY CO LTD
Filing Date
2026-04-09
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing multimodal anomaly detection algorithms struggle to fully utilize the probabilistic mutual information in the cross-modal weight matrix when dealing with communication anomaly detection in complex network environments. This leads to fixed false alarm boundaries when handling non-stationary burst communication anomalies, increases the computational overhead of adaptive localization, and ignores the complex temporal correlation structure in the original traffic data and log data, thus limiting the accuracy of detection.

Method used

By constructing an anomaly-triggered refined diagnostic mechanism based on the complex field Hermitian matrix and fractional Fourier transform gradient, a multimodal joint feature basis of traffic and logs is extracted. An improved DVAE learning model is used for feature reconstruction and anomaly localization. The communication anomaly score is output by combining the mutual information trace norm and the second-order nuclear norm of the Kronecker product. Time-frequency mask suppression and instantaneous frequency mutation energy extremum search are performed through cross-modal weight matrix.

Benefits of technology

It achieves improved accuracy in locating sudden communication anomalies while maintaining high-efficiency detection, overcomes the limitations of traditional methods such as shallow modal fusion, fixed thresholds, and neglect of time-frequency change constraints, and provides an efficient intelligent operation and maintenance solution for communication networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122339987A_ABST
    Figure CN122339987A_ABST
Patent Text Reader

Abstract

The application discloses a communication anomaly detection method based on multi-modal learning, relates to the technical field of communication networks, and comprises the following steps: S1, outputting a traffic feature matrix and a log feature matrix; S2, generating a cross-modal weight matrix and outputting a multi-modal fusion feature matrix; S3, outputting a communication anomaly contradiction score sequence; S4, improving a DVAE learning model, using local variance to guide anisotropic noise and variance condition adaptive compensation decoding, measuring a reconstruction deviation degree through a Huber loss, and outputting a communication anomaly reconstruction error score sequence; S5, outputting a communication anomaly score; S6, outputting an anomaly detection result; and S7, outputting a communication anomaly positioning result. The application overcomes the limitations of traditional methods, such as shallow modal fusion, fixed threshold and neglecting time-frequency mutation constraints, and provides an efficient solution for intelligent operation and maintenance of communication networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication network technology, and in particular to a communication anomaly detection method based on multimodal learning. Background Technology

[0002] With the explosive growth of multimodal communication data, classical computers face severe computational challenges in joint feature extraction and real-time anomaly detection from massive traffic and logs. Existing multimodal anomaly detection algorithms, such as improved dual-stream attention networks, while enhancing feature fusion efficiency through cross-attention mechanisms, primarily rely on independent Euclidean distances between traffic and log feature vectors for similarity calculation and reconstruction error assessment. This method, based solely on independent spatial distance, ignores the complex temporal correlation structures (such as autocorrelation lag characteristics and non-stationary frequency abrupt changes) and deep cross-modal contradictory semantic associations implicit in the original traffic and log data. This leads to structural distortion of the high-dimensional feature space when constructing multimodal joint representations, thus limiting the accuracy of communication anomaly detection in complex network environments. Furthermore, classical anomaly detection methods often struggle to fully utilize the probabilistic mutual information in the cross-modal weight matrix to dynamically constrain anomaly scoring thresholds, resulting in fixed false alarm boundaries when handling non-stationary burst communication anomalies, increasing the computational overhead of adaptive localization.

[0003] Therefore, how to provide a communication anomaly detection method based on multimodal learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] This invention proposes a communication anomaly detection method based on multimodal learning. It employs an anomaly-triggered refined diagnostic mechanism based on complex-domain Hermitian matrices and fractional Fourier transform gradients. Complex-domain eigenvalue decomposition is performed on the traffic autocorrelation matrix and log autocorrelation matrix within the current time window to extract a multimodal joint feature basis containing interwoven traffic and log features. A multimodal fusion feature matrix is ​​constructed based on a cross-modal attention mechanism with relative entropy minimization constraints. This multimodal fusion feature matrix is ​​input into an improved DVAE learning model, generating an anisotropic noise matrix based on local variance and adaptively compensating for decoding with variance as a condition. A communication anomaly score is output by combining the mutual information trace norm and the second-order kernel norm of the Kronecker product. When the anomaly detection result is positive, the multimodal fusion feature matrix undergoes complex-domain dimensionality reduction along the feature dimension. A cross-modal weight matrix is ​​used as a time-frequency mask to suppress interference. The energy extrema of instantaneous frequency mutations are searched based on fractional Fourier transform gradients to output the communication anomaly localization result. This mechanism effectively eliminates single-modal semantic bias by establishing a two-level triggering path from "multimodal joint error measurement" to "time-frequency domain extreme value search," achieving improved technical performance in accurately locating sudden communication anomalies while maintaining high-efficiency detection. This invention overcomes the limitations of traditional methods, such as shallow modal fusion, fixed thresholds, and neglect of time-frequency abrupt change constraints, providing an efficient solution for intelligent operation and maintenance of communication networks.

[0005] A communication anomaly detection method based on multimodal learning according to an embodiment of the present invention specifically includes: S1. Construct the autocorrelation matrix of traffic and logs within the current time window into a complex field Hermitian matrix, extract the multimodal joint feature basis through eigenvalue decomposition and project it to output the traffic feature matrix and log feature matrix. S2. Map the traffic feature matrix and the log feature matrix to a query, key and value matrix. Calculate the similarity of the matrix under the constraint of minimizing relative entropy to generate a cross-modal weight matrix. Then, sum the weighted values ​​of the cross-modal weight matrix to output a multimodal fusion feature matrix. S3. Extract the cross-modal weight matrix to construct the conditional mutual information matrix of the log dimension, calculate the trace norm of the conditional mutual information matrix, and output the communication anomaly contradiction score sequence by taking the negative value. S4. The multimodal fusion feature matrix and the traffic feature matrix are concatenated to generate a multimodal joint feature matrix and input into the improved DVAE learning model. Local variance is used to guide anisotropic noise addition and variance condition adaptive compensation decoding. The deviation is reconstructed through Huber loss metric, and the communication anomaly reconstruction error score sequence is output. S5. Extract the communication anomaly contradiction score sequence and the communication anomaly reconstruction error score sequence within the historical time window, perform the Kronecker product and solve the nuclear norm to output the communication anomaly score; S6. Extract historical communication anomaly scores to construct a Hankel matrix sequence and solve the nuclear norm sequence. Take the local minimum point as the communication anomaly score threshold, compare it with the current communication anomaly score, and output the anomaly detection result. S7. If the anomaly detection result is abnormal, the multimodal fusion feature matrix is ​​reduced in dimensionality along the feature dimension in the complex domain. The cross-modal weight matrix is ​​used as a time-frequency mask. The instantaneous frequency change energy extremum is searched based on the fractional Fourier transform gradient, and the communication anomaly location result is output.

[0006] Optionally, S1 specifically includes: S11. Based on the traffic data within the current time window, calculate the cross-correlation coefficient of the lag order along the time dimension, and construct the traffic autocorrelation matrix; S12. Based on the log data within the current time window, extract the log template sequence and convert it into a log term frequency vector. Calculate the lag order cross-correlation coefficient along the time dimension and construct the log autocorrelation matrix. S13. Perform zero-mean standardization on the traffic autocorrelation matrix and the log autocorrelation matrix to generate a standardized traffic matrix and a standardized log matrix; S14. Construct a complex field Hermitian matrix by using the standardized traffic matrix as the real part and the standardized log matrix as the imaginary part; S15. Perform eigenvalue decomposition on the complex field Hermitian matrix to obtain a descending sequence of eigenvalues ​​and an eigenvector matrix. Calculate the ratio sequence of adjacent eigenvalues ​​in the eigenvalue sequence. Use the position where the ratio sequence first falls below a preset decay limit as the truncation index. Based on the truncation index, extract a feature submatrix composed of traffic features and log features from the eigenvector matrix. Use the feature submatrix as a multimodal joint feature basis. S16. Perform matrix multiplication on the standardized traffic matrix and the multimodal joint feature basis to output the dimensionality-reduced traffic feature matrix. Perform matrix multiplication on the standardized log matrix and the multimodal joint feature basis to output the dimensionality-reduced log feature matrix.

[0007] Optionally, the cross-modal weight matrix includes a first cross-modal sub-weight matrix and a second cross-modal sub-weight matrix: Perform matrix multiplication on the traffic feature matrix with the initialized first query weight matrix, first key weight matrix, and first value weight matrix respectively, and output the traffic query matrix, traffic key matrix, and traffic value matrix; perform matrix multiplication on the log feature matrix with the initialized second query weight matrix, second key weight matrix, and second value weight matrix respectively, and output the log query matrix, log key matrix, and log value matrix. Perform matrix multiplication on the transpose of the traffic query matrix and the log key matrix to output the original cross-score matrix; calculate the natural logarithm of the sum of the exponents of each row of the original cross-score matrix, and subtract the original cross-score matrix element by element from the natural logarithm as a penalty term to output the constrained cross-score matrix; input the constrained cross-score matrix into the Softmax activation function to output the first cross-modal sub-weight matrix. Perform a weighted summation operation on the first cross-modal sub-weight matrix and the log value matrix to output a multimodal log side context enhancement matrix; perform matrix multiplication operation on the log query matrix and the transpose of the traffic key matrix to output an inverse original cross-score matrix; calculate the natural logarithm of the sum of the exponent values ​​of each row of the inverse original cross-score matrix as a penalty term and subtract the inverse original cross-score matrix element by element to output an inverse constrained cross-score matrix; input the inverse constrained cross-score matrix into the Softmax activation function to output a second cross-modal sub-weight matrix; Perform a weighted summation operation on the second cross-modal sub-weight matrix and the traffic value matrix to output a multimodal traffic-side context enhancement matrix; perform a concatenation operation on the multimodal traffic-side context enhancement matrix and the multimodal log-side context enhancement matrix to output a multimodal fusion feature matrix.

[0008] Optionally, S3 specifically includes: S31. Read the cross-modal weight matrix, extract the first cross-modal sub-weight matrix and the second cross-modal sub-weight matrix of the cross-modal weight matrix; calculate the mean of the first cross-modal sub-weight matrix along the row dimension, and output the log edge probability column vector; calculate the mean of the second cross-modal sub-weight matrix along the column dimension, and output the traffic edge probability row vector; perform element-wise multiplication on the transpose of the first cross-modal sub-weight matrix and the second cross-modal sub-weight matrix, and output the joint probability matrix; calculate the outer product of the log edge probability column vector and the traffic edge probability row vector, add a preset outer product divide-by-zero constant, and output the independent probability benchmark matrix; divide each element in the joint probability matrix by the corresponding element in the independent probability benchmark matrix, take the natural logarithm of the division result, and multiply it by the corresponding element in the joint probability matrix to output the multimodal conditional mutual information matrix; S32. Perform singular value decomposition on the multimodal conditional mutual information matrix, extract all singular values ​​after decomposition and calculate the sum of their absolute values, and output the trace norm of the multimodal conditional mutual information matrix. S33. Add a negative sign to the trace norm, and use the trace norm after adding the negative sign as the multimodal communication conflict score of the current time window. Arrange the multimodal communication conflict scores of all time windows in chronological order and output the communication anomaly conflict score sequence.

[0009] Optionally, the improved DVAE learning model includes an input projection layer, a joint probabilistic coding layer, an anisotropic noise injection layer, a conditional decoding layer, a feature reconstruction layer, and a Huber deviation calculation layer: The input projection layer is used to input the multimodal joint feature matrix into a linear mapping operator composed of a weight matrix and a bias vector to perform dimensionality reduction and nonlinear activation operations, and output a multimodal projection latent variable matrix. The joint probabilistic coding layer is used to input the multimodal projected latent variable matrix into two independent linear mapping operators composed of weight matrices and bias vectors to perform linear operations. The two output matrices are used as the latent variable prior mean matrix and the latent variable prior variance matrix, respectively. Based on the latent variable prior mean matrix and the latent variable prior variance matrix, a reparameterized sampling operation is performed to output the multimodal standard latent variable matrix. The anisotropic noise injection layer is used to generate a Gaussian noise matrix with the same dimension as the multimodal joint feature matrix. The local variance distribution matrix of each column of the multimodal joint feature matrix is ​​calculated along the feature dimension. After adding a preset matrix constant to each element of the local variance distribution matrix, the reciprocal operation is performed to output the reciprocal variance matrix. The Gaussian noise matrix and the reciprocal variance matrix are multiplied element by element and then superimposed on the multimodal joint feature matrix to output the multimodal noise-added joint feature matrix. The conditional decoding layer is used to calculate the local variance vectors of the multimodal denoising joint feature matrix and the multimodal standard latent variable matrix along the feature dimension, respectively. Each element in the local variance vector of the multimodal denoising joint feature matrix is ​​divided by the sum of the corresponding element in the local variance vector of the multimodal standard latent variable matrix and a preset element (excluding zero constants), outputting a variance compensation coefficient vector. Each column element in the multimodal standard latent variable matrix is ​​multiplied by the corresponding column coefficient in the variance compensation coefficient vector, outputting a multimodal recalibrated latent variable matrix. The multimodal recalibrated latent variable matrix is ​​input into a deconvolutional network and an activation function to perform upsampling operations, outputting a multimodal denoising latent variable decoding matrix. The feature reconstruction layer is used to input the multimodal denoising latent variable decoding matrix into a linear mapping operator composed of a weight matrix and a bias vector to perform a dimension mapping operation, restore the feature dimension to the dimension of the multimodal joint feature matrix, and output the multimodal reconstruction joint feature matrix. The Huber deviation calculation layer is used to calculate the absolute value of the element-wise difference between the multimodal reconstruction joint feature matrix and the multimodal joint feature matrix. Based on the preset threshold boundary, it sequentially performs a combination of quadratic and linear piecewise mappings to generate an element-level Huber loss matrix. It then performs a global mean operation on the matrix and outputs a communication anomaly reconstruction error score.

[0010] Optionally, S5 specifically includes: S51. Extract the communication anomaly contradiction score sequence and the communication anomaly reconstruction error score sequence within the historical time window, align and concatenate the communication anomaly contradiction score sequence and the communication anomaly reconstruction error score sequence according to the time step, and output the multimodal joint score matrix. S52. Calculate the Kronecker product between the multimodal joint score matrix and itself, and use the output of the Kronecker product operation as the second-order joint score matrix of the multimodal model. S53. Using the multimodal second-order joint score matrix as input, perform singular value decomposition to extract all singular values, use all singular values ​​as input to perform summation, and use the output of the summation as the communication anomaly score.

[0011] Optionally, S6 specifically includes: S61. Extract communication anomaly scores within the historical time window that do not include the current time. Input the communication anomaly scores into the sliding window, extract them sequentially by time step, and perform matrix permutation operations. Use the output of the matrix permutation operations as the Hankel matrix sequence. S62. Take each Hankel matrix in the Hankel matrix sequence as input and perform singular value decomposition to extract all singular values. Take all singular values ​​as input and perform summation. Arrange the output of the summation operation according to the time step as the nuclear norm sequence. S63. Perform a local minimum point search operation on the nuclear norm sequence as input, and use the output of the local minimum point search operation as the communication anomaly scoring threshold. S64. Extract the communication anomaly score calculated independently at the current time, perform a size comparison operation with the communication anomaly score threshold calculated independently at the current time as input, and use the output result of the size comparison operation as the anomaly detection result.

[0012] Optionally, S7 specifically includes: S71. When the anomaly detection result is abnormal, the multimodal fusion feature matrix is ​​used as input to perform a complex domain dimension reduction projection operation along the feature dimension, and the output result of the complex domain dimension reduction projection operation is used as the multimodal complex projection feature matrix. S72. Read the cross-modal weight matrix, extract the first cross-modal sub-weight matrix and the second cross-modal sub-weight matrix of the cross-modal weight matrix; take the multimodal complex projection feature matrix, the first cross-modal sub-weight matrix and the second cross-modal sub-weight matrix as input and perform element-wise multiplication operation, and take the output result of the element-wise multiplication operation as the multimodal time-frequency mask suppression feature matrix; S73. Perform fractional Fourier transform operation on the multimodal time-frequency mask suppression feature matrix as input, and use the output of the fractional Fourier transform operation as the multimodal time-frequency transform matrix. S74. Perform gradient calculation operation along the feature dimension using the multimodal time-frequency transformation matrix as input, and use the output result of the gradient calculation operation as the multimodal fractional Fourier transform gradient matrix. S75. The instantaneous frequency mutation energy calculation operation is performed using the multimodal fractional Fourier transform gradient matrix as input, and the output result of the instantaneous frequency mutation energy calculation operation is used as the multimodal instantaneous frequency mutation energy sequence. S76. The multimodal instantaneous frequency mutation energy sequence is used as input to perform an extreme point search operation, and the output result of the extreme point search operation is used as the communication anomaly location result.

[0013] The beneficial effects of this invention are: (1) This invention achieves deep interweaving and dimensionality reduction representation of traffic time-series features and log term frequency features by constructing a complex-domain Hermitian matrix and an eigenvalue decomposition truncation mechanism. The complex-domain space is constructed using the traffic autocorrelation matrix and the log autocorrelation matrix as the real and imaginary parts, respectively. Eigenvalue decomposition is performed on the Hermitian matrix, and the truncation index is accurately located based on the lower bound of adjacent eigenvalue decay. This mechanism maps heterogeneous modal information to a complex conjugate symmetric space, accurately separates the strongly coupled multimodal joint basis through mathematical truncation, completely eliminates redundant noise, and provides a low-dimensional projection space containing the core temporal correlation of dual modes.

[0014] (2) This invention achieves high-fidelity reconstruction of multimodal features and quantification of anomaly errors by constructing an improved DVAE learning model with an anisotropic noise injection and variance conditional adaptive compensation decoding mechanism. The improved DVAE learning model generates a standard latent variable matrix through a joint probabilistic coding layer, generates a noisy joint feature matrix by modulating Gaussian noise with the inverse of local variance using an anisotropic noise injection layer, calculates compensation coefficients by extracting the local variance of the noisy matrix and the latent variable matrix, recalibrates the latent variables and performs upsampling decoding, and finally outputs the communication anomaly reconstruction error score by a Huber deviation calculation layer. This mechanism uses the variance ratio condition to dynamically stretch the prior latent variables to counteract anisotropic noise interference, ensuring that the output reconstruction error score can quantify the micro-deviation degree of multimodal data with extremely high accuracy.

[0015] (3) This invention achieves high-resolution physical localization of communication anomalies triggered by anomalies by constructing a time-frequency mask suppression and fractional Fourier transform gradient search mechanism. When an anomaly is triggered, the multimodal fusion feature matrix is ​​subjected to complex domain dimensionality reduction projection, and the time-frequency mask suppression feature matrix is ​​generated by element-wise multiplication using the cross-modal weight matrix. The suppression feature matrix is ​​subjected to fractional Fourier transform and the gradient is calculated along the feature dimension. Based on the transformed gradient matrix, the extreme value of the instantaneous frequency change energy is searched and the localization result is output. This mechanism implements cross-attention mask filtering through cross-modal weights and extracts the gradient of non-stationary instantaneous frequency change by fractional Fourier transform, ensuring that the output localization result has extremely high capture resolution for the physical boundary of sudden anomalies. Attached Figure Description

[0016] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is an overall flowchart of a communication anomaly detection method based on multimodal learning proposed in this invention; Figure 2 This is a flowchart illustrating the working principle of the improved DVAE learning model for a communication anomaly detection method based on multimodal learning proposed in this invention. Detailed Implementation

[0017] The invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0018] refer to Figure 1 and Figure 2 A communication anomaly detection method based on multimodal learning, specifically including: S1. Construct the autocorrelation matrix of traffic and logs within the current time window into a complex field Hermitian matrix, extract the multimodal joint feature basis through eigenvalue decomposition and project it to output the traffic feature matrix and log feature matrix. S2. Map the traffic feature matrix and log feature matrix to query, key and value matrices. Calculate the similarity of the matrix under the constraint of minimizing relative entropy to generate a cross-modal weight matrix. Then, sum the weighted values ​​of the matrix to output the multimodal fusion feature matrix. S3. Extract the cross-modal weight matrix to construct the conditional mutual information matrix of the log dimension, calculate the trace norm of the conditional mutual information matrix, and output the communication anomaly contradiction score sequence by taking the negative value. S4. Concatenate the multimodal fusion feature matrix and the traffic feature matrix to generate a multimodal joint feature matrix and input it into the improved DVAE learning model. Generate an anisotropic noise matrix and a prior latent variable matrix based on local variance. Adaptively compensate the decoding prior latent variable matrix with the variance of the anisotropic noise matrix as a condition. Calculate the Huber loss deviation between the decoding matrix and the multimodal joint feature matrix and output the communication anomaly reconstruction error score sequence. S5. Extract the communication anomaly contradiction score sequence and the communication anomaly reconstruction error score sequence within the historical time window, perform the Kronecker product and solve the nuclear norm to output the communication anomaly score; S6. Extract historical communication anomaly scores to construct a Hankel matrix sequence and solve for the nuclear norm sequence. Take the local minimum point as the communication anomaly score threshold, compare it with the current communication anomaly score, and output the anomaly detection result. S7. If the anomaly detection result is anomaly, the multimodal fusion feature matrix is ​​reduced in dimensionality along the feature dimension in the complex domain. The cross-modal weight matrix is ​​used as a time-frequency mask. The instantaneous frequency change energy extremum is searched based on the fractional Fourier transform gradient, and the communication anomaly location result is output.

[0019] In this embodiment, S1 specifically includes: S11. Based on the traffic data within the current time window, set the maximum lag order to 10, calculate the Pearson cross-correlation coefficients from order 1 to order 10 along the time dimension, and calculate the 0th order autocorrelation coefficient; fill the 0th order autocorrelation coefficient into all 10 positions on the main diagonal of a 10x10 blank matrix, fill the 1st order cross-correlation coefficient into all 9 positions on the secondary diagonal shifted 1 position upward from the main diagonal, fill the 2nd order cross-correlation coefficient into all 8 positions on the secondary diagonal shifted 2 positions upward from the main diagonal, and so on, fill the 9th order cross-correlation coefficient into 1 position on the secondary diagonal shifted 9 positions upward, and fill the 10th order cross-correlation coefficient into 0 positions on the secondary diagonal shifted 10 positions upward and discard them, fill all other matrix positions except the coefficients with the value 0, and construct a 10x10 square matrix as the traffic autocorrelation matrix; S12. Based on the log data within the current time window, extract the log template sequence using regular expression parsing. Utilize the bag-of-words model to transform the log template sequence into a term frequency vector. Set the maximum lag order to 10. Calculate the Pearson cross-correlation coefficients from order 1 to order 10 along the time dimension, and calculate the 0th-order autocorrelation coefficient. Fill the 0th-order autocorrelation coefficients into all 10 positions on the main diagonal of a 10x10 blank matrix. Fill the 1st-order cross-correlation coefficients into the sub-matrix shifted one position upwards from the main diagonal. For all 9 positions on the diagonal, fill the 2nd order cross-correlation coefficient into all 8 positions on the secondary diagonal shifted 2 positions upward from the main diagonal. Similarly, fill the 9th order cross-correlation coefficient into 1 position on the secondary diagonal shifted 9 positions upward, and fill the 10th order cross-correlation coefficient into 0 positions on the secondary diagonal shifted 10 positions upward, then discard them. Fill all other matrix positions except for the coefficients with the value 0, and construct a 10-row, 10-column square matrix as the log autocorrelation matrix. S13. Traverse each element in the traffic autocorrelation matrix, calculate the difference between the current element value and the global mean of the traffic autocorrelation matrix, divide the difference by the global standard deviation of the traffic autocorrelation matrix to generate a standardized traffic matrix, and process each element in the log autocorrelation matrix using the same subtraction of the global mean and division by the global standard deviation to generate a standardized log matrix. S14. Read the first row and first column element of the standardized traffic matrix as the complex real part, read the first row and first column element of the standardized log matrix as the complex imaginary part, combine them to generate the complex elements of the first row and first column of the complex field Hermit matrix, and traverse all row and column positions according to the corresponding position of the row and column to construct a 10-row and 10-column complex field Hermit matrix. S15. Call the eigenvalue decomposition algorithm to perform eigenvalue decomposition on the complex field Hermit matrix, obtain the eigenvalue sequence and the corresponding eigenvector matrix arranged in descending order of value, divide the second eigenvalue in the sequence by the first eigenvalue to obtain the first ratio, and so on to calculate the ratio sequence of adjacent eigenvalues. Set the preset decay lower limit to the value of 0.15, traverse and compare from the first ratio, and take the position of the first ratio less than 0.15 as the truncation index. Extract all column vectors from the first column to the column corresponding to the truncation index in the eigenvector matrix, concatenate them to construct a feature submatrix, and use the feature submatrix as the multimodal joint feature basis. S16. Perform matrix multiplication on the standardized traffic matrix and the multimodal joint feature basis, multiply and accumulate the elements of each row of the standardized traffic matrix with the corresponding elements of each column of the multimodal joint feature basis, and output the dimensionality-reduced traffic feature matrix. Perform the same multiplication and accumulation on the standardized log matrix and the multimodal joint feature basis, and output the dimensionality-reduced log feature matrix.

[0020] In this embodiment, the cross-modal weight matrix includes a first cross-modal sub-weight matrix and a second cross-modal sub-weight matrix: The system reads the number of columns from the traffic feature matrix as the first feature dimension and the number of columns from the log feature matrix as the second feature dimension. It constructs and initializes the first feature dimension by multiplying the first query weight matrix, the first key weight matrix, and the first value weight matrix by a preset hidden layer value of 64. It also constructs and initializes the second feature dimension by multiplying the second query weight matrix, the second key weight matrix, and the second value weight matrix by a preset hidden layer value of 64. The system performs matrix multiplication operations (multiplying row elements and column elements) with the first query weight matrix, the first key weight matrix, and the first value weight matrix, respectively, and then sums the results. The output dimensions are the traffic query matrix, the traffic key matrix, and the traffic value matrix, all with a preset hidden layer value of 64. Similarly, the system performs the same matrix multiplication operations (multiplying row elements and column elements) with the second query weight matrix, the second key weight matrix, and the second value weight matrix, respectively, and then sums the results. The output dimensions are the log query matrix, the log key matrix, and the log value matrix, all with a preset hidden layer value of 64. The process involves multiplying each element in the traffic query matrix with the corresponding element in each column of the transpose of the log key matrix and summing the results to output the original cross-score matrix. Then, iterating through each row of the original cross-score matrix, each element in the current row is used as the base for exponential operations with the natural constant e. The results of all exponential operations in the current row are summed, and the logarithm of the sum is extracted. This logarithm is then used as the element of the current row in the logarithmic vector. After traversing all rows, a logarithmic vector with the same number of rows and columns as the original cross-score matrix is ​​generated. Each element in the original cross-score matrix is ​​subtracted from the element in the corresponding row of the logarithmic vector to generate the penalty constraint matrix. Simultaneously, each element in the original cross-score matrix is ​​subjected to exponential operations with the natural constant e and divided by the sum of all exponential operations in the current row. A Softmax activation function mapping is then performed to generate the initial sub-weight matrix. Finally, the penalty constraint matrix and the initial sub-weight matrix are multiplied element-wise to output the first cross-modal sub-weight matrix. Perform matrix multiplication by multiplying each element of the first cross-modal sub-weight matrix with the corresponding element of each column of the log value matrix and summing the results. Output the multimodal log side context enhancement matrix. Perform matrix multiplication by multiplying each element of the log query matrix with the corresponding element of each column of the transpose of the flow key matrix and summing the results. Output the inverse original cross-score matrix. Traverse each row of the inverse original cross-score matrix, perform exponential operations with the natural constant e as the base for each element in the current row, sum the results of all exponential operations in the current row, and extract the logarithm of the summation result with the natural constant e as the base. Use the logarithm calculated for the current row as... For each element in the current row of the inverse logarithmic vector, after traversing all rows, a reverse logarithmic vector with the same number of rows and columns as the original inverse cross-score matrix is ​​generated. Each element in the original inverse cross-score matrix is ​​subtracted from the element in the corresponding row of the inverse logarithmic vector to generate the inverse penalty constraint matrix. At the same time, each element in the original inverse cross-score matrix is ​​subjected to an exponential operation with the natural constant e as the base and divided by the sum of the results of all exponential operations in the current row. The Softmax activation function is then applied to generate the inverse initial sub-weight matrix. The inverse penalty constraint matrix and the inverse initial sub-weight matrix are multiplied element-wise to output the second cross-modal sub-weight matrix. Perform a matrix multiplication operation by multiplying the elements of each row in the second cross-modal sub-weight matrix with the corresponding elements in each column of the flow value matrix and accumulating the results. Output a multimodal flow-side context enhancement matrix. Perform a first-to-last concatenation operation by combining each row of the multimodal flow-side context enhancement matrix with the corresponding row of the multimodal log-side context enhancement matrix. Expand the number of columns in the concatenated matrix to twice the number of columns in the original matrix and output a multimodal fusion feature matrix.

[0021] In this embodiment, S3 specifically includes: S31. Read the cross-modal weight matrix, extract the first cross-modal sub-weight matrix generated by the traffic query log key and the second cross-modal sub-weight matrix generated by the log query traffic key. Traverse each row of the first cross-modal sub-weight matrix, sum all column elements in the current row and divide by the total number of columns in the current row. Use the calculated mean as the element of the current row in the log edge probability column vector. After traversing all rows, output the log edge probability column vector. Traverse each column of the second cross-modal sub-weight matrix, sum all row elements in the current column and divide by the total number of rows in the current column. Use the calculated mean as the element of the current column in the traffic edge probability row vector. After traversing all columns, output the traffic edge probability row vector. Combine the element in the i-th row and j-th column of the first cross-modal sub-weight matrix with the element in the second cross-modal sub-weight matrix. The elements in the i-th row and j-th column of the modal sub-weight matrix are multiplied one-to-one. The joint probability matrix is ​​output by traversing all rows and columns. The elements in the i-th row of the log edge probability column vector and the elements in the j-th column of the traffic edge probability row vector are multiplied and a preset outer product constant of 1e-8 is added. The independent probability baseline matrix is ​​output by traversing all rows and columns. The ratio of the elements in the i-th row and j-th column of the joint probability matrix to the elements in the i-th row and j-th column of the independent probability baseline matrix is ​​obtained. The logarithm of this ratio is extracted to the base e. The logarithm is multiplied by the original element in the i-th row and j-th column of the joint probability matrix. The result of the multiplication is used as the element in the i-th row and j-th column of the mutual information local matrix. The mutual information local matrix is ​​generated by traversing all rows and columns. The mutual information local matrix is ​​used as the multimodal conditional mutual information matrix. S32. Perform matrix decomposition on the multimodal conditional mutual information matrix through singular value decomposition to obtain the entire sequence of singular values ​​arranged in descending order of numerical value. Traverse each singular value in the singular value sequence, determine the sign of the current singular value. If it is positive, keep it directly. If it is negative, remove the negative sign and convert it to a positive number. Extract the absolute value of the current singular value. Perform addition and summation on all extracted absolute values. Use the summation result as the trace norm of the multimodal conditional mutual information matrix. S33. Add a negative sign before the value of the trace norm to convert it to a negative value. Define the trace norm after adding the negative sign as the multimodal communication conflict score of the degree of cross-modal interaction conflict between traffic and logs within the current time window. Obtain multiple time windows arranged in chronological order. Concatenate the multimodal communication conflict scores calculated for each time window as sequence elements in chronological order to output a communication anomaly conflict score sequence that reflects the evolution trend of communication conflicts in the time dimension.

[0022] The multimodal contradiction score extraction process proposed in this step is similar to the traditional cross-modal attention fusion mechanism in that both are based on probabilistic statistical learning theory, quantify the feature association between modalities by calculating the cross-modal interaction weight matrix, and use this interaction matrix as the basis for measuring the consistency of multi-source data.

[0023] The difference lies in that this invention breaks through the limitations of traditional methods that merely utilize the weight matrix to perform weighted summation of features or ignore their probabilistic topological meaning. This invention adds a probability distribution marginalization reconstruction step, extracting marginal probability vectors by averaging the sub-weight matrices along the row and column dimensions, rather than discarding their statistical meaning; in the mutual information calculation step, an independent benchmark is constructed through the outer product of marginal probabilities, and the conditional mutual information matrix is ​​output by performing element-wise logarithmic and product operations combined with the joint probability, instead of directly outputting the fused features; finally, in the contradiction score output step, the mutual information matrix is ​​subjected to singular value decomposition to solve for the trace norm and then negatively evaluated, outputting a scalar contradiction score sequence, rather than a high-dimensional feature matrix.

[0024] The beneficial effect of this improvement lies in the fact that, by probabilistic marginalization of the weight matrix and solving for the trace norm of mutual information, the implicit attention similarity is explicitly transformed into a contradictory measure in the information theory dimension. This breaks through the limitation of traditional methods that only perform feature selection but cannot quantify modal logical conflicts, achieving a precise mapping from feature-level weighting to semantic-level contradictory distance. This design significantly enhances the ability to perceive multimodal inconsistencies, accurately captures abrupt changes in traffic and log decoupling caused by anomalies using the trace norm, and effectively filters out high-frequency disturbances based on negative trace norm scalarization, significantly improving the system's accuracy in identifying hidden anomalies.

[0025] In this embodiment, the improved DVAE learning model includes an input projection layer, a joint probabilistic coding layer, an anisotropic noise injection layer, a conditional decoding layer, a feature reconstruction layer, and a Huber deviation calculation layer: The input projection layer is used to read the multimodal joint feature matrix and construct a linear mapping operator consisting of a dimensionality reduction weight matrix with the number of columns set to a preset latent variable dimension of 128 and a dimensionality reduction bias vector. The operator performs a multiplication and accumulation operation on the elements of each row in the multimodal joint feature matrix and the elements of the corresponding column in the dimensionality reduction weight matrix. The accumulated result is added to the bias element of the corresponding row in the dimensionality reduction bias vector. The summed value is input into the modified linear unit activation function to perform a nonlinear mapping operation and output the multimodal projected latent variable matrix. The joint probabilistic coding layer is used to construct a first independent linear mapping operator consisting of a prior mean weight matrix and a prior mean bias vector, and a second independent linear mapping operator consisting of a prior variance weight matrix and a prior variance bias vector. The multimodal projected latent variable matrix is ​​input into the first and second independent linear mapping operators respectively to perform matrix multiplication and bias addition, and outputs the latent variable prior mean matrix and the latent variable prior variance matrix. The standard Gaussian distribution random number generation algorithm is called to generate a standard Gaussian noise matrix with the same dimension as the latent variable prior mean matrix. The square root operation is performed on each element in the latent variable prior variance matrix to generate the standard deviation matrix. Each element in the standard deviation matrix is ​​multiplied element-wise with the corresponding element in the standard Gaussian noise matrix and added to the corresponding element in the latent variable prior mean matrix, and the multimodal standard latent variable matrix is ​​output. The anisotropic noise injection layer is used to call the Gaussian random number generator to generate a Gaussian noise matrix with the same number of rows and columns as the multimodal joint feature matrix. It iterates through each column of the multimodal joint feature matrix, performs addition on all row elements in the current column, and divides the sum by the total number of rows in the current column. The calculated mean is used as the element of the current column in the local variance distribution matrix. After iterating through all columns, the local variance distribution matrix is ​​output. Each element in the local variance distribution matrix is ​​added to a preset matrix constant of 1e-8 and then its reciprocal is calculated to output the reciprocal variance matrix. The element in the i-th row and j-th column of the Gaussian noise matrix is ​​multiplied element-wise with the element in the i-th row and j-th column of the reciprocal variance matrix. The result of the multiplication is added to the original element in the i-th row and j-th column of the multimodal joint feature matrix. After iterating through all rows and columns, the multimodal noise joint feature matrix is ​​output. The conditional decoding layer iterates through each column of the multimodal-noise joint feature matrix, summing all row elements in the current column and dividing by the total number of rows in that column. The calculated mean is used as the element of the current column in the local variance vector of the multimodal-noise joint feature matrix. The same averaging operation is then performed to iterate through each column of the multimodal standard latent variable matrix, outputting its local variance vector. Finally, the element of the j-th column in the local variance vector of the multimodal-noise joint feature matrix is ​​divided by the element of the j-th column in the local variance vector of the multimodal standard latent variable matrix, and a preset element of 1e-8 is added. The sum of zero constants is used as the coefficient of the j-th column in the variance compensation coefficient vector. All columns are traversed to output the variance compensation coefficient vector. Each column in the multimodal standard latent variable matrix is ​​traversed, and each row element in the current column is multiplied by the coefficient of the corresponding column in the variance compensation coefficient vector. All rows and columns are traversed to output the multimodal recalibrated latent variable matrix. The multimodal recalibrated latent variable matrix is ​​input into the deconvolution network to perform a transposed convolution upsampling operation with a stride of 2. Each element in the upsampled output matrix is ​​input into the modified linear unit activation function to perform a nonlinear mapping operation, and the multimodal denoising latent variable decoding matrix is ​​output. The feature reconstruction layer is used to construct a linear mapping operator consisting of a reconstruction weight matrix and a reconstruction bias vector. It performs multiplication and accumulation on the elements of each row in the multimodal denoising latent variable decoding matrix and the elements of the corresponding column in the reconstruction weight matrix. The accumulated result is added to the bias element of the corresponding row in the reconstruction bias vector. The number of columns in the reconstruction weight matrix is ​​set to the number of columns in the multimodal joint feature matrix to restore the feature dimension to the dimension of the multimodal joint feature matrix. The multimodal reconstruction joint feature matrix is ​​then output. The Huber deviation calculation layer is used to subtract the element in the i-th row and j-th column of the multimodal reconstruction joint feature matrix from the element in the i-th row and j-th column of the multimodal reconstruction joint feature matrix. It determines whether the difference obtained by the subtraction is negative. If it is negative, the negative sign is removed and converted to a positive number. The absolute value of the element-wise difference is extracted. A preset threshold boundary value of 1.0 is set. Each absolute value of the element-wise difference is traversed. When the current absolute value is less than or equal to 1.0, the square of the current absolute value is divided by 2 and used as the element at the corresponding position in the Huber loss matrix. When the current absolute value is greater than 1.0, the current absolute value is subtracted by 0.5 and used as the element at the corresponding position in the Huber loss matrix. After traversing all elements, an element-level Huber loss matrix is ​​generated. All elements in the Huber loss matrix are summed by addition and divided by the total number of elements in the matrix. The communication anomaly reconstruction error score is output.

[0026] The improved DVAE reconstruction error extraction process proposed in this step is similar to the traditional DVAE anomaly detection mechanism in that it is based on variational inference and generative reconstruction learning theory. That is, by mapping the input features to the latent probability space for parameterized sampling, the nonlinear approximation and reconstruction of the data is achieved by using the encoder-decoder network structure, and the feature deviation before and after reconstruction is used as the basis for judging the abnormal state.

[0027] The difference lies in that this invention breaks away from the limitations of traditional methods that rely solely on the isotropic Gaussian assumption to inject noise or ignore the local heteroscedasticity of features. Instead of directly generating globally fixed variance noise in traditional models, this invention adds an anisotropic noise modulation step. It calculates the local variance distribution along the feature dimension and takes its reciprocal to modulate the Gaussian noise matrix, rather than injecting homogeneous noise of uniform intensity. In the latent variable decoding step, a variance compensation coefficient vector is constructed using the ratio of the local variance of the noisy features to the local variance of the latent variables. This vector is then used to perform adaptive recalibration on the standard latent variable matrix, rather than directly inputting the latent variables into the decoder. Finally, in the error calculation step, a Huber loss combining quadratic and linear piecewise mappings is used to replace the traditional mean square error or absolute error, outputting a communication anomaly reconstruction error score instead of a single Euclidean distance mean.

[0028] The beneficial effects of this improvement are that, through anisotropic noise modulation and adaptive compensation of local variance, the local fluctuation characteristics of the input features can be deeply integrated into the generative game process of the latent space. This breaks the limitation of traditional DVAE in reconstructing the baseline distortion caused by the isotropic assumption in non-stationary communication data, and achieves accurate mapping from global fixed noise suppression to adaptive tracking of local feature structures. This design significantly enhances the model's ability to perceive non-stationary abrupt changes in multimodal communication data, and can more accurately capture reconstruction mismatch caused by anomalies under heteroscedastic conditions. The piecewise deviation calculation based on Huber loss effectively balances the robustness to small perturbations and the sensitivity to extreme anomalies, enhancing the reliability and discrimination accuracy of the reconstruction error score in complex communication anomaly detection tasks.

[0029] In this embodiment, S5 specifically includes: S51. Extract the communication anomaly contradiction score sequence and the communication anomaly reconstruction error score sequence arranged in chronological order within the historical time window. Perform a head-to-tail alignment and splicing operation with the communication anomaly contradiction score sequence as the first row and the communication anomaly reconstruction error score sequence as the second row to construct a multimodal joint score matrix with 2 rows and the number of columns equal to the total number of time steps. S52. Set the multimodal joint score matrix as the first input matrix and the multimodal joint score matrix as the second input matrix. Iterate through each element in the i-th row and j-th column of the first input matrix and each element in the k-th row and l-th column of the second input matrix. Multiply the element in the i-th row and j-th column of the first input matrix with the element in the k-th row and l-th column of the second input matrix. Use the result of the multiplication as the element in the (2 multiplied by i minus 2 plus k)-th row and (2 multiplied by j minus 2 plus l)-th column of the second-order intermediate matrix. After iterating through all rows and columns, generate the second-order intermediate matrix. Use this second-order intermediate matrix as the multimodal second-order joint score matrix. S53. The multimodal second-order joint score matrix is ​​decomposed into a first orthogonal matrix, a singular value diagonal matrix and a second orthogonal matrix by singular value decomposition. All non-zero values ​​on the main diagonal of the singular value diagonal matrix are extracted as all singular values. Each singular value in the total singular values ​​is traversed and an addition accumulation operation is performed. The sum obtained is used as the communication anomaly score.

[0030] In this embodiment, S6 specifically includes: S61. Extract the communication anomaly score sequence within the historical time window that does not include the current time. Set a sliding window with a single extraction length of 3. Align the starting position of the sliding window with the first value of the communication anomaly score sequence. Extract 3 consecutive values ​​within the coverage area of ​​the sliding window as the first row of the matrix. Slide the sliding window backward by 1 time step and extract 3 consecutive values ​​within the coverage area as the second row of the matrix. Slide the sliding window backward by 1 time step and extract 3 consecutive values ​​as the third row of the matrix. Perform a top-bottom concatenation operation on the first, second, and third rows to generate a Hankel matrix with 3 rows. Then, move the starting position of the sliding window backward by 1 time step. Repeat the above operation of extracting consecutive values ​​and concatenating top-bottom rows to generate the next Hankel matrix with 3 rows that moves with the time step. Iterate backward in sequence until the end of the sliding window exceeds the end of the communication anomaly score sequence. Arrange all the generated matrices in the order of the time steps to form a Hankel matrix sequence. S62. Traverse each Hankel matrix in the Hankel matrix sequence. For the currently traversed Hankel matrix, call the singular value decomposition algorithm to decompose it into a first orthogonal matrix, a singular value diagonal matrix, and a second orthogonal matrix. Extract all values ​​on the main diagonal of the singular value diagonal matrix as all singular values ​​of the current Hankel matrix. Traverse all extracted singular values ​​and perform addition to sum them. Use the summed value as the nuclear norm value of the current Hankel matrix at the corresponding time step. After traversing all matrices in the Hankel matrix sequence, obtain the nuclear norm values ​​of each time step. Concatenate all nuclear norm values ​​according to the chronological order of the corresponding time steps and output the nuclear norm sequence. S63. Traverse each intermediate nuclear norm value in the nuclear norm sequence except for the first and last two elements. For the currently traversed intermediate nuclear norm value, extract the nuclear norm value of its previous adjacent position in the nuclear norm sequence as the left neighbor value, and extract the nuclear norm value of its next adjacent position in the nuclear norm sequence as the right neighbor value. Determine whether the currently traversed intermediate nuclear norm value is simultaneously less than the left neighbor value and simultaneously less than the right neighbor value. If it is simultaneously less, mark the currently traversed intermediate nuclear norm value as a local minimum point. After traversing all intermediate nuclear norm values, extract all marked local minimum points, perform addition on all local minimum points, and divide by the total number of local minimum points to calculate the mean value, which is used as the communication anomaly scoring threshold. S64. Extract the latest communication anomaly score generated independently at the current moment as the value to be detected. Perform a subtraction operation between the value to be detected and the communication anomaly score threshold. Determine whether the difference obtained by the subtraction is greater than the value 0. If it is greater than the value 0, it is determined that there is a communication anomaly at the current moment, and the output result of the size comparison operation is set to the value 1 to represent that the anomaly detection result is abnormal. If it is less than or equal to the value 0, it is determined that there is no communication anomaly at the current moment, and the output result of the size comparison operation is set to the value 0 to represent that the anomaly detection result is normal. Output the anomaly detection result.

[0031] In this embodiment, S7 specifically includes: S71. Determine if the anomaly detection result is the value 1, which indicates an anomaly. If it is the value 1, trigger the anomaly localization process. Read the real element in the i-th row and j-th column of the multimodal fusion feature matrix, calculate the cosine function value of the real element as the complex real part, calculate the sine function value of the real element and multiply it by the preset imaginary unit to generate the complex imaginary part, add the complex real part and the complex imaginary part to construct the complex element, traverse all rows and columns to convert the multimodal fusion feature matrix into a complex domain matrix, construct the complex domain dimension reduction projection weight matrix, perform complex multiplication on the complex element in each row of the complex domain matrix and the complex element in the corresponding column of the complex domain dimension reduction projection weight matrix, and perform cumulative summation on the multiplied complex result column by column to output the multimodal complex projection feature matrix. S72. Read the cross-modal weight matrix, extract the first cross-modal sub-weight matrix and the second cross-modal sub-weight matrix, traverse each complex element in the i-th row and j-th column of the multimodal complex projection feature matrix, perform complex multiplication of the current complex element with the real element in the i-th row and j-th column of the first cross-modal sub-weight matrix to obtain the first mask complex number, perform complex multiplication of the current complex element with the real element in the i-th row and j-th column of the second cross-modal sub-weight matrix to obtain the second mask complex number, perform complex addition of the first mask complex number and the second mask complex number, and use the complex number result obtained by addition as the element in the i-th row and j-th column of the multimodal time-frequency mask suppression feature matrix. After traversing all rows and columns, output the multimodal time-frequency mask suppression feature matrix. S73. Set the fractional Fourier transform rotation order to 0.75, traverse each column of the multimodal time-frequency mask suppression feature matrix as the complex column vector to be transformed, calculate the index value of each complex element in the complex column vector to be transformed, substitute the index value into the formula of the fractional Fourier transform kernel function containing the rotation order of 0.75 to calculate and generate the rotation kernel complex matrix, perform complex domain matrix multiplication between the rotation kernel complex matrix and the complex column vector to be transformed, and use the complex column vector output by the multiplication as the element of the corresponding column in the transformed matrix. After traversing all columns, output the multimodal time-frequency transform matrix. S74. Traverse each row of the multimodal time-frequency transform matrix. For the current row, extract the real and imaginary parts of the complex element in the j-th column of the current row. Extract the real and imaginary parts of the complex element in the (j+1)-th column of the current row. Subtract the real part of the (j+1)-th column from the real part of the j-th column to obtain the real part difference. Subtract the imaginary part of the (j+1)-th column from the imaginary part of the j-th column to obtain the imaginary part difference. Use the real part difference as the real part and the imaginary part difference as the imaginary part to construct the feature dimension gradient complex number. Use the feature dimension gradient complex number as the element in the i-th row and j-th column of the multimodal fractional Fourier transform gradient matrix. After traversing all rows and columns, output the multimodal fractional Fourier transform gradient matrix. S75. Traverse each column of the multimodal fractional Fourier transform gradient matrix. For the current column, extract the real and imaginary parts of the complex element in the i-th row and j-th column. Square the real part and the imaginary part. Add the square root of the squared real and imaginary parts to extract the modulus of the complex element. Sum the modulus of all rows in the current column. Use the sum as the energy value of the current column at the corresponding time step in the multimodal instantaneous frequency mutation energy sequence. After traversing all columns, output the multimodal instantaneous frequency mutation energy sequence. S76. Traverse each intermediate energy value in the multimodal instantaneous frequency mutation energy sequence except for the first and last elements. For the current intermediate energy value, extract the energy value of its previous adjacent position in the sequence as the previous neighbor value, and extract the energy value of its next adjacent position in the sequence as the next neighbor value. Determine whether the current intermediate energy value is simultaneously greater than both the previous neighbor value and the next neighbor value. If both are greater, mark the current intermediate energy value as a maximum value point. Extract the index position of all marked maximum values ​​points in the multimodal instantaneous frequency mutation energy sequence, convert the index position back to the timestamp on the original time axis, and concatenate all the converted timestamps in chronological order to output the communication anomaly location result.

[0032] Example 1: To verify the feasibility of this invention in the operation and maintenance supervision of complex network environments, the method of this invention was applied to the intelligent monitoring system of the core nodes of the backbone network of a provincial telecommunications operator group company (hereinafter referred to as "Company C"). In traditional communication network anomaly monitoring systems, traffic baseline comparison based on fixed thresholds or simple single-modal log keyword matching are usually used. These methods not only have difficulty in accurately identifying hidden congestion and malicious tampering risks in cross-modal data interaction, but also cannot accurately obtain the deep time-frequency mutation characteristics when anomalies occur, which can easily lead to misjudgment or missed detection of sudden anomalies. To solve the above problems, Company C decided to adopt a communication anomaly detection method based on multimodal learning proposed in this invention.

[0033] During implementation, Company C first used probes deployed on core routers and log collection servers to acquire network traffic time-series data and device operation log data streams. After preprocessing operations such as zero-mean normalization and term frequency-inverse document frequency mapping, an aligned input space containing traffic autocorrelation matrices and log autocorrelation matrices was constructed. Simultaneously, Company C's network operations experts precisely labeled abnormal time periods and identified fault types in the collected multi-source historical data, serving as the benchmark for model training and anomaly scoring.

[0034] Company C constructs a complex-domain Hermitian matrix to map the traffic autocorrelation matrix and log autocorrelation matrix to their real and imaginary parts, respectively. It then utilizes an eigenvalue decomposition truncation mechanism to extract a multimodal joint feature basis containing core temporal correlations of both modalities, effectively eliminating redundant noise interference. Next, a cross-modal attention mechanism based on relative entropy minimization constraints calculates the mutual information attention weights of traffic and log features, generating a cross-modal weight matrix. This matrix, combined with the multimodal joint feature basis, outputs a multimodal fusion feature matrix. Subsequently, the multimodal fusion feature matrix and the traffic feature matrix are concatenated and input into an improved DVAE learning model. An anisotropic noise injection layer modulates Gaussian noise based on the inverse of local variance, and a conditional decoding layer recalibrates prior latent variables using variance compensation coefficients and performs upsampling decoding, outputting a high-fidelity denoised reconstructed feature matrix.

[0035] In the core detection and localization phase, this invention calculates the element-level Huber loss between the reconstructed matrix and the original matrix through a Huber deviation calculation layer. It then combines the trace norm of the multimodal conflict score with the nuclear norm of the second-order joint score matrix to simultaneously decouple and output the communication anomaly conflict score and the communication anomaly reconstruction error score. Subsequently, the anomaly scoring module extracts historical scores to construct a Hankel matrix sequence and solves for the nuclear norm sequence. Local minimum points are used as dynamically adaptive communication anomaly scoring thresholds, and anomaly detection results are determined accordingly. Finally, when an anomaly is triggered, the system performs complex-domain dimensionality reduction on the multimodal fusion feature matrix, uses a cross-modal weight matrix as a time-frequency mask to suppress interference, searches for the instantaneous frequency mutation energy extremum based on the fractional Fourier transform gradient, and outputs the communication anomaly localization result, achieving a closed-loop transition from anomaly detection to physical localization.

[0036] During implementation, Company C's technical team discovered that, compared to traditional manual inspection and conventional single-modal detection methods, the method of this invention significantly improves the accuracy and real-time performance of communication network anomaly identification. Traditional methods cannot quantify the degree of contradiction and reconstruction error in cross-modal data, and have poor time-frequency localization performance for non-stationary sudden anomalies. In contrast, the method of this invention effectively achieves accurate detection and high-resolution physical localization of complex communication anomalies through complex domain basis interleaving, improved DVAE variance compensation, and fractional-order time-frequency mask extremum search.

[0037] To further verify the actual performance of the method of the present invention, Company C conducted a detailed comparative test between the method of the present invention and the traditional method. The specific performance data is shown in Table 1: Table 1 Performance Comparison of Company C's Communication Network Anomaly Detection Methods

[0038] As shown in Table 1, the performance of the communication network anomaly monitoring system was comprehensively improved after applying the method of this invention. The accuracy of concealed anomaly detection increased from 78.4% of the traditional method to 95.6%, and the accuracy of anomaly reconstruction error quantification increased from 65.2% to 92.3%, significantly improving the accuracy of anomaly perception and providing a reliable basis for subsequent risk assessment. The false judgment rate of dynamic threshold adaptation decreased from 15.6% to 2.1%, effectively avoiding the storm of massive false alarms. The anomaly trigger response time was significantly shortened from 45 seconds to 4 seconds, significantly enhancing the system's timeliness. In addition, the malicious tampering event interception rate increased from 74.5% to 97.2%, and the manual cost of fault diagnosis decreased from 2 million yuan / year to 900,000 yuan / year, significantly reducing operation and maintenance costs. The satisfaction rate of operation and maintenance scheduling also significantly improved, from 82.0% to 96.5%.

[0039] Through the method of this invention, Company C has successfully achieved accurate detection and physical location of hidden anomalies and time-frequency mutations in complex communication networks, effectively reducing the risk of large-scale network paralysis, ensuring the safe operation of the backbone network, significantly improving the intelligence and digitalization level of communication network operation and maintenance, significantly reducing the workload of operation and maintenance personnel, enhancing the stability and robustness of the monitoring system, and providing strong technical support for the construction of intelligent communication networks.

[0040] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A communication anomaly detection method based on multimodal learning, characterized in that, Includes the following steps: S1. Construct the autocorrelation matrix of traffic and logs within the current time window into a complex field Hermitian matrix, extract the multimodal joint feature basis through eigenvalue decomposition and project it to output the traffic feature matrix and log feature matrix. S2. Map the traffic feature matrix and the log feature matrix to a query, key and value matrix. Calculate the similarity of the matrix under the constraint of minimizing relative entropy to generate a cross-modal weight matrix. Then, sum the weighted values ​​of the cross-modal weight matrix to output a multimodal fusion feature matrix. S3. Extract the cross-modal weight matrix to construct the conditional mutual information matrix of the log dimension, calculate the trace norm of the conditional mutual information matrix, and output the communication anomaly contradiction score sequence by taking the negative value. S4. The multimodal fusion feature matrix and the traffic feature matrix are concatenated to generate a multimodal joint feature matrix and input into the improved DVAE learning model. Local variance is used to guide anisotropic noise addition and variance condition adaptive compensation decoding. The deviation is reconstructed through Huber loss metric, and the communication anomaly reconstruction error score sequence is output. S5. Extract the communication anomaly contradiction score sequence and the communication anomaly reconstruction error score sequence within the historical time window, perform the Kronecker product and solve the nuclear norm to output the communication anomaly score; S6. Extract historical communication anomaly scores to construct a Hankel matrix sequence and solve the nuclear norm sequence. Take the local minimum point as the communication anomaly score threshold, compare it with the current communication anomaly score, and output the anomaly detection result. S7. If the anomaly detection result is abnormal, the multimodal fusion feature matrix is ​​reduced in dimensionality along the feature dimension in the complex domain. The cross-modal weight matrix is ​​used as a time-frequency mask. The instantaneous frequency change energy extremum is searched based on the fractional Fourier transform gradient, and the communication anomaly location result is output.

2. The communication anomaly detection method based on multimodal learning according to claim 1, characterized in that, S1 specifically includes: S11. Based on the traffic data within the current time window, calculate the cross-correlation coefficient of the lag order along the time dimension, and construct the traffic autocorrelation matrix; S12. Based on the log data within the current time window, extract the log template sequence and convert it into a log term frequency vector. Calculate the lag order cross-correlation coefficient along the time dimension and construct the log autocorrelation matrix. S13. Perform zero-mean standardization on the traffic autocorrelation matrix and the log autocorrelation matrix to generate a standardized traffic matrix and a standardized log matrix; S14. Construct a complex field Hermitian matrix by using the standardized traffic matrix as the real part and the standardized log matrix as the imaginary part; S15. Perform eigenvalue decomposition on the complex field Hermitian matrix to obtain a descending sequence of eigenvalues ​​and an eigenvector matrix. Calculate the ratio sequence of adjacent eigenvalues ​​in the eigenvalue sequence. Use the position where the ratio sequence first falls below a preset decay limit as the truncation index. Based on the truncation index, extract a feature submatrix composed of traffic features and log features from the eigenvector matrix. Use the feature submatrix as a multimodal joint feature basis. S16. Perform matrix multiplication on the standardized traffic matrix and the multimodal joint feature basis to output the dimensionality-reduced traffic feature matrix. Perform matrix multiplication on the standardized log matrix and the multimodal joint feature basis to output the dimensionality-reduced log feature matrix.

3. The communication anomaly detection method based on multimodal learning according to claim 1, characterized in that, The cross-modal weight matrix includes a first cross-modal sub-weight matrix and a second cross-modal sub-weight matrix: Perform matrix multiplication on the traffic feature matrix with the initialized first query weight matrix, first key weight matrix, and first value weight matrix respectively, and output the traffic query matrix, traffic key matrix, and traffic value matrix; perform matrix multiplication on the log feature matrix with the initialized second query weight matrix, second key weight matrix, and second value weight matrix respectively, and output the log query matrix, log key matrix, and log value matrix. Perform matrix multiplication on the transpose of the traffic query matrix and the log key matrix to output the original cross-score matrix; calculate the natural logarithm of the sum of the exponents of each row of the original cross-score matrix, and subtract the original cross-score matrix element by element from the natural logarithm as a penalty term to output the constrained cross-score matrix; input the constrained cross-score matrix into the Softmax activation function to output the first cross-modal sub-weight matrix. Perform a weighted summation operation on the first cross-modal sub-weight matrix and the log value matrix to output the multimodal log side context enhancement matrix; Perform matrix multiplication on the transpose of the log query matrix and the traffic key matrix to output the inverse original cross score matrix. Calculate the natural logarithm of the sum of the exponents of each row of the inverse original cross score matrix as a penalty term and subtract the inverse original cross score matrix element by element to output the inverse constrained cross score matrix. Input the inverse constrained cross score matrix into the Softmax activation function to output the second cross-modal subweight matrix. Perform a weighted summation operation between the second cross-modal sub-weight matrix and the flow value matrix to output the multimodal flow-side context enhancement matrix; The multimodal traffic-side context enhancement matrix and the multimodal log-side context enhancement matrix are concatenated to output a multimodal fusion feature matrix.

4. The communication anomaly detection method based on multimodal learning according to claim 1, characterized in that, S3 specifically includes: S31. Read the cross-modal weight matrix, extract the first cross-modal sub-weight matrix and the second cross-modal sub-weight matrix of the cross-modal weight matrix; calculate the mean of the first cross-modal sub-weight matrix along the row dimension, and output the log edge probability column vector; calculate the mean of the second cross-modal sub-weight matrix along the column dimension, and output the traffic edge probability row vector; perform element-wise multiplication on the transpose of the first cross-modal sub-weight matrix and the second cross-modal sub-weight matrix, and output the joint probability matrix; calculate the outer product of the log edge probability column vector and the traffic edge probability row vector, add a preset outer product divide-by-zero constant, and output the independent probability benchmark matrix; divide each element in the joint probability matrix by the corresponding element in the independent probability benchmark matrix, take the natural logarithm of the division result, and multiply it by the corresponding element in the joint probability matrix to output the multimodal conditional mutual information matrix; S32. Perform singular value decomposition on the multimodal conditional mutual information matrix, extract all singular values ​​after decomposition and calculate the sum of their absolute values, and output the trace norm of the multimodal conditional mutual information matrix. S33. Add a negative sign to the trace norm, and use the trace norm after adding the negative sign as the multimodal communication conflict score of the current time window. Arrange the multimodal communication conflict scores of all time windows in chronological order and output the communication anomaly conflict score sequence.

5. The communication anomaly detection method based on multimodal learning according to claim 1, characterized in that, The improved DVAE learning model includes an input projection layer, a joint probabilistic coding layer, an anisotropic noise injection layer, a conditional decoding layer, a feature reconstruction layer, and a Huber deviation calculation layer. The input projection layer is used to input the multimodal joint feature matrix into a linear mapping operator composed of a weight matrix and a bias vector to perform dimensionality reduction and nonlinear activation operations, and output a multimodal projection latent variable matrix. The joint probabilistic coding layer is used to input the multimodal projected latent variable matrix into two independent linear mapping operators composed of weight matrices and bias vectors to perform linear operations. The two output matrices are used as the latent variable prior mean matrix and the latent variable prior variance matrix, respectively. Based on the latent variable prior mean matrix and the latent variable prior variance matrix, a reparameterized sampling operation is performed to output the multimodal standard latent variable matrix. The anisotropic noise injection layer is used to generate a Gaussian noise matrix with the same dimension as the multimodal joint feature matrix. The local variance distribution matrix of each column of the multimodal joint feature matrix is ​​calculated along the feature dimension. After adding a preset matrix constant to each element of the local variance distribution matrix, the reciprocal operation is performed to output the reciprocal variance matrix. The Gaussian noise matrix and the reciprocal variance matrix are multiplied element by element and then superimposed on the multimodal joint feature matrix to output the multimodal noise-added joint feature matrix. The conditional decoding layer is used to calculate the local variance vectors of the multimodal denoising joint feature matrix and the multimodal standard latent variable matrix along the feature dimension, respectively. Each element in the local variance vector of the multimodal denoising joint feature matrix is ​​divided by the sum of the corresponding element in the local variance vector of the multimodal standard latent variable matrix and a preset element (excluding zero constants), outputting a variance compensation coefficient vector. Each column element in the multimodal standard latent variable matrix is ​​multiplied by the corresponding column coefficient in the variance compensation coefficient vector, outputting a multimodal recalibrated latent variable matrix. The multimodal recalibrated latent variable matrix is ​​input into a deconvolutional network and an activation function to perform upsampling operations, outputting a multimodal denoising latent variable decoding matrix. The feature reconstruction layer is used to input the multimodal denoising latent variable decoding matrix into a linear mapping operator composed of a weight matrix and a bias vector to perform a dimension mapping operation, restore the feature dimension to the dimension of the multimodal joint feature matrix, and output the multimodal reconstruction joint feature matrix. The Huber deviation calculation layer is used to calculate the absolute value of the element-wise difference between the multimodal reconstruction joint feature matrix and the multimodal joint feature matrix. Based on the preset threshold boundary, it sequentially performs a combination of quadratic and linear piecewise mappings to generate an element-level Huber loss matrix. It then performs a global mean operation on the matrix and outputs a communication anomaly reconstruction error score.

6. The communication anomaly detection method based on multimodal learning according to claim 1, characterized in that, S5 specifically includes: S51. Extract the communication anomaly contradiction score sequence and the communication anomaly reconstruction error score sequence within the historical time window, align and concatenate the communication anomaly contradiction score sequence and the communication anomaly reconstruction error score sequence according to the time step, and output the multimodal joint score matrix. S52. Calculate the Kronecker product between the multimodal joint score matrix and itself, and use the output of the Kronecker product operation as the second-order joint score matrix of the multimodal model. S53. Using the multimodal second-order joint score matrix as input, perform singular value decomposition to extract all singular values, use all singular values ​​as input to perform summation, and use the output of the summation as the communication anomaly score.

7. The communication anomaly detection method based on multimodal learning according to claim 1, characterized in that, S6 specifically includes: S61. Extract communication anomaly scores within the historical time window that do not include the current time. Input the communication anomaly scores into the sliding window, extract them sequentially by time step, and perform matrix permutation operations. Use the output of the matrix permutation operations as the Hankel matrix sequence. S62. Take each Hankel matrix in the Hankel matrix sequence as input and perform singular value decomposition to extract all singular values. Take all singular values ​​as input and perform summation. Arrange the output of the summation operation according to the time step as the nuclear norm sequence. S63. Perform a local minimum point search operation on the nuclear norm sequence as input, and use the output of the local minimum point search operation as the communication anomaly scoring threshold. S64. Extract the communication anomaly score calculated independently at the current time, perform a size comparison operation with the communication anomaly score threshold calculated independently at the current time as input, and use the output result of the size comparison operation as the anomaly detection result.

8. The communication anomaly detection method based on multimodal learning according to claim 1, characterized in that, Specifically, S7 includes: S71. When the anomaly detection result is abnormal, the multimodal fusion feature matrix is ​​used as input to perform a complex domain dimension reduction projection operation along the feature dimension, and the output result of the complex domain dimension reduction projection operation is used as the multimodal complex projection feature matrix. S72. Read the cross-modal weight matrix, extract the first cross-modal sub-weight matrix and the second cross-modal sub-weight matrix of the cross-modal weight matrix; take the multimodal complex projection feature matrix, the first cross-modal sub-weight matrix and the second cross-modal sub-weight matrix as input and perform element-wise multiplication operation, and take the output result of the element-wise multiplication operation as the multimodal time-frequency mask suppression feature matrix; S73. Perform fractional Fourier transform operation on the multimodal time-frequency mask suppression feature matrix as input, and use the output of the fractional Fourier transform operation as the multimodal time-frequency transform matrix. S74. Perform gradient calculation operation along the feature dimension using the multimodal time-frequency transformation matrix as input, and use the output result of the gradient calculation operation as the multimodal fractional Fourier transform gradient matrix. S75. The instantaneous frequency mutation energy calculation operation is performed using the multimodal fractional Fourier transform gradient matrix as input, and the output result of the instantaneous frequency mutation energy calculation operation is used as the multimodal instantaneous frequency mutation energy sequence. S76. The multimodal instantaneous frequency mutation energy sequence is used as input to perform an extreme point search operation, and the output result of the extreme point search operation is used as the communication anomaly location result.