Communication network node classification method and device, terminal equipment and storage medium
By encoding the topological connection data and traffic timing data of communication network nodes, combined with reconstruction and cluster loss functions, the problem of low node classification accuracy caused by inaccurate node feature representation in communication network is solved, and a higher node classification accuracy is achieved.
Patent Information
- Application Number
- CN202510717403.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-22
AI Technical Summary
Due to the inaccurate representation of the characteristic nodes of the communication network, the node classification accuracy of the distribution communication network is low.
By obtaining the topological connection data and traffic timing data of the node equipment of the communication network, using an automatic encoder for encoding processing, generating topological embedding representations and attribute embedding representations, combining the reconstruction loss function and the cluster loss function, model training and update model parameters until the loss function converges, and obtaining the classification results of the node.
The node classification accuracy of the distribution communication network is improved, the optimal performance of the automatic encoder is ensured, and the classification results of the node equipment of the communication network are obtained through information mining.
Smart Images

Figure CN120358148A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and in particular, to a method, device, terminal device, and storage medium for classifying communication network nodes. Background Art
[0002] In the field of communication technologies for smart power grids, fault location in distribution communication networks is one of the key technologies to ensure the stable operation of the power grid. With the complexity and high voltage of the power system, the importance of power line fault detection and location problems has become increasingly prominent. As the power system becomes more and more complex, the feature representation of each communication node becomes less and less accurate. Due to the inaccurate feature representation of communication network equipment nodes, the problem of low accuracy in classifying nodes in the distribution communication network occurs. Summary of the Invention
[0003] The present invention provides a method, device, terminal device, and storage medium for classifying communication network nodes, which can solve the technical problem of low accuracy in classifying nodes in the distribution communication network due to inaccurate feature representation of communication network nodes.
[0004] The method for classifying communication network nodes provided by the present invention includes:
[0005] Obtain the topological connection data and traffic time series data of each communication network node device;
[0006] Based on the topological connection data and traffic time series data of each communication network node device, repeatedly perform model training operations. When the sum of the loss function values converges, stop performing the model training operations, and based on the output of the last model training operation, obtain the classification result of each communication network node device;
[0007] Wherein, the model training operation includes:
[0008] Through an autoencoder with current model parameters, respectively perform encoding processing on each topological connection data to obtain the communication network topological embedding representation of each communication network node device, respectively perform encoding processing on each traffic time series data to generate the communication network node attribute embedding representation of each communication network node device, and calculate the reconstruction loss function value based on the communication network topological embedding representation, the communication network node attribute embedding representation, and the reconstruction loss function;
[0009] According to the communication network topological embedding representation and the communication network node attribute embedding representation, form the topological and traffic comprehensive embedding representation vector of each communication network node device;
[0010] Respectively perform clustering on the topological and traffic comprehensive embedding representation vectors of each communication network node device to obtain the clustering result of each communication network node device, and calculate the clustering loss function value based on the clustering result and the clustering loss function;
[0011] Calculate the sum of the loss functions based on the clustering loss function value and the reconstruction loss function value, and make a judgment on the sum of the loss functions: If the sum of the loss functions converges, update the model parameters of the autoencoder, use the updated model parameters as the model parameters of the autoencoder in the next model training operation, and perform the next model training operation; if the sum of the loss functions converges, use the clustering result of each current communication network node device as the output.
[0012] Furthermore, the model parameters include: the first model parameter and the second model parameter; the autoencoder includes: a basic autoencoder and a graph autoencoder; the reconstruction loss function includes: the first reconstruction loss function and the second reconstruction loss function; the reconstruction loss function values include: the first reconstruction function loss value and the second reconstruction function loss value;
[0013] Through the autoencoder with the current model parameters, encode each of the topology connection data to obtain the communication network topology embedding representation of each communication network node device, and encode each of the traffic time series data to generate the communication network node attribute embedding representation of each communication network node device. Based on the communication network topology embedding representation, the communication network node attribute embedding representation, and the reconstruction loss function, calculate the reconstruction loss function value, including:
[0014] Through the graph autoencoder corresponding to the current first model parameter, encode each of the topology connection data to generate the communication network topology embedding representation of each communication network node device, and calculate the first reconstruction loss function value of the graph autoencoder based on the communication network topology embedding representation and the first reconstruction loss function;
[0015] Through the basic autoencoder corresponding to the current second model parameter, encode each of the traffic time series data to generate the communication network node attribute embedding representation of each communication network node device, and calculate the second reconstruction loss function value of the basic autoencoder based on the communication network node attribute embedding representation and the second reconstruction loss function.
[0016] Furthermore, the calculating the first reconstruction loss function value of the graph autoencoder based on the communication network topology embedding representation and the first reconstruction loss function includes:
[0017] Through the graph autoencoder, decode each of the communication network topology embedding representations to obtain the communication network topology reconstruction data of each communication network node device, and substitute the communication network topology and the communication network topology reconstruction data into the first reconstruction loss function to calculate the first reconstruction loss function value.
[0018] Further, calculating the second reconstruction loss function value of the basic autoencoder based on the communication network node attribute embedding representation and the second reconstruction loss function includes:
[0019] Through the basic autoencoder, decoding processing is respectively performed on each communication network node attribute embedding representation to obtain the communication network node attribute reconstruction data of each communication network node device, and substituting the traffic time series data and the communication network node attribute reconstruction data into the second reconstruction loss function to calculate the second reconstruction loss function value.
[0020] Further, the first reconstruction loss function includes:
[0021]
[0022] In the formula, L gae is the first reconstruction loss function value, N is the number of communication network device nodes, Ai is the communication network topology of the i-th communication network node device, is the communication network topology reconstruction data of the i-th communication network node device.
[0023] Further, the second reconstruction loss function includes:
[0024]
[0025] In the formula, L ae is the second reconstruction loss function value, N is the number of communication network device nodes, X is the traffic time series data, is the communication network node attribute reconstruction data.
[0026] Further, the clustering loss function includes:
[0027]
[0028] In the formula, L KL is the clustering loss function value, p ij is the probability that node i belongs to cluster j, q ij is the similarity between the feature representation z i of node i and the cluster center u j P is the set of p ij and Q is the set of q ij of.
[0029] Another embodiment of the present invention further provides a classification device for communication network nodes, including: a data acquisition module and a data classification module;
[0030] The data acquisition module is used to acquire the topological connection data and traffic time series data of each communication network node device;
[0031] The data classification module is configured to repeatedly perform model training operations based on the topological connection data and traffic time-series data of each communication network node device. When the sum of the loss function values converges, it stops performing the model training operations and obtains the classification result of each communication network node device based on the output of the last model training operation;
[0032] Among them, the model training operation includes:
[0033] Through the autoencoder with the current model parameters, encode each piece of the topological connection data respectively to obtain the communication network topological embedding representation of each communication network node device, encode each piece of the traffic time-series data respectively to generate the communication network node attribute embedding representation of each communication network node device, and calculate the reconstruction loss function value based on the communication network topological embedding representation, the communication network node attribute embedding representation, and the reconstruction loss function;
[0034] According to the communication network topological embedding representation and the communication network node attribute embedding representation, form the topological and traffic comprehensive embedding representation vector of each communication network node device;
[0035] Cluster the topological and traffic comprehensive embedding representation vectors of each communication network node device respectively to obtain the clustering result of each communication network node device, and calculate the clustering loss function value based on the clustering result and the clustering loss function;
[0036] Calculate the sum of the loss functions according to the clustering loss function value and the reconstruction loss function value, and judge the sum of the loss functions: If the sum of the loss functions converges, update the model parameters of the autoencoder, use the updated model parameters as the model parameters of the autoencoder in the next model training operation, and perform the next model training operation; If the sum of the loss functions converges, use the current clustering result of each communication network node device as the output.
[0037] Another embodiment of the present invention also provides a terminal device, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the steps of the classification method of the communication network node provided by the present invention.
[0038] Another embodiment of the present invention also provides a computer-readable storage medium item, including: a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the steps of the classification method of the communication network node provided by the present invention.
[0039] By implementing the present invention, the following beneficial effects are achieved:
[0040] The present invention performs model training operations on the topological connection data and traffic time series data of each communication network node device. The topological connection data is encoded through an autoencoder to obtain a communication network topology embedding representation, and the traffic time series data is encoded through an autoencoder to obtain a communication network node attribute embedding representation. Then, the reconstruction loss function value is calculated based on the data before and after the encoding process. The communication network topology embedding representation and the communication network node attribute embedding representation are combined to form a topology and traffic comprehensive embedding representation vector. Clustering is performed on the topology and traffic comprehensive embedding representation vector, and clustering is performed based on the features of the two embedding representations to obtain the clustering result of the communication network node device. Then, the clustering loss function value corresponding to the clustering result is calculated. Thus, by determining whether the sum of the loss functions corresponding to the clustering loss function value and the reconstruction loss function value converges, the encoding result and clustering result of the autoencoder are supervised, and based on the judgment result, the model parameters are updated. When the sum of the loss functions converges, an accurate classification result of the communication network node device is obtained. The present invention uses the communication network topology embedding representation and the communication network node attribute embedding representation as the feature representations of the communication network device nodes, and updates the model parameters of the autoencoder participating in the classification through the loss function, so as to ensure the optimal performance of the autoencoder when the sum of the loss functions converges. Information mining is performed on the communication network topology embedding representation and the communication network node attribute embedding representation through the autoencoder with optimal performance, which greatly improves the node classification accuracy of the distribution communication network. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions of the present application, the drawings required for the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 It is a flowchart of a method for classifying communication network nodes provided by an embodiment of the present invention;
[0043] Figure 2 It is a structural diagram of a device for classifying communication network nodes provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs; the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above description of the drawings are intended to cover non-exclusive inclusion.
[0046] In the description of the embodiments of this application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity, specific order, or primary-secondary relationship of the indicated technical features. In the description of the embodiments of this application, the meaning of "a plurality" is more than two, unless otherwise specifically defined.
[0047] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0048] In the description of the embodiments of this application, the term "and / or" is merely a description of the associated relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0049] In the description of the embodiments of this application, the term "a plurality" refers to more than two (including two). Similarly, "multiple groups" refers to more than two groups (including two groups), and "multiple pieces" refers to more than two pieces (including two pieces).
[0050] In the description of the embodiments of this application, unless otherwise clearly specified and limited, technical terms such as "installation", "connection", "connection", and "fixation" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can also be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and can be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of this application can be understood according to specific circumstances.
[0051] See Figure 1, to solve the technical problem that the node classification accuracy of the distribution communication network is low due to inaccurate feature representation of communication network nodes, a classification method for communication network nodes provided by an embodiment of the present invention includes:
[0052] 101. Obtain the topological connection data and traffic time series data of each communication network node device.
[0053] 102. Based on the topological connection data and traffic time series data of each communication network node device, repeatedly execute the model training operation. When the sum of the loss function values converges, stop executing the model training operation, and based on the output of the last model training operation, obtain the classification result of each communication network node device;
[0054] Wherein, the model training operation includes:
[0055] Through the autoencoder with the current model parameters, respectively encode each topological connection data to obtain the communication network topological embedding representation of each communication network node device, respectively encode each traffic time series data to generate the communication network node attribute embedding representation of each communication network node device, and calculate the reconstruction loss function value based on the communication network topological embedding representation, the communication network node attribute embedding representation, and the reconstruction loss function;
[0056] According to the communication network topological embedding representation and the communication network node attribute embedding representation, form the topological and traffic comprehensive embedding representation vector of each communication network node device;
[0057] Cluster the topological and traffic comprehensive embedding representation vectors of each communication network node device respectively to obtain the clustering result of each communication network node device, and calculate the clustering loss function value based on the clustering result and the clustering loss function;
[0058] Calculate the sum of the loss functions according to the clustering loss function value and the reconstruction loss function value, and judge the sum of the loss functions: If the sum of the loss functions converges, update the model parameters of the autoencoder, use the updated model parameters as the model parameters of the autoencoder in the next model training operation, and execute the next model training operation; If the sum of the loss functions converges, use the current clustering result of each communication network node device as the output.
[0059] In a specific embodiment, the traffic time series data is specifically: The traffic time series data of the communication network node device is a kind of dynamic network traffic information recorded in the form of time series, and its core feature is that the data points are strictly arranged according to the time stamps and contain multi-dimensional network state indicators.
[0060] Furthermore, the model parameters include: a first model parameter and a second model parameter; the autoencoder includes: a basic autoencoder and a graph autoencoder; the reconstruction loss function includes: a first reconstruction loss function and a second reconstruction loss function; the reconstruction loss function values include: a first reconstruction function loss value and a second reconstruction function loss value;
[0061] Encoding each of the topology connection data through the autoencoder with the current model parameters to obtain the communication network topology embedding representation of each communication network node device, encoding each of the traffic time series data to generate the communication network node attribute embedding representation of each communication network node device, and calculating the reconstruction loss function value based on the communication network topology embedding representation, the communication network node attribute embedding representation, and the reconstruction loss function, includes:
[0062] Encoding each of the topology connection data through the graph autoencoder corresponding to the current first model parameter to generate the communication network topology embedding representation of each communication network node device, and calculating the first reconstruction loss function value of the graph autoencoder based on the communication network topology embedding representation and the first reconstruction loss function;
[0063] Encoding each of the traffic time series data through the basic autoencoder corresponding to the current second model parameter to generate the communication network node attribute embedding representation of each communication network node device, and calculating the second reconstruction loss function value of the basic autoencoder based on the communication network node attribute embedding representation and the second reconstruction loss function.
[0064] Furthermore, calculating the first reconstruction loss function value of the graph autoencoder based on the communication network topology embedding representation and the first reconstruction loss function includes:
[0065] Decoding each of the communication network topology embedding representations through the graph autoencoder to obtain the communication network topology reconstruction data of each communication network node device, substituting the communication network topology and the communication network topology reconstruction data into the first reconstruction loss function, and calculating the first reconstruction loss function value.
[0066] In a specific embodiment, the topology connection data can be encoded through the graph attention network layer (GraphAttention Network, GAT) in the graph autoencoder GAE, and the graph representation of the communication network node device i is denoted as It can be obtained from the representation of its previous layer:
[0067]
[0068] where W is a trainable weight matrix for linear feature transformation, Ni is the set of neighbor nodes of node i, and α ij is the attention coefficient between nodes i and j. The attention coefficient measures the degree of association or importance of different neighbors to the target node;
[0069] To calculate the attention coefficient α ij , first, the mutual relationship between nodes i and j is learned through a linear transformation:
[0070]
[0071] where a is the weight vector, and || is the vector concatenation operation, and are the original features of nodes i and j respectively, and [·||·] concatenates the transformed features of vertices i and j. Next, the coefficients m of all neighbors of node i are normalized through the softmax function with the activation function LeakyReLU(·), and then the attention coefficient α of nodes ij can be obtained ij : ij :
[0072]
[0073] It should be noted that the original features of a node refer to the in and out traffic of the node, but are not limited to traffic. In a switch communication network, in addition to in and out traffic, there will be many time series data such as delay, bandwidth, and packet loss, all of which can be part of the original features of the node. The topological connection relationship is given by the adjacency matrix of the switch network.
[0074] To fully explore the complex relationships in graph data, this method takes into account higher-order neighbors rather than first-order neighbor nodes. A proximity matrix M is introduced, and the higher-order topological correlation is calculated by considering t-order neighbor nodes:
[0075] M = (B + B 2 +... + B t ) / t;
[0076] Here B is the transition matrix. For the matrix element B ij , if e ij ∈ edge set E, then B ij = 1 / d ij , otherwise B ij = 0. For an undirected graph, here d ij is the degree of node i or node j. Mij reflects the t-order (t value is set according to needs) topological correlation from node j to node i. The attention coefficient α after adding the proximity matrix M ij is updated to:
[0077]
[0078] The GAT contains L layers connected in sequence. Therefore, the rewritten representation of embedding the communication network topology of communication network node i is as follows:
[0079]
[0080] where and are the attention coefficient and the feature transformation weight of the th layer respectively, and σ(·) is an activation function, such as the ReLU function or the Tanh function. Specifically, is the input of the GAT encoder.
[0081] The graph decoder after the graph encoder is used to reconstruct the graph topology. The topology of the graph is represented by the adjacency matrix A. The communication network topology embedding representation Z obtained by the graph encoder is used as the input of this module, and a simple inner product decoder is used to reconstruct the communication network topology reconstruction data
[0082]
[0083] Then the reconstruction loss function (i.e., the first reconstruction loss function) can be defined by cross entropy:
[0084]
[0085] Taking minimizing this loss function as the objective, the graph encoder is trained round by round to generate a more appropriate feature representation.
[0086] Furthermore, calculating the second reconstruction loss function value of the basic autoencoder based on the communication network node attribute embedding representation and the second reconstruction loss function includes:
[0087] Through the basic autoencoder, each communication network node attribute embedding representation is decoded respectively to obtain the communication network node attribute reconstruction data of each communication network node device, and the traffic time series data and the communication network node attribute reconstruction data are substituted into the second reconstruction loss function to calculate the second reconstruction loss function value.
[0088] In a specific embodiment, fully learning the attribute information of the data is of great significance for the graph clustering task, because this information can provide important decision-making information for node classification. For this purpose, a basic autoencoder (AE) and multiple fully connected layers are used to learn the hierarchical attribute information from the communication data and embed the learned content into a compact feature representation in a low-dimensional space. The basic autoencoder model consists of two important parts: an encoder function and a decoder function. The encoder function projects the traffic time series data into the latent space In
[0089] X′ = f e (W e , X);
[0090] Among them, the conversion result X′ is called the communication network node attribute embedding representation, and its dimension satisfies d′ < d. f e represents the encoder function, and W e is the weight matrix, which can perform a linear transformation on the traffic time series data. For simplicity, it is assumed that all weight matrices {W} in the formulas of this patent carry biases {b}.
[0091] The decoder is designed after the encoder, and based on X′, the reconstructed data of the communication network node attributes is
[0092]
[0093] where f d represents the decoder function, and W d is the weight matrix of this decoder. In order to obtain the hierarchical attribute information of the data, we construct the encoder and decoder through L fully connected layers, where the specific depth layer adaptively processes the corresponding hierarchical information hidden in the data. In this encoder, the hierarchical attribute information representation learned by the neural layer can be expressed as
[0094]
[0095] where is the weight matrix of the layer, and σ(·) is a non-linear activation function, such as ReLU or Tanh. is the attribute information representation of the previous layer. In particular, represents the traffic time series data X, represents the encoder output X′. Similarly, the decoder function f d after the encoder also contains L fully connected layers, and the hierarchical attribute information representation of the layer is denoted as
[0096]
[0097] where is the weight matrix of the layer decoder layer. Similarly, in particular, represents the feature representation X′ generated by the encoder, represents the decoder output, that is, the reconstructed data of the communication network node attributes
[0098] The representations of each neural layer show the hierarchical nature of the features. Therefore, the attribute information extracted by the autoencoder layers at different network depths is labeled with hierarchical meanings and properties. In this way, it provides rich semantic information for distinguishing faulty nodes from other nodes in the downstream clustering task. To ensure the effectiveness of the data representation in the latent space, this module defines the reconstruction loss of the basic AE (i.e., the second reconstruction loss function) through the mean squared error (MSE):
[0099]
[0100] where N is the number of device nodes in the distribution communication network. The reconstruction loss is used as the objective function to minimize the difference between the reconstructed data and the original input data.
[0101] The acquisition of traffic time-series data is specifically as follows: The time-series data extracted from the network management system of a company's switch contains data of each node over a relatively long period of time, which is then used as traffic time-series data.
[0102] Furthermore, the first reconstruction loss function includes:
[0103]
[0104] where L gae is the value of the first reconstruction loss function, N is the number of device nodes in the communication network, Ai is the communication network topology of the i-th communication network node device, is the reconstructed data of the communication network topology of the i-th communication network node device.
[0105] Furthermore, the second reconstruction loss function includes:
[0106]
[0107] where L ae is the value of the second reconstruction loss function, N is the number of device nodes in the communication network, X is the traffic time-series data, is the reconstructed data of the communication network node attributes.
[0108] Furthermore, the clustering loss function includes:
[0109]
[0110] where L KL is the value of the clustering loss function, p ij is the probability that node i belongs to cluster j, q ij is the feature representation z of node i i and the cluster center u jThe similarity between, where P is p ij The set of, where Q is q ij The set of.
[0111] In a specific embodiment, the communication network node attribute embedding representation H of the basic autoencoder AE e Is combined with the communication network topology embedding representation output from the graph autoencoder GAT To enhance the feature representation, as follows:
[0112]
[0113] Among them, λ balances the weights of the two modules, which indicates the proportion of the feature attributes in each neural layer of AE that are transmitted to the corresponding layer of GAT. The hierarchical attribute information is integrated into the graph encoder layer by layer. Then we use As the input of the Layer in GAT to generate the representation of the next layer:
[0114]
[0115] The basic AE can transmit the attribute information of different layers to the graph model, helping it to fully learn the representation of the topology and traffic comprehensive embedding representation vector Z = Z i For clustering tasks.
[0116] In a specific embodiment, λ is a hyperparameter in the deep learning model and can be trained through multiple attempts. It can also be set according to expert experience, such as 0.1
[0117] Two autoencoders (basic AE and graph GAE) are organically combined into a unified model, aiming to generate enhanced feature representations in the latent space, and then perform subsequent node classification clustering, fault detection, etc. tasks based on the topology and traffic comprehensive embedding representation vector Z. However, both the model and the downstream clustering are unsupervised, resulting in a lack of strong supervision information to guide the entire optimization process.
[0118] A self-supervised mechanism is designed to unify the model and the clustering task. For the two autoencoders, the reconstruction loss is minimized to optimize their feature representations. To generate more reliable guidance for node clustering, the KL divergence (Kullback-Leibler divergence) is used as the objective function. The specific content is as follows.
[0119] First, the t-distribution is used to calculate the similarity between the topology and traffic comprehensive embedding representation vector z of the communication network node i i And the clustering center u j As follows:
[0120]
[0121] where v is the degree of freedom of the t-distribution, and these cluster centers u are initialized using Z-based K-means. The similarity q ij is also known as soft cluster assignment and is regarded as the probability of assigning node i to cluster j.
[0122] Then, the target cluster assignment is defined based on the soft cluster assignment as follows:
[0123]
[0124] P = {p i,j | 0 ≤ i, j ≤ N};
[0125] Square and normalize q ij to make the feature representation closer to the cluster center.
[0126] Finally, the KL divergence is used to make the soft cluster assignment close to the target assignment:
[0127]
[0128] The KL divergence is regarded as the clustering loss. The overall learning objective consists of three parts, namely the reconstruction loss of the basic autoencoder AE and the reconstruction loss of the graph autoencoder GAE, as well as the clustering loss:
[0129] L = L ae + L gae + εL KL ;
[0130] where the hyperparameter ε balances the weights of the reconstruction and clustering losses. The parameters of the neural network are jointly optimized and updated by stochastic gradient descent. When the model is trained to convergence, we can obtain the final predicted assignment of each communication network node according to the soft assignment Q
[0131]
[0132] As Figure 2 shown, based on the above method item embodiments, corresponding device item embodiments are provided;
[0133] An embodiment of the present invention provides a classification device for communication network nodes, including: a data acquisition module 201 and a data classification module 202;
[0134] The data acquisition module is used to acquire the topological connection data and traffic time series data of each communication network node device;
[0135] The data classification module is configured to repeatedly perform model training operations based on the topological connection data and traffic time series data of each communication network node device. When the sum of the loss function values converges, it stops performing the model training operations and obtains the classification result of each communication network node device based on the output of the last model training operation.
[0136] Among them, the model training operation includes:
[0137] Through the autoencoder with the current model parameters, encode each piece of the topological connection data respectively to obtain the communication network topological embedding representation of each communication network node device, encode each piece of the traffic time series data respectively to generate the communication network node attribute embedding representation of each communication network node device, and calculate the reconstruction loss function value based on the communication network topological embedding representation, the communication network node attribute embedding representation, and the reconstruction loss function.
[0138] According to the communication network topological embedding representation and the communication network node attribute embedding representation, form the topological and traffic comprehensive embedding representation vector of each communication network node device.
[0139] Cluster the topological and traffic comprehensive embedding representation vectors of each communication network node device respectively to obtain the clustering result of each communication network node device, and calculate the clustering loss function value based on the clustering result and the clustering loss function.
[0140] Calculate the sum of the loss functions according to the clustering loss function value and the reconstruction loss function value, and judge the sum of the loss functions: If the sum of the loss functions converges, update the model parameters of the autoencoder, use the updated model parameters as the model parameters of the autoencoder in the next model training operation, and perform the next model training operation; If the sum of the loss functions converges, use the current clustering result of each communication network node device as the output.
[0141] It can be understood that the above device item embodiments correspond to the method item embodiments of the present invention, and can implement the classification method of communication network nodes provided in any one of the above method item embodiments of the present invention.
[0142] It should be noted that the device embodiments described above are only illustrative. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement without creative efforts.
[0143] Based on the embodiments of the classification method of the above communication network nodes, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the classification method of the communication network nodes according to any embodiment of the present invention is implemented.
[0144] Exemplarily, in this embodiment, the computer program may be divided into one or more modules. The one or more modules are stored in the memory and executed by the processor to complete the present invention. The one or more module elements may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device.
[0145] The terminal device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory.
[0146] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the terminal device and connects various parts of the entire terminal device through various interfaces and lines.
[0147] Based on the above method item embodiments, another embodiment of the present invention provides a computer-readable storage medium, including a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the classification method of the communication network nodes described in any of the above method item embodiments of the present invention.
[0148] Among them, if the modules / units integrated in the device / terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0149] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A classification method for communication network nodes, characterized in that, Including: Obtaining the topological connection data and traffic time-series data of each communication network node device; Based on the topological connection data and traffic time-series data of each communication network node device, repeatedly performing model training operations. When the sum of the loss function values converges, stop performing the model training operations, and based on the output of the last model training operation, obtain the classification result of each communication network node device; Wherein, the model training operation includes: Through the autoencoder with the current model parameters, respectively encoding each topological connection data to obtain the communication network topological embedding representation of each communication network node device, respectively encoding each traffic time-series data to generate the communication network node attribute embedding representation of each communication network node device, and calculating the reconstruction loss function value based on the communication network topological embedding representation, the communication network node attribute embedding representation, and the reconstruction loss function; According to the communication network topological embedding representation and the communication network node attribute embedding representation, forming the topological and traffic comprehensive embedding representation vector of each communication network node device; Respectively clustering the topological and traffic comprehensive embedding representation vectors of each communication network node device to obtain the clustering result of each communication network node device, and calculating the clustering loss function value based on the clustering result and the clustering loss function; Calculating the sum of the loss functions according to the clustering loss function value and the reconstruction loss function value, and judging the sum of the loss functions: If the sum of the loss functions converges, update the model parameters of the autoencoder, use the updated model parameters as the model parameters of the autoencoder in the next model training operation, and perform the next model training operation; If the sum of the loss functions converges, use the current clustering result of each communication network node device as the output.
2. The classification method of a communication network node according to claim 1, characterized in that, The model parameters include: the first model parameter and the second model parameter; the autoencoder includes: the basic autoencoder and the graph autoencoder; the reconstruction loss function includes: the first reconstruction loss function and the second reconstruction loss function; the reconstruction loss function value includes: the first reconstruction function loss value and the second reconstruction function loss value; The step of, through the autoencoder with the current model parameters, respectively encoding each topological connection data to obtain the communication network topological embedding representation of each communication network node device, respectively encoding each traffic time-series data to generate the communication network node attribute embedding representation of each communication network node device, and calculating the reconstruction loss function value based on the communication network topological embedding representation, the communication network node attribute embedding representation, and the reconstruction loss function, includes: Through the graph autoencoder corresponding to the current first model parameter, respectively encoding each topological connection data to generate the communication network topological embedding representation of each communication network node device, and calculating the first reconstruction loss function value of the graph autoencoder based on the communication network topological embedding representation and the first reconstruction loss function; Encode each piece of the traffic time series data through the basic autoencoder corresponding to the current second model parameters to generate the communication network node attribute embedding representation of each communication network node device, and calculate the second reconstruction loss function value of the basic autoencoder based on the communication network node attribute embedding representation and the second reconstruction loss function.
3. The classification method of the communication network node according to claim 2, characterized in that, The calculating of the first reconstruction loss function value of the graph autoencoder based on the communication network topology embedding representation and the first reconstruction loss function includes: Decode each piece of the communication network topology embedding representation through the graph autoencoder to obtain the communication network topology reconstruction data of each communication network node device, substitute the communication network topology and the communication network topology reconstruction data into the first reconstruction loss function, and calculate the first reconstruction loss function value.
4. The classification method of a communication network node according to claim 3, characterized in that, The calculating of the second reconstruction loss function value of the basic autoencoder based on the communication network node attribute embedding representation and the second reconstruction loss function includes: Decode each piece of the communication network node attribute embedding representation through the basic autoencoder to obtain the communication network node attribute reconstruction data of each communication network node device, substitute the traffic time series data and the communication network node attribute reconstruction data into the second reconstruction loss function, and calculate the second reconstruction loss function value.
5. The classification method of a communication network node according to claim 4, characterized in that, The first reconstruction loss function includes: where L gae is the value of the first reconstruction loss function, N is the number of communication network device nodes, and Ai is the communication network topology of the i-th communication network node device, is the communication network topology reconstruction data of the i-th communication network node device.
6. The classification method of a communication network node according to claim 4, characterized in that, The second reconstruction loss function includes: where L ae is the value of the second reconstruction loss function, N is the number of communication network device nodes, X is the traffic time series data, is the communication network node attribute reconstruction data.
7. The classification method of a communication network node according to claim 4, characterized in that, The clustering loss function includes: Where L KL is the value of the clustering loss function, p ij is the probability that node i belongs to cluster j, q ij is the feature representation z i of node i and the similarity j with the cluster center u ij P is the set of p ij , and Q is the set of q 8. A classification device for a communication network node, characterized in that, Includes: A data acquisition module and a data classification module; The data acquisition module is used to acquire the topological connection data and the traffic time series data of each communication network node device; The data classification module is used to repeatedly perform model training operations based on the topological connection data and the traffic time series data of each communication network node device, stop performing the model training operations when the sum of the loss function values converges, and obtain the classification result of each communication network node device based on the output of the last model training operation; Wherein, the model training operation includes: Encode each piece of the topological connection data through the autoencoder with the current model parameters to obtain the communication network topology embedding representation of each communication network node device, encode each piece of the traffic time series data to generate the communication network node attribute embedding representation of each communication network node device, and calculate the reconstruction loss function value based on the communication network topology embedding representation, the communication network node attribute embedding representation and the reconstruction loss function; According to the communication network topology embedding representation and the communication network node attribute embedding representation, form the topological and traffic comprehensive embedding representation vector of each communication network node device; Cluster the topological and traffic comprehensive embedding representation vectors of each communication network node device respectively to obtain the clustering result of each communication network node device, and calculate the clustering loss function value based on the clustering result and the clustering loss function; Calculate the sum of the loss functions based on the clustering loss function value and the reconstruction loss function value, and make a judgment on the sum of the loss functions: If the sum of the loss functions converges, update the model parameters of the autoencoder, use the updated model parameters as the model parameters of the autoencoder in the next model training operation, and perform the next model training operation; If the sum of the loss functions converges, use the clustering result of each current communication network node device as the output.
9. A terminal device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the classification method of the communication network node according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, It includes: A stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the classification method of the communication network node according to any one of claims 1-7.