A Method for Identifying Service Network Function Nodes Based on Subgraph Information Bottleneck
By adopting a bottleneck method based on subgraph information in the service resource network, combining the subgraph multi-level node semantic information and multi-scale topological information, the feature encoding model is trained, and the problem of inaccurate feature capture of functional nodes in the prior art is solved, and the accuracy of node recognition is improved.
Patent Information
- Application Number
- CN202510096925.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The prior art is difficult to accurately capture the characteristics of functional nodes in the service resource network, resulting in a low node recognition accuracy.
Using a method based on subgraph information bottleneck, combining the multi-level node semantic information and multi-scale topological information of subgraphs from the perspective of information theory, the network node features with multi-scale constraint information in hidden space are obtained through the joint optimization function of multi-order subgraph information bottleneck and the feature coding model trained by the global topological constraint function.
The accuracy of the identification of service network nodes is improved, and the characteristics of functional nodes in the service network can be captured more accurately.
Smart Images

Figure CN119557748B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network technologies, and in particular, to a method for identifying service network function nodes based on subgraph information bottleneck. Background Art
[0002] A service resource network refers to a comprehensive computer network system integrating an information service system, a network management system, and resource information users. By integrating different types of resources, this system realizes the efficient management and dynamic allocation of resources, thereby providing comprehensive information services to various resource information users, and further enhancing the efficiency in aspects such as data storage, network services, and system management. The service resource network allocates and schedules resources according to real-time demands, significantly improving resource utilization. An intelligent service resource network can also monitor and analyze resource usage in real time, predict future demands, and thus help managers formulate resource plans in advance.
[0003] In a service resource network, a typical service node is usually an integrated and multi-functional service node. For such multi-functional service nodes, it is necessary to perform function identification of the service resource network to obtain function labels in different application scenarios, so as to prepare data identification for subsequent downstream tasks. In a service resource network, the task of node identification is to identify and classify key nodes with different functions by analyzing the characteristics of each node in the network and their connection relationships. This process usually requires comprehensive consideration of the topological attributes of nodes (such as degree, path length, etc.), functional attributes (such as bandwidth, latency, etc.), and the interaction relationships between nodes. Through node identification, researchers can better understand the structure of the network, optimize resource allocation, and enhance the security and robustness of the network.
[0004] Common related technologies include: (1) a node identification method based on feature engineering and classification algorithms; (2) a node identification method based on deep learning or graph neural networks. However, these methods are all difficult to accurately capture the characteristics of function nodes in the service resource network, resulting in a low identification accuracy. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method for identifying service network function nodes based on subgraph information bottleneck, which, in a data-driven form, combines subgraph multi-level node semantic information and multi-scale topological information from the perspective of information theory to obtain network node features with multi-scale constraint information in the latent space, can capture the more accurate characteristics of function nodes in the service resource network, and improve the accuracy of service network node identification.
[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0007] In a first aspect, the present invention provides a method for identifying service network function nodes based on subgraph information bottleneck. The method includes:
[0008] Obtain the topology graph of the service network to be measured, and obtain the feature matrix and adjacency matrix of the topology graph; wherein, the topology graph includes multiple function nodes;
[0009] Input the feature matrix and the adjacency matrix into a feature encoding model, and use the feature encoding model to obtain the node features of the service network to be measured according to the feature matrix and the adjacency matrix; wherein, the feature encoding model is a model obtained by training an optimization function of joint multi-order subgraph information bottleneck and a global topology constraint function;
[0010] Obtain the node recognition result of the service network to be measured according to the node features.
[0011] Optionally, the step of obtaining the feature encoding model includes:
[0012] Obtain the matrix information and real features of the sample service network, and the real connection relationship between each function node in the sample service network; wherein, the matrix information includes the adjacency matrix and the feature matrix;
[0013] Use the real features and the real connection relationship as the labels of the matrix information;
[0014] Use the matrix information as the input of the target network, and perform iterative training using the joint function to obtain the feature encoding model; wherein, the target network includes a feature encoding network for learning the features of each function node and an attention network for learning the connection relationship between function nodes; the joint function includes an optimization function of multi-order subgraph information bottleneck and a global topology constraint function, the global topology constraint function is to calculate the loss value between the real connection relationship and the output result of the attention network, and the optimization function is to calculate the loss value between the real features and the output result of the feature encoding network.
[0015] Optionally, the step of using the matrix information as the input of the target network and performing iterative training using the joint function to obtain the feature encoding model includes:
[0016] Input the matrix information into the feature encoding network and the attention network respectively, and obtain the first prediction result output by the target network and the second prediction result output by the attention network;
[0017] Based on the real connection relationship and the second prediction result, use the global topology constraint function to obtain the global topology loss value;
[0018] Based on the true features and the first prediction result, use the optimization function to obtain the information bottleneck optimization value;
[0019] Integrate the global topology loss value and the information bottleneck optimization value to obtain the joint loss value;
[0020] According to the joint loss value and backpropagation, update the parameters of the feature encoding network and the attention network;
[0021] When the joint loss value converges, use the current feature encoding network as the feature encoding model;
[0022] When the joint loss value does not converge, return to execute the step of inputting the matrix information into the feature encoding network and the attention network respectively to obtain the first prediction result output by the target network and the second prediction result output by the attention network.
[0023] Optionally, the optimization function includes a variational cross-entropy function and a KL divergence function, the true features include the true adjacency matrices of the first-order subgraphs and second-order subgraphs of each functional node, and the first prediction result includes the learned adjacency matrices of the first-order subgraphs and second-order subgraphs of each functional node;
[0024] The step of obtaining the information bottleneck optimization value by using the optimization function based on the true features and the first prediction result includes:
[0025] According to the learned adjacency matrix and the true adjacency matrix of the first-order subgraph of each functional node, use the KL divergence function to obtain the first-order distribution difference value;
[0026] According to the learned adjacency matrix and the true adjacency matrix of the second-order subgraph of each functional node, use the variational cross-entropy function to obtain the second-order distribution difference value;
[0027] Integrate the first-order distribution difference value and the second-order distribution difference value to obtain the information bottleneck optimization value.
[0028] Optionally, the optimization function includes: ;
[0029] Wherein, represents the total number of functional nodes of the sample service network, characterizes the learned adjacency matrix of the second-order subgraph of the th functional node, characterizes the true adjacency matrix of the second-order subgraph of the th functional node, represents the second-order distribution difference value, represents the hyperparameter of the feature encoding network, Characterize the learning adjacency matrix of the first-order subgraph representing the th functional node, true adjacency matrix of the first-order subgraph representing the th functional node,
[0030] Optionally, the second prediction result includes low-dimensional node representations of each functional node;
[0031] The step of obtaining the global topological loss value by using the global topological constraint function based on the true connection relationship and the second prediction result includes:
[0032] Obtaining the learning connection relationship between each functional node in the sample service network according to each low-dimensional node representation;
[0033] Obtaining a connection difference value according to the learning connection relationship and the true connection relationship;
[0034] Obtaining a connection distribution difference value based on each low-dimensional node representation and graph variational inference;
[0035] Obtaining the global topological loss value according to the connection difference value and the connection distribution difference value.
[0036] Optionally, the step of obtaining the connection distribution difference value based on each low-dimensional node representation and graph variational inference includes:
[0037] Calculating the mean and standard deviation of each low-dimensional node representation, and obtaining the variational distribution of each low-dimensional node representation according to the mean and standard deviation;
[0038] Obtaining a forward propagation result according to each low-dimensional node representation, the feature matrix, and a multi-layer perceptron;
[0039] Obtaining the connection distribution difference value according to the variational distribution, the forward propagation result, and the KL divergence function.
[0040] Optionally, the step of obtaining the learning connection relationship between each functional node in the sample service network according to each low-dimensional node representation includes:
[0041] Inputting any two low-dimensional node representations into an inlier function to obtain the learning connection relationship between the functional nodes corresponding to the two low-dimensional nodes.
[0042] Optionally, the global topological constraint function includes: ;
[0043] wherein, represents the functional node and the functional node The true connection relationship between , represents the th low-dimensional node representation, the in-point function of the representation, represents the functional node and the functional node the learning connection relationship between them, , represents the total number of low-dimensional node representations, the representation feature matrix, the forward propagation model of the representation, and respectively represent the mean and standard deviation of the low-dimensional node representations.
[0044] Optionally, the step of obtaining the node recognition result of the service network to be measured according to the node features includes:
[0045] Input the node features into a support vector machine to obtain the classification results of each functional node in the service network to be measured.
[0046] In a second aspect, the present invention provides a service network functional node recognition device based on subgraph information bottleneck, including a network modeling module, a feature extraction module, and a node recognition module;
[0047] The network modeling module is used to obtain the topological graph of the service network to be measured and obtain the feature matrix and adjacency matrix of the topological graph; wherein, the topological graph includes multiple functional nodes;
[0048] The feature extraction module is used to input the feature matrix and the adjacency matrix into a feature encoding model, and use the feature encoding model to obtain the node features of the service network to be measured according to the feature matrix and the adjacency matrix; wherein, the feature encoding model is a model obtained by training an optimization function of joint multi-order subgraph information bottleneck and a global topological constraint function;
[0049] The node recognition module is used to obtain the node recognition result of the service network to be measured according to the node features.
[0050] In a third aspect, the present invention provides an electronic device, including a processor and a memory, the memory stores a computer program that can be executed by the processor, and the processor can execute the computer program to implement the method for recognizing functional nodes of a service network based on subgraph information bottleneck as described in the first aspect.
[0051] Fourthly, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for identifying service network function nodes based on subgraph information bottleneck as described in the first aspect.
[0052] The method for identifying service network function nodes based on subgraph information bottleneck provided by the embodiments of the present invention includes: obtaining a topology graph of a service network to be measured, and obtaining a feature matrix and an adjacency matrix of the topology graph; wherein, the topology graph includes a plurality of function nodes; inputting the feature matrix and the adjacency matrix into a feature encoding model, and using the feature encoding model to obtain node features of the service network to be measured according to the feature matrix and the adjacency matrix, and the feature encoding model is a model obtained by training a joint optimization function of multi-order subgraph information bottleneck and a global topology constraint function; obtaining a node recognition result of the service network to be measured according to the node features. In this way, in a data-driven form, from the perspective of information theory, combining multi-level node semantic information and multi-scale topology information of subgraphs, network node features with multi-scale constraint information in the latent space are obtained, and more accurate function node features in the service network can be captured, thereby greatly improving the accuracy of service network node recognition.
[0053] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0055] Figure 1 Shows a schematic diagram of the system architecture of the service network function node recognition system provided by the embodiments of the present invention.
[0056] Figure 2 Shows a schematic diagram of the module architecture of the electronic device provided by the embodiments of the present invention.
[0057] Figure 3 Shows a schematic flowchart of the method for identifying service network function nodes provided by the embodiments of the present invention.
[0058] Figure 4 Shows a schematic flowchart of the method for identifying service network function nodes provided by the embodiments of the present invention.
[0059] Figure 5 Shows Figure 4Flow diagram of some sub-steps of step 25 in
[0060] Figure 6 shows Figure 5 Flow diagram of some sub-steps of step 252 in
[0061] Figure 7 shows Figure 5 Flow diagram of some sub-steps of step 253 in
[0062] Figure 8 shows Figure 7 Flow diagram of some sub-steps of step 253-5 in
[0063] Figure 9 shows the block diagram of the service network function node recognition device provided by the embodiment of the present invention.
[0064] Explanation of reference numerals: 10 - service network function node recognition system; 110 - recognition device; 120 - model training device; 20 - electronic device; 210 - memory; 220 - processor; 230 - communication module; 30 - service network function node recognition device; 310 - network modeling module; 320 - feature extraction module; 330 - node recognition module. Detailed implementation manners
[0065] In different fields, a service network is a type of data with obvious non-Euclidean characteristics. In addition to the semantic features of the network, the topological features of the network are also aspects that need to be considered key. However, in the actual process of service network function recognition, it is still rare to simultaneously consider both the attribute features and topological features of the network to jointly complete the service network node recognition task. Only relying on the single-attribute features of nodes to obtain node representations, or only relying on the topological structure statistical features to analyze the network, which still faces the problems of low computational efficiency in large-scale networks and difficulty in capturing node features in dynamic network environments.
[0066] To solve the above problems, the embodiment of the present invention provides a service network function node recognition method based on subgraph information bottleneck. In a data-driven form, combining subgraph multi-level node semantic information and multi-scale topological information from the perspective of information theory, obtaining network node features with multi-scale constraint information in the latent space, which can capture the features of more accurate functional nodes in the service resource network and improve the accuracy of service network node recognition.
[0067] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0068] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0069] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0070] The method for identifying service network function nodes based on subgraph information bottleneck provided by the embodiments of the present invention can be applied to Figure 1 the service network function node identification system 10 shown in the figure. The service network function node identification system 10 includes an identification device 110 and a model training device 120. The identification device 110 can be communicatively connected to the model training device 120 in a wired or wireless manner such as through a network, a data line, etc.
[0071] The model training device 120 is used to train and obtain a feature encoding model and deploy the feature encoding model to the identification device 110.
[0072] An identification device 110, which is used to implement the service network function node identification method based on subgraph information bottleneck provided by the embodiments of the present invention according to a feature encoding model, that is: obtain a topology graph of a service network to be measured, and obtain a feature matrix and an adjacency matrix of the topology graph; wherein, the topology graph includes multiple function nodes; input the feature matrix and the adjacency matrix into the feature encoding model, and use the feature encoding model to obtain node features of the service network to be measured according to the feature matrix and the adjacency matrix; wherein, the feature encoding model is a model obtained by training an optimization function combining multi-order subgraph information bottleneck and a global topology constraint function; obtain a node identification result of the service network to be measured according to the node features.
[0073] Among them, the identification device 110 can be, but is not limited to: a personal computer, a laptop computer, a tablet computer, an independent server, a server cluster, a mobile terminal, a wearable portable device, etc. The model training device 120 can be, but is not limited to: an independent server, a server cluster, a personal computer, etc.
[0074] Exemplarily, the identification device 110 and the model training device 120 can be the same device, that is, model training and node identification are both performed on the same device. The identification device 110 and the model training device 120 can also be independent devices.
[0075] Please refer to Figure 2 , which is a block diagram of an electronic device 20. The electronic device 20 can be Figure 1 the identification device 110 in the service network function node identification system 10 shown in the figure, or can also be the model training device 120. The electronic device 20 includes a memory 210, a processor 220, and a communication module 230. The elements of the memory 210, the processor 220, and the communication module 230 are directly or indirectly electrically connected to each other to realize data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines.
[0076] Among them, the memory 210 is used to store programs or data. The memory 210 can be, but is not limited to, a random access memory, a read-only memory, a programmable read-only memory, an erasable read-only memory, an electrically erasable read-only memory, etc.
[0077] The processor 220 is used to read / write data or programs stored in the memory 210 and execute corresponding functions. For example, Figure 1 in the service network function node identification system 10 shown in the figure, the processor 220 of the identification device 110 executes the computer program stored in the memory 210 to implement the service network function node identification method provided by the embodiments of the present invention, and the processor 220 of the model training device 120 executes the computer program stored in the memory 210 to obtain the feature encoding model.
[0078] The communication module 230 is used to establish a communication connection between the electronic device 20 and other communication terminals through a network, and is used to send and receive data through the network. For example, Figure 1 In the service network function node recognition system 10 shown, the communication module 230 of the recognition device 110 or the model training device 120 is used to send and receive data through the network.
[0079] It should be understood that Figure 2 The structure shown is only a schematic diagram of the structure of the electronic device 20, and the electronic device 20 may also include more or fewer components than those shown in Figure 2 or have a different configuration from that shown in Figure 2 . Figure 2 Each component shown in can be implemented by hardware, software, or a combination thereof.
[0080] Referring to Figure 3 , a service network function node recognition method provided by an embodiment of the present invention includes steps 11 to 15. Thus, it is made that: Figure 1 In the service network function node recognition system 10 shown, the recognition device 110 Figure 2 with the structure shown, when the processor 220 executes the computer program stored in the memory, steps 11 to 15 are implemented.
[0081] Step 11, obtain the topology graph of the service network to be measured, and obtain the feature matrix and adjacency matrix of the topology graph.
[0082] Among them, the topology graph includes multiple function nodes, that is, it is composed of function nodes and the connection relationships between function nodes. A function node is an integrated and multi-functional service node in the service network. For example, it can be a client, a web server, an application server, a data server, a switch, a router, a firewall, a DNS server, etc.
[0083] Step 13, input the feature matrix and the adjacency matrix into the feature encoding model, and use the feature encoding model to obtain the node features of the service network to be measured according to the feature matrix and the adjacency matrix.
[0084] Among them, the feature encoding model is a model obtained by training the optimization function of the joint multi-order subgraph information bottleneck and the global topology constraint function. That is, in the process of operation and processing, the feature encoding model can combine the subgraph multi-level node semantic information and multi-scale topology information from the perspective of information theory to obtain the function node features with multi-scale constraint information in the hidden space.
[0085] Step 15, obtain the node recognition result of the service network to be measured according to the node features.
[0086] The feature matrix includes the preliminary features of each functional node in the topology graph (which can be represented as ), such as functions (such as computing, storage, retrieval, etc.), topological structure, connection distance, operating status, service resources, IP addresses, etc. The adjacency matrix includes the adjacency relationships of each functional node, that is, an n×n matrix, where each element represents the situation of the edge between the functional node and the functional node . For example, if there is an edge between the functional node and the functional node , then the value of is 1; if there is no edge between the functional node and the functional node , then the value of is 0.
[0087] Exemplarily, in combination with the service network function node recognition system 10 shown in Figure 1 , after the model training device 120 trains the target network (such as a graph convolutional neural network) by combining the optimization function of the multi-order subgraph information bottleneck and the global topological constraint function, a feature encoding model is obtained, and the feature encoding model is downloaded and deployed to the recognition device 110. When it is necessary to recognize the service network to be measured, the recognition device 110 models the service network to be measured to obtain a topology graph, and obtains the preliminary feature matrix and adjacency matrix of the topology graph by using a preset feature extraction rule or model. Thus, the recognition device 110 inputs the feature matrix and the adjacency matrix into the feature encoding model, and uses the feature encoding model to combine the subgraph multi-level node semantic information (i.e., the feature matrix) and the multi-scale topological information (i.e., the adjacency matrix) from the perspective of information theory according to the feature matrix and the adjacency matrix, so as to obtain the node features of the service network to be measured.
[0088] At this time, the node features include the node features of each functional node in the service network to be measured. Furthermore, the recognition device 110 obtains the node recognition result of the service network to be measured, that is, the recognition result of each functional node.
[0089] Compared with the method in the related technology that only relies on the single-attribute features of nodes to obtain node representations or only depends on the topological structure statistical features to analyze the network, in the above service network function node recognition method based on the subgraph information bottleneck, in a data-driven form, from the perspective of information theory, it combines the subgraph multi-level node semantic information and the multi-scale topological information to obtain the network node features with multi-scale constraint information in the hidden space, which can capture the more accurate functional node features in the service network, thus greatly improving the accuracy of service network node recognition.
[0090] Among them, the method for training the feature encoding model can be flexibly selected. For example, a joint function composed of an optimization function of multi-order subgraph information bottleneck and a global topology constraint function can be used to iteratively train the graph convolutional neural network, or it can be trained according to preset rules. Moreover, the above methods are only examples, and their implementation methods are not limited.
[0091] To enable the feature encoding model to learn: from the perspective of information theory, combining subgraph multi-level node semantic information and multi-scale topology information to obtain functional node features with multi-scale constraint information in the hidden space, a joint function is introduced during the model training process, and attention network is used for auxiliary reinforcement learning to train the feature encoding model. Refer to Figure 4 , the process of obtaining the feature encoding model includes steps 21 to 25. Thus, it is ensured that: Figure 1 In the service network function node recognition system 10 shown in Figure 2 , when the model training device 120 has the structure shown in
[0092] Step 21, obtain the matrix information and true features of the sample service network, and the true connection relationships between the functional nodes in the sample service network.
[0093] Among them, the matrix information includes the adjacency matrix and feature matrix of the topology graph of the sample service network. The topology graph includes each functional node in the sample service network and the connection relationships between the functional nodes in at least one application scenario. The true connection relationships include the connection relationships shown in the topology graph and the implicit connection relationships not shown in the topology graph, that is, the connection relationships between the functional nodes in various application scenarios (such as storage scenarios, instant messaging, remote login services, data query services, content distribution services, etc.). The true features include the features of each functional node in different scenarios, for example, functions (such as computing, storage, retrieval, etc.), topology structures, connection distances, operating states, service resources, IP addresses, etc.
[0094] Step 23, use the true features and true connection relationships as the labels of the matrix information.
[0095] Step 25, use the matrix information as the input of the target network, and use the joint function for iterative training to obtain the feature encoding model.
[0096] The target network includes a feature encoding network for learning the features of each functional node and an attention network for learning the connection relationships between the functional nodes. Among them, the feature encoding network can be any graph convolutional neural network, and the attention network can be any graph attention network.
[0097] The joint function includes the optimization function of the multi-order subgraph information bottleneck and the global topology constraint function. Among them, the global topology constraint function is to calculate the loss value between the real connection relationship and the output result of the attention network. The output result of the attention network includes the learned connection relationship (that is, the connection relationship between each functional node in the learned sample service network). The optimization function is to calculate the loss value between the real feature and the output result of the feature encoding network. The output result of the feature encoding network includes the features of each functional node in the learned sample service network in each application scenario.
[0098] During the iterative training process, the gradient of the model parameters is calculated based on the loss value between the real features and the output results of the feature encoding network, as well as the loss value between the real connection relationship and the output results of the attention network. Then, the model parameters of the feature encoding network (i.e., graph convolutional neural network) and the attention network are continuously adjusted according to the gradient and learning rate, such as the weights of each layer, optimizer parameters, regularization parameters, dropout rate, etc.
[0099] In the above steps 21 to 25, the attention network is introduced to assist the feature encoding network in enhanced learning, and the feature encoding network is continuously optimized according to the loss value between the real feature and the output result of the feature encoding network, and the loss value between the real connection relationship and the output result of the attention network, so that the feature encoding network (i.e., the feature encoding model) can quickly learn: from the perspective of information theory, combining the multi-level node semantic information and multi-scale topological information of the subgraph, to obtain the functional node features with multi-scale constraint information in the latent space. Therefore, the feature encoding model can mine the node features in the latent space, making the node features more accurate.
[0100] The iterative training method in step 25 can be flexibly set, for example, the number of iterations can be set, and the training is completed when the number of iterations reaches the set value, or the training is terminated when the joint function converges. The above method is only an example, and its implementation method is not limited.
[0101] In order to ensure that the feature encoding model with the required performance (such as accuracy) is obtained, the training is ended when the joint function converges in the iterative training of step 25 to obtain the concept of the feature encoding model. Figure 5 , the iterative training process in step 25 includes steps 251 to 257.
[0102] Step 251, input the matrix information into the feature encoding network and the attention network respectively to obtain the first prediction result output by the target network and the second prediction result output by the attention network.
[0103] Step 252: Based on the real features and the first prediction result, an optimization function is used to obtain an information bottleneck optimization value.
[0104] Step 253: Based on the true connection relationship and the second prediction result, use the global topology constraint function to obtain the global topology loss value.
[0105] Step 254: Synthesize the global topology loss value and the information bottleneck optimization value to obtain the combined loss value.
[0106] Step 255: Update the parameters of the feature encoding network and the attention network according to the combined loss value and backpropagation.
[0107] Step 256: Determine whether the combined loss value converges. If it does, execute Step 257; if not, return to execute Step 251.
[0108] Step 257: Take the current feature encoding network as the feature encoding model.
[0109] To enable the final feature encoding model to extract the node features in the latent space. Therefore, during initialization, the component needs to initialize the network parameters of the feature encoding model and initialize the first-order subgraph and multi-order subgraphs of each functional node according to the adjacency matrix in the matrix information. For example, it should at least include the second-order subgraph.
[0110] The first-order subgraph of a functional node includes the node itself (i.e., the vertex), the adjacent nodes of the vertex, and the edges between the vertex and the adjacent nodes. The second-order subgraph of a functional node includes the node itself (i.e., the vertex), the adjacent nodes of the vertex, the adjacent nodes of the adjacent nodes, and the edges between these nodes. After determining a functional node, the first-order subgraph and multi-order subgraphs of this functional stage can be obtained from the adjacency matrix of the matrix information.
[0111] Input the matrix information into the feature encoding network and the attention network respectively. The target network outputs the first prediction result after processing, and the attention network outputs the second prediction result. The first prediction result includes the learned adjacency matrices of the first-order subgraph and the second-order subgraph of each functional node, and the second prediction result includes the low-dimensional node representations of each functional node.
[0112] The setting of the optimization function for the multi-order subgraph information bottleneck can be flexibly selected. For example, it can be any loss function that can learn the difference between the learned and the true features and minimize the difference, such as the KL divergence, cross-entropy loss function, etc. Its implementation method is not restricted.
[0113] To enable the feature encoding network to learn the minimum sufficient features of each functional node by maximizing the mutual information between the node features and the target, while constraining the mutual information between the node features and the matrix information. For a given functional node and its subgraph , define the multi-order subgraph information bottleneck optimization problem as: , where and Respectively characterize the learned adjacency matrix and subgraph learned by the feature encoding network and the true adjacency matrix of the subgraph The predicted connection relationship between functional nodes in is directly described as the learned self-supervised label. In addition The elements of are expressed as where is the low-dimensional node representation of the functional node represents the Sigmoid function
[0114] In order to enable the feature encoding network to effectively learn the node topology information at multiple scales, in the subgraph information bottleneck optimization problem, the embodiments of the present invention introduce a multi-order subgraph to effectively learn the node topology information at multiple scales. And taking the functional node as the central node, the second-order subgraph is defined as where the node set of the second-order subgraph is defined as and the edge set of the second-order subgraph is defined as: represents the edge set of the topological graph represents the node set of the topological graph
[0115] Furthermore, assuming that the first-order subgraph and the second-order subgraph of the given functional node are given, the multi-order subgraph information bottleneck optimization problem is reformulated as: represents the mutual information between the second-order subgraph of the th functional node and the learned adjacency matrix represents the mutual information between the first-order subgraph of the th functional node and the learned adjacency matrix represents the hyperparameter of the feature encoding network. The first-order subgraph has the same definition as a general subgraph and respectively represent the learned adjacency matrices learned by the feature encoding network from and respectively
[0116] Since the mutual information between two variables in a high-dimensional space cannot be directly calculated, variational inference is used to approximate the mutual information. Therefore, for is the true posterior The variational approximation is such that: , represents entropy.
[0117] From this, it can be deduced that The lower bound of is: , characterizes the topological graph.
[0118] Let be the variational approximation representing the true prior distribution , then the upper bound of the second term in the expression of the multi-section subgraph information bottleneck optimization problem can be calculated as: .
[0119] Combining the lower bound of the first term and the upper bound of the second term, the overall lower bound of the multi-order topological information bottleneck optimization problem can be obtained, and this expression is represented as: .
[0120] For the above, the first term can be obtained by calculating the variational cross-entropy function, and the second term is quantified by the Kullback-Leibler divergence (i.e., KL divergence) between the true posterior distribution and the prior distribution. Therefore, the loss function of the multi-order topological information bottleneck optimization problem, that is, the optimization function of the multi-order subgraph information bottleneck, includes the variational cross-entropy function and the KL divergence function. At this time, referring to Figure 6 , the process of obtaining the information bottleneck optimization value in step 252 can include steps 252-1 to step 252-5.
[0121] Step 252-1, according to the learned adjacency matrix and the true adjacency matrix of the first-order subgraphs of each functional node, use the KL divergence function to obtain the first-order distribution difference value.
[0122] Step 252-3, according to the learned adjacency matrix and the true adjacency matrix of the second-order subgraphs of each functional node, use the variational cross-entropy function to obtain the second-order distribution difference value.
[0123] Step 252-5, synthesize the first-order distribution difference value and the second-order distribution difference value to obtain the information bottleneck optimization value.
[0124] At this time, the process of obtaining the information bottleneck optimization value in the above steps 252-1 to 252-5 can be expressed by a formula, that is, the optimization function includes: .
[0125] Among them, represents the total number of functional nodes of the sample service network, characterizes the th learned adjacency matrix of the second-order subgraph of the th functional node, characterizes the true adjacency matrix of the second-order subgraph of the Represents the second-order distribution difference value, Represents the hyperparameters of the feature encoding network, Characterizes the Learning adjacency matrix of the first-order subgraph of the Characterizes the True adjacency matrix of the first-order subgraph of the Represents the first-order distribution difference value.
[0126] Meanwhile, the setting of the global topological constraint function can be flexibly selected. For example, it can be the mean squared error loss function, KL divergence, cross-entropy loss function, or any other function that can measure the difference between the second prediction result of the attention network output and the true connection relationship, and its implementation method is not restricted.
[0127] When the second prediction result includes the low-dimensional node representations of each functional node, that is, the attention network outputs the low-dimensional node representations of each functional node, a global topological constraint function is constructed by combining graph variational inference and a function that measures the difference between the learned connection relationship and the true connection relationship.
[0128] Thus, by minimizing the difference between the approximate distribution of the connection relationship and the true posterior distribution through graph variational inference, and measuring the difference between the learned connection relationship and the true connection relationship, the information capture of the global topology and local topology in the subgraph information bottleneck learning process is balanced. On this basis, referring to Figure 7 In step 253, the process of obtaining the global topological loss value includes steps 253-1 to 253-7.
[0129] Step 253-1: Based on each low-dimensional node representation, obtain the learned connection relationship between each functional node in the sample service network.
[0130] Step 253-3: According to the learned connection relationship and the true connection relationship, obtain the connection difference value.
[0131] Step 253-5: Based on each low-dimensional node representation and graph variational inference, obtain the connection distribution difference value.
[0132] Step 253-7: According to the connection difference value and the connection distribution difference value, obtain the global topological loss value.
[0133] Exemplarily, in step 253-1, input any two low-dimensional node representations into the inlier function or MLP model, and obtain the connection relationship between the functional nodes corresponding to these two low-dimensional nodes output by the inlier function or MLP model.
[0134] Among them, the learned connection relationship between two functional nodes can be expressed as: , Characterizes the The low-dimensional node representation of a functional node, which can be an inner-point function or an MLP model, represents the low-dimensional node representation of the corresponding functional node and the learning connection relationship between the low-dimensional node representation and the corresponding functional node.
[0135] Furthermore, in step 253-3, any difference measurement function can be used to calculate the connection difference value between the learning connection relationship and the true connection relationship. For example, it can be the L1 norm, the L2 norm, or any other loss function.
[0136] In the case of using the L2 norm, the learning connection relationship and the true connection relationship between any two functional nodes can be expressed as: , where represents the true connection relationship between functional node and functional node represents the learning connection relationship between functional node and functional node.
[0137] Meanwhile, referring to Figure 8 , in step 253-5, the connection distribution difference value is obtained through the following steps 25-51 to 25-55, that is, by minimizing the difference between the approximate distribution of the connection relationship and the true posterior distribution through graph variational inference.
[0138] Step 25-51, calculate the mean and standard deviation of each low-dimensional node representation, and obtain the variational distribution of each low-dimensional node representation according to the mean and standard deviation.
[0139] Step 25-53, obtain the forward propagation result according to each low-dimensional node representation, the feature matrix, and the multi-layer perception.
[0140] Step 25-55, obtain the connection distribution difference value according to the variational distribution, the forward propagation result, and the KL divergence function.
[0141] The process of the above steps 25-51 to 25-55 can be expressed by the formula: , where , represents the th low-dimensional node representation, represents the feature matrix, represents the forward propagation model, and respectively represent the mean and standard deviation of the low-dimensional node representation.
[0142] Therefore, the calculation formula of the global topology loss value, that is, the global topology constraint function, can be expressed as: .
[0143] Furthermore, the joint function can be expressed as: .
[0144] Through the above steps 21 to 25, that is, the corresponding sub-steps, a feature encoding model is trained to effectively learn the node topology information at multiple scales and balance the information capture of the global topology and the local topology in the subgraph information bottleneck learning process, so that the feature encoding model can combine the subgraph multi-level node semantic information (i.e., the feature matrix) and the multi-scale topology information (i.e., the adjacency matrix) from the perspective of information theory to perform feature mining and obtain network node features with multi-scale constraint information in the latent space, so as to obtain more accurate node features (node features of a single functional node) of each functional node in the service network to be measured.
[0145] After obtaining the node features of each functional node in the service network to be measured by using the trained feature encoding model, in step 15, the way to obtain the node recognition result can be flexibly selected. For example, classification can be performed according to a preset rule, or a pre-trained classification model can be used for classification to obtain the classification results of each functional node. And the above methods are only examples, and their implementation methods are not limited.
[0146] Exemplarily, the node features are input into a support vector machine to obtain the classification results of each functional node in the service network to be measured. That is, after training a graph convolutional neural network following the gradient descent method and completing the end-to-end network training to obtain the feature encoding model. Using this feature encoding model, the final node features of the service network functional nodes are obtained, and a support vector machine is used to classify and identify the final node features of the service network functional nodes.
[0147] Based on the same concept as the above service network functional node recognition method based on subgraph information bottleneck, referring to Figure 9 , the embodiment of the present invention further provides a service network functional node recognition device 30, including a network modeling module 310, a feature extraction module 320, and a node recognition module 330.
[0148] The network modeling module 310 is configured to obtain the topology graph of the service network to be measured and obtain the feature matrix and the adjacency matrix of the topology graph. Among them, the topology graph includes multiple functional nodes.
[0149] The feature extraction module 320 is configured to input the feature matrix and the adjacency matrix into a feature encoding model, and use the feature encoding model to obtain the node features of the service network to be measured according to the feature matrix and the adjacency matrix. Among them, the feature encoding model is a model obtained by training an optimization function for jointly multi-order subgraph information bottleneck and a global topology constraint function.
[0150] The node recognition module 330 is configured to obtain the node recognition result of the service network to be measured according to the node features.
[0151] The above-mentioned network modeling module 310, feature extraction module 320, and node recognition module 330 can be deployed in Figure 1 the recognition system of the service network function node recognition system 10 as shown.
[0152] Exemplarily, it may further include a sample processing module and a model training module to obtain the feature encoding model.
[0153] The sample processing module is configured to: obtain the matrix information and real features of the sample service network, and the real connection relationship between each functional node in the sample service network; use the real features and the real connection relationship as the labels of the matrix information. Among them, the matrix information includes the adjacency matrix and the feature matrix.
[0154] The model training module is configured to use the matrix information as the input of the target network, and perform iterative training using the joint function to obtain the feature encoding model.
[0155] Among them, the target network includes a feature encoding network for learning the features of each functional node and an attention network for learning the connection relationship between functional nodes. The joint function includes an optimization function for multi-order subgraph information bottleneck and a global topology constraint function. The global topology constraint function is to calculate the loss value between the real connection relationship and the output result of the attention network, and the optimization function is to calculate the loss value between the real features and the output result of the feature encoding network.
[0156] Under the action of the network modeling module 310, feature extraction module 320, and node recognition module 330 of the above-mentioned service network function node recognition device 30, in a data-driven form, from the perspective of information theory, combining subgraph multi-level node semantic information and multi-scale topology information, the network node features with multi-scale constraint information in the latent space are obtained, and more accurate functional node features in the service network can be captured, thereby greatly improving the accuracy of service network node recognition.
[0157] For the specific implementation and effects of the service network function node recognition device 30, reference can be made to the description of the implementation of the service network function node recognition method in the above text. For example, for the specific implementation and effects of the network modeling module 310, reference can be made to the description of the relevant content in step 11 above; for the specific implementation and effects of the feature extraction module 320, reference can be made to the description of the relevant content in step 13 above; for the specific implementation and effects of the node recognition module 330, reference can be made to the description of the relevant content in step 15 above, which will not be elaborated here.
[0158] In addition, each module of the above service network function node recognition device 30 can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor 220 of the electronic device 20 in hardware form or be independent of it, or can be stored in the memory 210 of the electronic device 20 in software form, so as to facilitate the processor 220 to call and execute the operations corresponding to each of the above modules to implement the service network function node recognition method provided in the above text.
[0159] An embodiment of the present invention also provides an electronic device 20, including a processor 220 and a memory 210. The memory 210 stores a computer program that can be executed by the processor 220, and the processor 220 can execute the computer program to implement the service network function node recognition method provided above.
[0160] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor 220, it implements the service network function node recognition method proposed in the embodiment of the present invention.
[0161] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and a module, a program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0162] In addition, in each embodiment of the present invention, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0163] If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory 210 (ROM, Read-Only Memory), random access memory 210 (RAM, Random Access Memory), magnetic disks, or optical discs and other various media that can store program codes.
[0164] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for identifying service network function nodes based on subgraph information bottleneck, characterized in that: The method comprises: Obtain a topology map of the service network to be tested, and obtain a feature matrix and an adjacency matrix of the topology map; wherein the topology map includes a plurality of functional nodes, and a functional node is an integrated, multifunctional service node in the service network, including a client, a web server, an application server, a data server, a switch, a router, a firewall, and a DNS server; Inputting the feature matrix and the adjacency matrix into a feature coding model, and using the feature coding model to obtain node features of the service network to be tested according to the feature matrix and the adjacency matrix; wherein the feature coding model is a model obtained by training the optimization function of the joint multi-order subgraph information bottleneck and the global topology constraint function; Obtaining a node identification result of the service network to be tested according to the node characteristics; The step of obtaining the feature encoding model comprises: Obtaining matrix information and real features of the sample service network, as well as real connection relationships between functional nodes in the sample service network; wherein the matrix information includes an adjacency matrix and a feature matrix; Using the real features and the real connection relationship as labels of the matrix information; The matrix information is used as the input of the target network, and iterative training is performed using a joint function to obtain a feature encoding model; wherein the target network includes a feature encoding network for learning the features of each functional node and an attention network for learning the connection relationship between functional nodes; the joint function includes an optimization function of a multi-order subgraph information bottleneck and a global topology constraint function, the global topology constraint function is used to calculate the loss value between the real connection relationship and the output result of the attention network, and the optimization function is used to calculate the loss value between the real feature and the output result of the feature encoding network.
2. The method for identifying service network function nodes based on subgraph information bottleneck according to claim 1, characterized in that: The step of using the matrix information as the input of the target network and performing iterative training using a joint function to obtain a feature encoding model includes: Inputting the matrix information into the feature encoding network and the attention network respectively, obtaining a first prediction result output by the target network and a second prediction result output by the attention network; Based on the real connection relationship and the second prediction result, the global topology constraint function is used to obtain a global topology loss value; Based on the real feature and the first prediction result, using the optimization function to obtain an information bottleneck optimization value; The global topology loss value and the information bottleneck optimization value are combined to obtain a joint loss value; According to the joint loss value and back propagation, updating the parameters of the feature encoding network and the attention network; When the joint loss value converges, using the current feature encoding network as a feature encoding model; When the joint loss value does not converge, return to the step of inputting the matrix information into the feature encoding network and the attention network respectively to obtain a first prediction result output by the target network and a second prediction result output by the attention network.
3. The method for identifying service network function nodes based on subgraph information bottleneck according to claim 2 is characterized in that: The optimization function includes a variable cross entropy function and a KL divergence function, the real feature includes a real adjacency matrix of a first-order subgraph and a second-order subgraph of each functional node, and the first prediction result includes a learned adjacency matrix of a first-order subgraph and a second-order subgraph of each functional node; The step of obtaining the information bottleneck optimization value by using the optimization function based on the real feature and the first prediction result comprises: According to the learned adjacency matrix and the real adjacency matrix of the first-order subgraph of each functional node, the KL divergence function is used to obtain a first-order distribution difference value; According to the learned adjacency matrix and the real adjacency matrix of the second-order subgraph of each functional node, the variable cross entropy function is used to obtain a second-order distribution difference value; The information bottleneck optimization value is obtained by combining the first-order distribution difference value and the second-order distribution difference value.
4. The method for identifying service network function nodes based on subgraph information bottleneck according to claim 3 is characterized in that: The optimization function includes: ; in, represents the total number of functional nodes of the sample service network, Characterization The learned adjacency matrix of the second-order subgraph of function nodes, Characterization The true adjacency matrix of the second-order subgraph of function nodes, represents the difference value of the second-order distribution, represents the hyperparameters of the feature encoding network, Characterization The learned adjacency matrix of the first-order subgraph of function nodes, Characterization The true adjacency matrix of the first-order subgraph of function nodes, Represents the first-order distribution difference value.
5. The method for identifying service network function nodes based on subgraph information bottleneck according to claim 2, characterized in that: The second prediction result includes a low-dimensional node representation of each functional node; The step of obtaining a global topology loss value based on the real connection relationship and the second prediction result by using the global topology constraint function comprises: Obtaining a learning connection relationship between each functional node in the sample service network according to each low-dimensional node representation; Obtaining a connection difference value according to the learned connection relationship and the real connection relationship; Based on the low-dimensional node representations and graph variational inference, obtaining a connection distribution difference value; A global topology loss value is obtained according to the connection difference value and the connection distribution difference value.
6. The method for identifying service network function nodes based on subgraph information bottleneck according to claim 5, characterized in that: The step of obtaining the connection distribution difference value based on each of the low-dimensional node representations and graph variational inference includes: Calculating the mean and standard deviation of each of the low-dimensional node representations, and obtaining the variational distribution of each of the low-dimensional node representations according to the mean and standard deviation; Obtaining forward propagation results according to the low-dimensional node representations, the feature matrix and the multi-layer perception; A connection distribution difference value is obtained according to the variational distribution, the forward propagation result and the KL divergence function.
7. The method for identifying service network function nodes based on subgraph information bottleneck according to claim 5, characterized in that: The step of obtaining the learning connection relationship between each functional node in the sample service network according to each low-dimensional node representation includes: Any two of the low-dimensional node representations are input into the interior point function to obtain the learning connection relationship between the functional nodes corresponding to the two low-dimensional nodes.
8. The method for identifying service network function nodes based on subgraph information bottleneck according to claim 6 or 7, characterized in that: The global topology constraint function includes: ; in, Represents a function node With function node The real connection between , Indicates A low-dimensional node representation, Characterize the interior point function, Represents a function node With function node The learning connection between , represents the total number of low-dimensional node representations, Characterize the feature matrix, Characterize the forward propagation model, and Represent the mean and standard deviation of the low-dimensional node representation respectively.
9. The method for identifying service network function nodes based on subgraph information bottleneck according to claim 1, characterized in that: The step of obtaining the node identification result of the service network to be tested according to the node characteristics includes: The node features are input into a support vector machine to obtain classification results of each functional node in the service network to be tested.
Citation Information
Patent Citations
Community structure identification method and device based on network embedding
CN111931023A
Training method and training device for node attribute information detection model in transaction platform
CN115965071A