A Clustering Aggregation Method for the Core-Periphery Stochastic Block Model in a Literature Citation Network
Through the Bayesian optimization-driven clustering and aggregation method of literature citation network core-peripheral random block model, the inconsistency and accuracy of literature citation network core-peripheral structure analysis in the existing technology is solved, and more accurate core papers and peripheral paper identification and academic influence analysis are achieved.
Patent Information
- Application Number
- CN202510440333.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The prior art has problems of inconsistency and accuracy when analyzing the core-peripheral structure of the literature citation network, making it difficult to accurately identify core papers and peripheral papers, affecting academic influence analysis and scientific research dynamic prediction.
A Bayesian optimization-driven literature citation network core-peripheral random block model cluster aggregation method is proposed. By constructing the core-peripheral random block model, Bayesian priorsum model minimum description length (MDL) optimization is introduced, which improves the accuracy and robustness of structure recognition, and forms a more representative and stable consensus structure through clustering aggregation scheme.
It significantly improves the accuracy and robustness of the core-peripheral structure analysis of the literature citation network, can more accurately identify core papers and peripheral papers, reveal the dissemination paths and influence distribution of academic knowledge, and provide more accurate theoretical support for academic influence analysis.
Smart Images

Figure CN119961617B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network science, especially to the technical field of literature citation network analysis and graph model research. A clustering aggregation method for the core-periphery stochastic block model of a literature citation network is proposed, specifically a core-periphery stochastic block model clustering aggregation method of a literature citation network driven by Bayesian optimization, which is used for literature influence analysis and scientific research trend prediction. Background Art
[0002] In the process of the continuous deepening of academic research and knowledge accumulation, the literature citation network, as an important carrier of knowledge dissemination and academic communication, has gradually received extensive attention. The literature citation network is composed of academic papers and their mutual citation relationships, reflecting the inheritance of academic thoughts, the cross-integration of disciplinary fields, and the evolution of research hotspots. The literature citation network is similar to a social network, where its nodes represent specific academic literatures, and the edges represent the citation relationships between literatures. Identifying influential literatures is of great significance for understanding and mastering the literature citation network.
[0003] The existing research still has limitations in understanding the structural characteristics and knowledge flow patterns of the literature citation network, especially in the aspect of how to accurately identify and analyze the core-periphery structure in the network. Although traditional network analysis methods such as centrality analysis and clustering algorithms can provide certain insights, when dealing with large-scale, complex and changeable literature citation networks, it is often difficult to effectively capture the global characteristics and local differences of the network. In addition, the core-periphery structures obtained by different methods may be inconsistent, making the comprehensive analysis and understanding of the network structure face challenges. Therefore, how to propose a method that can accurately depict the core-periphery structure of the literature citation network and perform effective structure aggregation and analysis based on this has become an important issue in current academic research.
[0004] The application of existing core-periphery structure analysis methods in academic citation networks has the following problems: 1) There are significant differences in the core-periphery structures obtained by different methods, lacking a unified evaluation standard; 2) It is difficult for traditional methods to introduce prior knowledge into the model, thus affecting the accuracy of the results; 3) The clustering aggregation scheme of existing methods does not fully consider the weights of local features. If directly applied to network analysis, it may lead to insufficient network representation ability. Therefore, in order to solve the above problems, a clustering aggregation method for the core-periphery stochastic block model of a literature citation network needs to be proposed. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a clustering aggregation method for the core-periphery stochastic block model of a literature citation network, specifically a clustering aggregation method for the core-periphery stochastic block model of a literature citation network driven by Bayesian optimization, which can be used for literature influence analysis and scientific research trend prediction, and solves the problems of inconsistency and accuracy existing in the analysis of the core-periphery structure of a literature citation network.
[0006] This method constructs a core-periphery stochastic block model to divide the nodes (papers) in the literature citation network into core nodes and peripheral nodes, so as to reveal the dissemination path and influence distribution of academic knowledge. By introducing Bayesian priors and combining network data and prior knowledge, the accuracy of core-periphery structure recognition is improved. At the same time, the minimum description length (MDL) of the model is used as the optimization objective to avoid model overfitting and enhance the robustness and generalization ability of the model. In addition, using the idea of clustering aggregation, a core-periphery structure aggregation scheme is designed. By calculating the uncertainty and reliability of the node set, the local aggregation weight is determined, and multiple core-periphery structures are aggregated into a comprehensive consensus core-periphery structure, which more accurately reflects the overall characteristics of the academic citation network and the influence of core papers. This method significantly improves the accuracy and robustness of the core-periphery structure analysis of the literature citation network, and provides more accurate theoretical support for academic influence analysis.
[0007] To achieve the above object, the present invention adopts the following technical solutions: A clustering aggregation method for the core-periphery stochastic block model of a literature citation network, comprising the following steps:
[0008] S1. Define the literature citation network; give the definition G(V, E) of the literature citation network, where G(V, E) includes a set V composed of N nodes and a set E composed of M unweighted directed edges. Each unweighted directed edge represents the specific citation record of a literature.
[0009] S2. Construct the core-periphery stochastic block model of the literature citation network;
[0010] S3. Construct a spoke model and a hierarchical model;
[0011] S4. Use Bayesian inference to infer the core-periphery stochastic block model of the literature citation network;
[0012] S5: Based on the principle of the minimum description length of the model, optimize and evaluate the core-periphery stochastic block model of the literature citation network;
[0013] S6. Aggregate the structure of the core-periphery stochastic block model of the literature citation network;
[0014] S7. Obtain the influential literature, that is, the core nodes.
[0015] Further, the construction of the core-periphery stochastic block model of the literature citation network in step S2 includes the following steps:
[0016] S21: Give the model assumptions of the core-periphery stochastic block model of the literature citation network, that is, the core nodes and the periphery nodes are assigned to different blocks. Based on the defined G(V, E), a literature citation network with nodes is given, and the corresponding adjacency matrix is , nodes are randomly assigned to M blocks;
[0017] S22: Represent the block assignment of nodes in the literature citation network through a block assignment vector D with a length of , specifically expressed as , where the node being assigned to block s is correspondingly expressed as ; Represent the connection probability between any two nodes in G(V, E) with the matrix , represents the probability that a node in block s is connected to a node in block t; in the core-periphery stochastic block model of the literature citation network, the connection probability between two nodes depends on the block assignment matrix D and the block connection matrix ;
[0018] S23: Based on the Bayesian method, relate to the prior probabilities of the block assignment vector D and the block connection matrix to obtain Formula 1: ; where the symbol represents proportionality, represents the joint probability distribution of the block assignment vector and the block connection matrix under the condition of observing the adjacency matrix , represents the prior distribution of the block assignment vector, represents the prior distribution of the block connection matrix, represents the probability of generating the adjacency matrix and given .
[0019] Further, the spoke model in S3 is obtained by combining the two-block model with the core-periphery stochastic block model of the literature citation network; the two-block model divides the network nodes into a core block and a periphery block, and defines the structural model of the connection probability between nodes through the block connection matrix ; in the spoke model, the network nodes are divided into two node sets, the core nodes and the periphery nodes. The core nodes are connected to each other, and some of the core nodes are connected to the periphery nodes, and the periphery nodes are not connected to each other;
[0020] Encode the core block of the spoke model as , and the peripheral block as ; The specific definition is and ( ), where represents the connection probability of the internal nodes of the core block. In the spoke model, set , that is, the core nodes are fully connected, represents the connection probability between the core block and the peripheral block, represents the connection probability of the internal nodes of the peripheral block;
[0021] The prior constraints of the block connection matrix R of the spoke model are represented by Equation 2 and Equation 3: Equation 2 is: ; Equation 3 is ; where is an indicator function used to define the legality constraint conditions of the block connection matrix in the spoke model, represents the prior probability distribution of the block connection matrix .
[0022] Furthermore, the hierarchical model in S3 is obtained by combining the k-core model with the core-periphery stochastic block model of the literature citation network; in the hierarchical model, the network nodes are stratified based on the k-core decomposition method, where the k-core represents the set of all nodes in the network that are connected to at least k other nodes;
[0023] The prior constraints of the block connection matrix R of the hierarchical model are represented by Equation 4 and Equation 5; Equation 4 is , and Equation 5 is ; where is an indicator function used to ensure that the block connection matrix of the hierarchical model conforms to its structural assumptions, represents the prior probability distribution of the block connection matrix .
[0024] Furthermore, Bayesian inference is used in S4 to infer the core-periphery stochastic block model of the literature citation network, which specifically includes the following sub-steps:
[0025] S41: Based on the idea of Gibbs sampling, alternately sample the block assignment vector D and the block connection matrix , that is, first fix the block assignment vector D and update the block connection matrix , and then fix the updated block connection matrix and further update the block assignment vector D;
[0026] S42: During the Gibbs sampling process, the nodes The most frequently assigned block s is used as the block assignment in the statistical sense, that is ; for two blocks s and t, two quantities are defined: the two quantities are specifically the number of edges actually present in the block and the maximum number of edges that could potentially be present ;
[0027] S43: According to and , the expected value of the edges starting from block s and connecting to other blocks in the core-periphery stochastic block model of the literature citation network is expressed as , and the expected value of no edges connected is ;
[0028] The posterior distribution of the core-periphery stochastic block model of the literature citation network is represented by Equation (6): ; where represents the probability of observing edges under the condition of the connection probability of the given block , represents the connection probability of block , represents the number of potential edges not connected to block , is an indicator function used to enforce the prior constraints of the hierarchical model;
[0029] S44: Use the Markov chain Monte Carlo simulation method to obtain the posterior distribution based on ; randomly assign blocks to the nodes in the literature citation network. After a certain number of iterations, randomly select a node and update its block label ;
[0030] S45: Adopt the method of random sampling to select a new block label for , that is , where L is the number of layers of the model; reverse to obtain a new block assignment ;
[0031] S46: Calculate the probability of accepting according to the Metropolis-Hastings criterion, specifically calculated by Equation (7), and Equation (7) is: ; where represents the probability of the block assignment under the condition of the given block connection matrix and the adjacency matrix The posterior probability, and respectively represent the proposed distribution probability of transferring from the old allocation to the new allocation and the probability of reverse transfer.
[0032] Furthermore, step S5 is specifically as follows: using a variation information measure to quantify the differences between different core - periphery structures, and using the minimum description length model to evaluate the fitness and network feature representation ability of each core - periphery structure; the evaluation metrics for the provided core - periphery structures are specifically as follows:
[0033] Select a node from the literature citation network , The probability of belonging to the node set is represented by Equation (VIII), and Equation (VIII) is: , where and are respectively and the number of nodes included in the literature citation network; define a discrete random variable of length , variables correspond to the number of sets included in the core - periphery partition ; the entropy of the discrete random variable is represented by Equation (IX), and Equation (IX) is: ;
[0034] Given two different core - periphery node partitions and , in the node set corresponds to the node set in ; the joint probability distribution of belonging to the node set in the partition and belonging to the node set in the partition is represented by Equation (X), and Equation (X) is ; where, represents the number of nodes that are partitioned into both and in and ;
[0035] Use mutual information to describe in about the information of the partition , in the network the uncertainty of ; The MI value with is represented by Equation (11), and Equation (11) is: ; where represents the shared information of two partitions. By subtracting the shared information from the total uncertainty, the difference between the two is obtained. represents that the node belongs to the c-th set in and the joint probability of the -th set in ; represents the probability that the node belongs to the c-th set in the partition ; represents the probability that the node belongs to the
[0036] -th set in the partition and are regarded as the uncertainties of the node sets in the partitions and ; represents the known information shared in the partitions and ; Calculate the sum of the uncertainties and eliminate the influence of the shared information to obtain the VI distance between different core - periphery node partitions. The value of the VI distance is represented by Equation (12), and Equation (12) is: .
[0037] Furthermore, when using the minimum description length model to evaluate the fitness and network feature representation ability of each core - periphery structure, the evaluation of the core - periphery stochastic block model is divided into the number of bits of the core - periphery stochastic block model itself and the number of bits required for the core - periphery stochastic block model to describe the network data, that is ; The model length for describing the network data is approximately represented by Equation (13), and Equation (13) is: ; where represents the minimum description length of the core - periphery stochastic block model; represents the coding length of the model M, that is, the number of bits required to describe the core - periphery stochastic block model structure itself; represents the number of bits required to encode the data block assignment , the block connection matrix , and the adjacency matrix under the given model M; represents the likelihood function of the adjacency matrix , the connection matrix , and the model M given the block assignment ;
[0038] Using Monte Carlo simulation in obtained from a sample, the integral of the description length of the approximate model; the sum of the intervals between the sampled samples is 1, and the distribution probabilities of the intervals are consistent and randomly combined. Through the intervals of the samples describe the sample , that is ;
[0039] Take the logarithm of the description length of the core-periphery random block model to obtain the formula for the description length of the core-periphery random block model, which is specifically represented by Formula XIV. Formula XIV is: ; where ; represents the log-likelihood function of the adjacency matrix A under the given block assignment and the block connection matrix , Under the given model M, the joint probability of observing the adjacency matrix A and the block assignment , represents the prior probability of the block assignment under the model M.
[0040] Furthermore, in step S6, structural aggregation of the core-periphery random block model of the literature citation network is performed, including estimating the uncertainty of the core-periphery structure in each core-periphery partition using the concept of information entropy. Specifically:
[0041] Given a set of nodes and the core-periphery partition , where and represent the set of all nodes and the set of core-periphery partitions respectively, , ; For the uncertainty is calculated by considering how the nodes in cluster in ; The distribution of each node in in the node set in is calculated through Formula XV. Formula XV is: ; where represents the number of nodes that belong to both the node set and the subset of the partition ; represents the number of nodes that belong to the set ;
[0042] Further obtain For The uncertainty, specifically represented by Formula XVI, which is: ; where is the node set and the partition in the subset distribution ratio;
[0043] Introduce the core-periphery structure into the calculation of uncertainty, and obtain For the uncertainty of the core-periphery partition set is represented by Formula XVII, and Formula XVII is: ; where represents the local uncertainty of the node set in the partition ; represents the minimum description length of the k-th core-periphery partition .
[0044] Furthermore, in step S6, the structure aggregation of the core-periphery stochastic block model of the literature citation network includes using the ensemble-driven clustering index to represent the reliability corresponding to each node set; given the core-periphery partition set and , There are a total of core-periphery partitions in The of the node set is represented by Formula XVIII, and Formula XVIII is: ; where the parameter is used to balance the impact of instability on growth, represents the comprehensive uncertainty of the node set
[0045] Based on the local weighted core-periphery structure aggregation method of the bipartite graph, the nodes, core node sets, and peripheral node sets in the literature citation network are all used as nodes of the bipartite graph, and the bipartite graph is combined with index and the of the core-periphery partition;
[0046] Define the bipartite graph , where , represents all nodes, S represents the union of the core node set and the peripheral node set, and W represents the weight matrix of the existence of edges in two different node sets in ;
[0047] Nodes belonging to different sets and the edge weight between is defined as shown in Equation (19): ;
[0048] Calculate the degree matrix U of the nodes in the bipartite graph, and the elements in the matrix ; According to the degree matrix U and the weight calculate the normalized weighted matrix , and find the eigenmatrix of length k ; Normalize each row of to unit length, each row of is the k-dimensional embedding of the nodes in the bipartite graph. Use K-means to cluster the nodes; the nodes in the same part in the clustering result are regarded as belonging to the same set, and the partitioning result of the bipartite graph is obtained; the partitioning result of the bipartite graph corresponds to the aggregated consensus core-periphery structure.
[0049] Compared with the prior art, the present invention can achieve the following beneficial effects:
[0050] 1. This solution can solve the problems of inconsistency and accuracy in the analysis of the core-periphery structure of the literature citation network. By introducing Bayesian prior and the model minimum description length (MDL) to optimize the aggregation scheme, the accuracy and robustness of the model are significantly improved. It can use Bayesian prior combined with network data to more accurately identify core papers and peripheral papers, and reveal the dissemination path and influence distribution of academic knowledge.
[0051] 2. This solution optimizes the aggregation scheme through MDL, effectively integrates the core-periphery structures obtained by different methods, and forms a more representative and stable consensus structure. This method can not only better evaluate the influence of core papers in the academic citation network, but also provide more accurate theoretical support for academic influence analysis, providing a strong basis for the analysis and decision-making of academic research.
[0052] 3. This solution determines the local aggregation weight by calculating the uncertainty and reliability of each node set, aggregates multiple core-periphery structures in the literature citation network analysis into a more representative and robust consensus core-periphery structure; by combining Bayesian inference with the stochastic block model, it can not only encode the prior knowledge in the literature citation network, enhance the model fitting ability, but also optimize the representation accuracy of the core-periphery structure through local weighted aggregation. The present invention provides an important theoretical basis for the importance measurement of literature, the accurate identification of core literature, and the analysis of literature influence.
[0053] In summary, the proposed method for clustering and aggregating the core-periphery stochastic block model of the literature citation network driven by Bayesian optimization, combined with Bayesian prior and model minimum description length (MDL) optimization, can accurately identify the core papers and peripheral papers in the literature citation network, reveal the dissemination path and influence distribution of academic knowledge. The introduction of Bayesian prior improves the model's ability to handle data uncertainty and noise, while MDL optimization effectively avoids model overfitting, enhances the robustness and generalization ability of the model. In addition, the clustering and aggregation scheme calculates the uncertainty and reliability of the node set, aggregates multiple core-periphery structures into a comprehensive consensus structure, and improves the stability and accuracy of the analysis results.
[0054] Compared with existing methods, the present invention can more accurately reflect the overall characteristics of the literature citation network and the influence of core papers. The practical significance of the present invention lies in providing more accurate theoretical support and tools for academic influence analysis, helping researchers better understand the dynamics and trends of academic research, providing a strong basis for academic decision-making and resource allocation, and promoting the in-depth development of academic research and knowledge innovation. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is the overall flowchart of the present invention.
[0056] Figure 2 is a schematic diagram of the spoke model in the present invention;
[0057] Figure 3 is a schematic diagram of the hierarchical model in the present invention;
[0058] Figure 4 is the aggregation flowchart of the local weighted core-periphery structure in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0059] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further describes the present invention in detail with reference to the accompanying Figures 1-4 drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and do not constitute a limitation to the present invention.
[0060] A method for clustering and aggregating the core-periphery stochastic block model of a literature citation network, as Figure 1 shown, includes the following steps:
[0061] S1. Define the literature citation network; give the definition G(V, E) of the literature citation network, where G(V, E) includes a set V composed of N nodes and a set E composed of M unweighted directed edges. Each unweighted directed edge represents the specific citation record of a literature.
[0062] S2. Construct the core-periphery stochastic block model of the literature citation network. The steps of constructing the core-periphery stochastic block model of the literature citation network in step S2 include the following steps:
[0063] S21: Give the model assumptions of the core-periphery stochastic block model of the literature citation network, that is, the core nodes and the periphery nodes are assigned to different blocks. Given a literature citation network with nodes based on the defined G(V, E), the corresponding adjacency matrix is , nodes are randomly assigned to M blocks;
[0064] S22: Represent the block assignment of the nodes in the literature citation network by a block assignment vector D of length N, specifically represented as , where the node being assigned to block s is correspondingly represented as ; Represent the connection probability between any two nodes in G(V, E) by the matrix , represents the probability that a node in block s is connected to a node in block t; In the core-periphery stochastic block model of the literature citation network, the connection probability between two nodes depends on the block assignment matrix D and the block connection matrix ;
[0065] S23: Based on the Bayesian method, relate to the prior probabilities of the block assignment vector D and the block connection matrix R to obtain Formula 1: ; Among them, The symbol represents proportionality, represents the joint probability distribution of the block assignment vector and the block connection matrix under the condition of observing the adjacency matrix , represents the prior distribution of the block assignment vector, represents the prior distribution of the block connection matrix, represents the probability of generating the adjacency matrix and given .
[0066] S3. Construct the spoke model and the hierarchical model. The spoke model in S3 is obtained by combining the two-block model with the core-periphery stochastic block model of the literature citation network; The two-block model divides the network nodes into a core block and a periphery block, and through the block connection matrix Define the structural model of the connection probability between nodes; in the spoke model, network nodes are divided into two node sets: core nodes and peripheral nodes. Core nodes are connected to each other, core nodes are connected to some peripheral nodes, and peripheral nodes are not connected to each other.
[0067] Encode the core block of the spoke model as , and the peripheral block is encoded as ; The specific definition is and ( ), where represents the connection probability between nodes inside the core block. In the spoke model, let , that is, core nodes are fully connected to each other, represents the connection probability between the core block and the peripheral block, represents the connection probability between nodes inside the peripheral block.
[0068] Figure 2 Figure Figure 2 is a schematic diagram of the spoke model. Among them, the connection between nodes in the core layer is the densest, and the number of connections between nodes in the peripheral layer, which is opposite to the core layer, and core nodes is the second, and the connection between nodes in the peripheral layer is the sparest.
[0069] The prior constraints of the block connection matrix R of the spoke model are represented by Formula 2 and Formula 3; Formula 2 is: ; Formula 3 is ; where is an indicator function used to define the legality constraint conditions of the block connection matrix in the spoke model, represents the prior probability distribution of the block connection matrix .
[0070] In S3, the hierarchical model is obtained by combining the k-core model with the core-periphery stochastic block model of the literature citation network; in the hierarchical model, network nodes are stratified based on the k-core decomposition method. Among them, the k-core represents the set of all nodes in the network that are connected to at least k other nodes. The k-shell hierarchical structure is defined in the k-core decomposition method. Each layer of the k-shell consists of all nodes in the k-core, but does not include the nodes in the (k + 1)-core. The connection degree of the hierarchical structure gradually increases from the outside to the inside. Figure 3 Figure Figure 3 is a schematic diagram of the hierarchical model. Nodes in the innermost first layer are likely to be connected to nodes in other layers, and the probability of connection between nodes in other layers and nodes in more peripheral layers is smaller.
[0071] The prior constraints of the block connection matrix R of the hierarchical model are represented by Formula 4 and Formula 5; Formula 4 is , and Formula 5 is ; where is an indicator function used to ensure the block connection matrix of the hierarchical model complies with its structural assumptions, represents the prior probability distribution of the block connection matrix .
[0072] S4. Use Bayesian inference to infer the core-periphery stochastic block model of the literature citation network. The specific steps of using Bayesian inference to infer the core-periphery stochastic block model of the literature citation network in S4 are as follows:
[0073] S41: Based on the idea of Gibbs sampling, alternately sample the block assignment vector D and the block connection matrix , that is, first fix the block assignment vector D and update the block connection matrix , then fix the updated block connection matrix and further update the block assignment vector D.
[0074] S42: Take the block s to which the node is most frequently assigned during the Gibbs sampling process as the block assignment in the statistical sense, that is ; for two blocks s and t, clarify two quantities: The two quantities are specifically the number of edges actually present in the block and the maximum number of possible edges .
[0075] S43: According to and , represent the expectation that the edges starting from block s in the core-periphery stochastic block model of the literature citation network are connected to other blocks as , and the expectation of no edge connection is .
[0076] The posterior distribution of the core-periphery stochastic block model of the literature citation network is represented by Equation (6): ; where represents the probability of observing edges under the condition of the connection probability of the given block , represents the connection probability of block , represents the number of potential edges where block is not connected, is an indicator function used to enforce the prior constraints of the hierarchical model.
[0077] S44: Use the Markov chain Monte Carlo simulation method to obtain the posterior distribution based on , randomly perform block assignment on the nodes in the literature citation network. After a certain number of iterations, randomly select a node and update its block label .
[0078] S45: Use random sampling to select its new block label , that is , where L is the number of model layers; reverse to obtain a new block assignment .
[0079] S46: According to the Metropolis-Hastings criterion, calculate the probability of accepting , specifically calculated through Formula Seven, and Formula Seven is: ; where represents the posterior probability of the block assignment under the given block connection matrix and the adjacency matrix , and respectively represent the proposed distribution probability of transferring from the old assignment to the new assignment and the probability of reverse transfer.
[0080] S5: Based on the principle of the minimum description length (MDL) of the model, optimize and evaluate the core-periphery stochastic block model of the literature citation network. Step S5 is specifically as follows: Use the variational information (VI) metric to quantify the differences between different core-periphery structures, and use the minimum description length (MDL) model to evaluate the fitness and network feature representation ability of each core-periphery structure; the evaluation indicators of the provided core-periphery structures are specifically as follows:
[0081] Select a node from the literature citation network , The probability of belonging to the node set is represented by Formula Eight, and Formula Eight is: , where and are respectively and the number of nodes included in the literature citation network; define a discrete random variable of length , variables correspond to the number of sets included in the core-periphery partition ; the entropy of the discrete random variable is represented by Formula Nine, and Formula Nine is: ;
[0082] Given two different core-periphery node partitions and , in the node set in corresponds to the node set in ; The joint probability distribution of the nodes that belong to the node set in the partition and also belong to the node set in the partition is represented by Formula X, and Formula X is ; where represents the number of nodes that are partitioned into both and in ; and ; and also partitioned into ;
[0083] Use mutual information (MI) to describe the information about the partition in , in the uncertainty is ; The MI value between and is represented by Formula XI, and Formula XI is: ; where represents the shared information volume of the two partitions. By subtracting the shared information from the total uncertainty, the difference between the two is obtained. represents the joint probability that a node belongs to the c-th set in and the -th set in ; represents the probability that a node belongs to the c-th set in the partition ; represents the probability that a node belongs to the -th set in the partition
[0084] Regard and as the uncertainties of the node sets in the partitions and the partition ; represents the known information shared in the partitions and the partition ; Calculate the sum of the uncertainties and eliminate the influence of the shared information to obtain the VI distance between different core-periphery node partitions. The value of the VI distance is represented by Formula XII, and Formula XII is: .
[0085] When evaluating the fitness and network feature representation ability of each core - periphery structure using the Minimum Description Length (MDL) model, the evaluation of the core - periphery stochastic block model is divided into the number of bits of the core - periphery stochastic block model itself and the number of bits required for the core - periphery stochastic block model to describe network data, that is ; the model length for describing network data is approximately represented by Equation (13), and Equation (13) is: ; where represents the minimum description length of the core - periphery stochastic block model; represents the coding length of model M, that is, the number of bits required to describe the structure of the core - periphery stochastic block model itself; represents the number of bits required to encode the data block assignment , the block connection matrix , and the adjacency matrix ; represents the likelihood function of the adjacency matrix given the block assignment , the connection matrix , and model M;
[0086] Use Monte Carlo simulation to obtain in samples of to approximate the integral of the model description length; the sum of the intervals between the sampled samples is 1, and the distribution probabilities of the intervals are consistent and randomly combined. The samples are described by the intervals of the samples , that is ;
[0087] Take the logarithm of the description length of the core - periphery stochastic block model to obtain the formula for the description length of the core - periphery stochastic block model, which is specifically represented by Equation (14), and Equation (14) is: ; where ; represents the log - likelihood function of the adjacency matrix A given the block assignment and the block connection matrix ; the joint probability of observing the adjacency matrix A and the block assignment given model M; represents the prior probability of the block assignment under model M.
[0088] S6: Aggregate the structure of the core - periphery stochastic block model of the literature citation network. Figure 4It is a flowchart for aggregating the local weighted core-periphery structure. By calculating the uncertainty and reliability of each node set, the local aggregation weight is determined, and multiple core-periphery structures in the literature citation network analysis are aggregated into a more representative and robust consensus core-periphery structure.
[0089] In step S6, structural aggregation is performed on the core-periphery stochastic block model of the literature citation network, including estimating the uncertainty of the core-periphery structure in each core-periphery partition using the concept of information entropy. Specifically:
[0090] Given a set of nodes and a core-periphery partition , where and represent all sets of nodes and the set of core-periphery partitions respectively, , ; For the uncertainty is calculated by considering how the nodes in are clustered in ; the distribution of each node in in the set of nodes in is calculated by formula fifteen, and formula fifteen is: ; where represents the number of nodes that belong to both the node set and the subset of the partition ; represents the number of nodes that belong to the set .
[0091] Based on the distribution, the uncertainty of for is further obtained, which is specifically represented by formula sixteen, and formula sixteen is: ; where is the distribution ratio of the subset of the node set and the partition .
[0092] The of the core-periphery structure is introduced in the calculation of uncertainty, and the uncertainty for the set of core-periphery partitions is obtained, which is represented by formula seventeen, and formula seventeen is: ; where represents the local uncertainty of the node set in the partition ; Denote the k-th core-periphery partition of the minimum description length.
[0093] In step S6, structure aggregation is performed on the core-periphery stochastic block model of the literature citation network, including using the ensemble-driven clustering index (ECI) to represent the reliability corresponding to each node set; given the core-periphery partition set and , There are a total of core-periphery partitions in , and the of the node set is represented by Equation XVIII, and Equation XVIII is:[[]] ; where the parameter is used to balance the influence of instability on growth, and represents the comprehensive uncertainty of the node set
[0094] Based on the global diversity and local reliability of the core-periphery structure, a local weighted core-periphery structure aggregation method based on a bipartite graph is proposed. In the local weighted core-periphery structure aggregation method based on a bipartite graph, the nodes, core node sets, and peripheral node sets in the literature citation network are all used as nodes of the bipartite graph, and the bipartite graph is combined with index and the of the core-periphery partition.
[0095] Define the bipartite graph , where , represents all nodes, S represents the union of the core node set and the peripheral node set, and W represents the weight matrix of the existence of edges between two different node sets in .
[0096] The edge weight and between nodes belonging to different sets is defined as Equation XIX:[[]] ;
[0097] Calculate the degree matrix U of the nodes in the bipartite graph, and the element in the matrix; calculate the normalized weighted matrix according to the degree matrix U and the weight , and find the characteristic matrix of length k of ; Normalize each row of to unit length, Each line is the k-dimensional embedding of the nodes in the bipartite graph, and K-means is used to cluster the nodes; the nodes in the same part in the clustering result are regarded as belonging to the same set, and the partitioning result of the bipartite graph is obtained; the partitioning result of the bipartite graph corresponds to the aggregated consensus core-periphery structure.
[0098] S7: Obtain the influential literature, i.e., the core nodes.
[0099] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A literature citation network core-periphery random block model clustering method, characterized in that: The following steps are involved: S1. Define the literature citation network; give the definition of the literature citation network G(V, E), G(V, E) includes a set V consisting of N nodes and a set E consisting of M unweighted directed edges, where each unweighted directed edge represents a specific citation record of the document; S2, constructing a core-periphery random block model of literature citation network; S3, construct hub-and-spoke model and hierarchical model; S4, using Bayesian inference to infer the core-periphery random block model of the literature citation network; S5: Based on the principle of minimum description length of the model, the core-periphery random block model of the literature citation network is optimized and evaluated; S6: Structural aggregation of the core-periphery random block model of the literature citation network; S7: Obtain influential documents, i.e. core nodes; The construction of the core-periphery random block model of the literature citation network in step S2 includes the following steps: S21: The model assumption of the core-periphery random block model of the literature citation network is given, that is, the core nodes and peripheral nodes are assigned to different blocks. Based on the defined G(V, E), given a literature citation network with N nodes, the corresponding adjacency matrix is A N×N , N nodes are randomly assigned to M blocks; S22: The block allocation status of nodes in the literature citation network is represented by a block allocation vector D of length N, which is specifically represented as D = {D1, D2, ...D N }, where node v i is assigned to block s and is represented as D i =s; the connection probability of any two nodes in G(V, E) is represented by the matrix R, R st represents the probability of connecting a node in block s to a node in block t; in the core-periphery random block model of the reference network, the connection probability between two nodes depends on the block assignment matrix D and the block connection matrix R; S23: Based on the Bayesian method, P(D, R|A) is linked to the prior probability of the block allocation vector D and the block connection matrix R, and Formula 1 is obtained: P(D, R|A)∝P(D)P(R)P(A|D, R); where the ∝ symbol indicates proportionality, P(D, R|A) represents the joint probability distribution of the block allocation vector D and the block connection matrix R under the condition that the adjacency matrix A is observed, P(D) represents the prior distribution of the block allocation vector, P(R) represents the prior distribution of the block connection matrix, and P(A|D,R) represents the probability of generating the adjacency matrix A given D and R.
2. The document citation network core-periphery random block model clustering method according to claim 1 is characterized in that: The hub-and-spoke model in S3 is obtained by combining the two-block model with the core-periphery random block model of the literature citation network; the two-block model divides the network nodes into core blocks and peripheral blocks, and defines the structural model of the connection probability between nodes through the block connection matrix R; in the hub-and-spoke model, the network nodes are divided into two sets of nodes, core nodes and peripheral nodes, the core nodes are connected to each other, the core nodes and some peripheral nodes are connected to each other, and the peripheral nodes are not connected to each other; The core block of the hub-and-spoke model is encoded as block1, and the peripheral block is encoded as block2; the specific definition is R 11 >R 12 >R 22 And R 11 =1,R 12 =a,R 22 =β(1≥α>>β≥0), where R 11 represents the connection probability of the internal nodes of the core block. In the hub-and-spoke model, R 11 =1, that is, the core nodes are fully connected, R 12 represents the connection probability between the core block and the peripheral block, R 22 represents the connection probability of the internal nodes of the peripheral block; The prior constraints of the block connection matrix R of the hub-and-spoke model are expressed by Formula 2 and Formula 3: Formula 2 is: Formula 3 is P(R)∝const hub-and-spoke (R); where const hub-and-spoke (R) is an indicator function used to define the legality constraints of the block connection matrix R in the hub-and-spoke model, and P(R) represents the prior probability distribution of the block connection matrix R.
3. The document citation network core-periphery random block model clustering method according to claim 2, characterized in that: The hierarchical model in S3 is obtained by combining the k-core model with the core-periphery random block model of the literature citation network. In the hierarchical model, the network nodes are layered based on the k-core decomposition method, where the k-core represents the set of nodes in the network that are connected to at least k other nodes. The prior constraints of the block connection matrix R of the hierarchical model are expressed by Formula 4 and Formula 5; Formula 4 is Formula 5 is P(R)∝const layered (R); where const layered (R) is an indicator function used to ensure that the block connection matrix R of the hierarchical model conforms to its structural assumptions, and P(R) represents the prior probability distribution of the block connection matrix R.
4. The document citation network core-periphery random block model clustering method according to claim 3, characterized in that: In S4, Bayesian reasoning is used to infer the core-periphery random block model of the literature citation network, which specifically includes the following sub-steps: S41: Based on the Gibbs sampling idea, alternately sample the block allocation vector D and the block connection matrix R, that is, first fix the block allocation vector D and update the block connection matrix R, then fix the updated block connection matrix R and further update the block allocation vector D; S42: The node node in the Gibbs sampling process i The most frequently allocated block s is used as the statistical block allocation, i.e. For two blocks s and t, two quantities are specified: the two quantities are specifically the number of edges actually present in the block e st and the maximum number of possible edges E st ; S43: According to e st and E st , the expectation of the edges from block s to other blocks in the core-periphery random block model of the literature citation network is expressed as Y s , the expectation without edge connection is N s ; The posterior distribution P(D|R, A) of the core-periphery random block model of the literature citation network is expressed by Formula 6: in, represents the connection probability R at a given block s s Under the condition of s The probability of an edge, R s represents the connection probability of block s, N s Indicates the number of potential edges that block s is not connected to, const layered is an indicator function used to enforce a priori constraints on the hierarchical model; S44: Use the Markov Chain Monte Carlo simulation method to obtain the posterior distribution P(D|R, A) based on P(D|R, A)∝P(A|D, R)P(D); randomly allocate blocks to nodes in the literature citation network, and after a certain number of iterations, randomly select a node node i And update its block tag D i ; S45: Random sampling is used for node i Select its new block label D i ,Right now Where L is the number of model layers; i Reverse and get the new block allocation D′; S46: According to the Metropolis-Hastings criterion, the probability of accepting D′ is calculated, which is specifically calculated by Formula 7, which is: where P(D′|R, A) represents the posterior probability of block assignment D′ given the block connectivity matrix R and the adjacency matrix A, P(D|D′) and P(D′|D) represent the proposed distribution probability of transferring from the old assignment D to the new assignment D′ and the probability of reverse transfer, respectively.
5. The document citation network core-periphery random block model clustering method according to claim 4 is characterized by: Step S5 specifically includes: using variation information metrics to quantify the differences between different core-periphery structures, and using the minimum description length model to evaluate the fit and network feature representation capabilities of each core-periphery structure; the evaluation indicators of the core-periphery structure provided are as follows: Select a node from the literature citation network i , node i Belongs to the node set S c The probability of is expressed by formula eight, which is: where n c and n are S c and the number of nodes contained in the literature citation network; define a discrete random variable of length C, where the C variables correspond to the number of sets contained in the core-periphery partition P; The entropy of a discrete random variable C is expressed by Formula 9, which is: Given two different core-periphery node partitions P i and P j , in P i The node set S in c Corresponding to P j The node set S in c′ ; node i In dividing P i belongs to the node set S c And in the division P j belongs to the node set S c′ The joint probability distribution of is expressed by formula 10, which is in, Indicates that in P i and P j It is divided into S c Also classified as S c′ The number of nodes; Using mutual information to describe P i About the division of P j Select node in the network i , node i In P i The uncertainty in E(P i );P i With P j The MI value is expressed by Formula 11, which is: Among them, MI(P i , P j ) represents the amount of shared information between the two partitions. By subtracting the shared information from the total uncertainty, we get the difference between the two. P(c, c′) represents the nodes that belong to both P i The cth set and P j The joint probability of the cth set in the partition P, P(c) indicates that the node belongs to the partition P i The probability of the cth set in the partition P, P(c′) indicates that the node belongs to the partition P j The probability of the c′th set in ; E(P i ) and E(P j ) is considered as dividing P i and partition P j The uncertainty of the node set, MI(P i , P j ) represents the partition P i and partition P j The known information shared in the network is used to calculate the sum of uncertainties and eliminate the influence of shared information to obtain the VI distance between different core-periphery node partitions. The value of the VI distance is expressed by Formula 12, which is: VI(P i ,P j )=E(P i )+E(P j )-2MI(P i ,P j )。 6. The document citation network core-periphery random block model clustering method according to claim 5, characterized in that: When using the minimum description length model to evaluate the fit of each core-periphery structure and the ability to represent network features, the evaluation of the core-periphery random block model is divided into the number of bits of the core-periphery random block model itself and the number of bits required by the core-periphery random block model to describe network data, namely, MDL M =L(M)+L(D, R, A|M); The model length describing the network data is approximately expressed by Formula 13, which is: L(D, R, A|M)≈P(A, D M |M); among them, MDL M represents the minimum description length of the core-periphery random block model; L(M) represents the coding length of model M, that is, the number of bits required to describe the core-periphery random block model structure itself; L(D, R, A|M) represents the number of bits required to encode the data block allocation D, block connection matrix R, and adjacency matrix A under a given model M; P(A, D M |M) represents the likelihood function of the adjacency matrix A under a given block assignment D, connectivity matrix R, and model M; Monte Carlo simulation is used to obtain n samples of R in P(R), approximating the integral of the model description length; the sum of the intervals between the samples obtained is 1, and the distribution probability of the intervals is consistent and randomly combined, through the blank interval of the samples i =R i+1 -R i Describe sample R s , that is, R s =1-Σ i≤s block i ; Taking the logarithm of the core-periphery random block model description length, the core-periphery random block model description length is obtained as a formula, which is specifically expressed by Formula 14. Formula 14 is: in, MAX=max(logP(A|D M ,R s ,M)); log P(A|D M , R S , M) indicates that D is allocated in a given block M and the block connection matrix R s Under this condition, the log-likelihood function of the adjacency matrix A, P(A, D M |M) Given a model M, we observe the adjacency matrix A and the block assignment D M The joint probability, P(D M |M) means that under model M, block allocation D M The prior probability of .
7. The document citation network core-periphery random block model clustering method according to claim 6, characterized in that: In step S6, the core-periphery random block model of the literature citation network is subjected to structural aggregation, including estimating the uncertainty of the core-periphery structure in each core-periphery partition using the concept of information entropy, specifically: Given a node set s i ∈S and core-periphery partition p k ∈P, where S and P represent the set of all nodes and the set of core-periphery partitions, respectively. P = {p1, p2, ..., p k ,…,p M };s i For p k The uncertainty of s is taken into account i How do the nodes in p gather together? k Calculate from ; Calculate s by formula 15 i Each node in p k The node collection in The distribution of , formula 15 is: in, Indicates that it belongs to the node set s at the same time i and divide p k A subset of The number of nodes; Indicates that it belongs to the set S i The number of nodes; Further obtain s through distribution i For p k The uncertainty of is specifically expressed by Formula 16, which is: in, is a set of nodes s i and divide p k A subset of The distribution ratio of Introducing the MDL of the core-periphery structure into the uncertainty calculation, we obtain s i For the uncertainty E(s) of the core-periphery partition set P i , P), is expressed by Formula 17, which is: Among them, E k (s i , p k ) represents the node set s i In dividing p k Local uncertainty in MDL(p k ) represents the kth core peripheral partition p k The minimum description length.
8. The document citation network core-periphery random block model clustering method according to claim 7, characterized in that: In step S6, the core-periphery random block model of the literature citation network is structurally aggregated, including using an integrated driven clustering index to represent the reliability corresponding to each node set; given the core-periphery partition set P and s i , there are M core-periphery partitions in P, and the node set s i The ECI is expressed by Formula 18, which is: The parameter δ is used to balance the effect of instability on the growth of ECI, E(s i , P) represents the node set s i The combined uncertainty in the core-periphery partition set P; Based on the local weighted core-periphery structure aggregation method of bipartite graph, the nodes in the literature citation network, the core node set and the peripheral node set are all used as nodes of the bipartite graph, and the bipartite graph is combined with the ECI index and the MDL of core-periphery division; Define a bipartite graph Where V = N∪S, N represents all nodes, S represents the union of the core node set and the peripheral node set, and W represents The weight matrix of the edges between two different sets of nodes in , In={in 11 ,In 12 ,…,In ij ,…,In mm }; Nodes v belonging to different sets i and v j The edge weight w between ij The definition of is shown in Formula 19: Calculate the degree matrix U of the nodes in the bipartite graph. The elements in the matrix Calculate the normalized weight matrix NW=D according to the degree matrix U and weight W -1 W, and find the length k of the characteristic matrix NW of NW F ={f1, f2, ..., f k }; NW F Each row of is normalized to unit length, NW F Each row of is the k-dimensional embedding of a node in the bipartite graph. K-means is used to cluster the nodes. Nodes in the same part of the clustering result are considered to belong to the same set, and the partition result of the bipartite graph is obtained. The partition result of the bipartite graph corresponds to the consensus core-periphery structure after aggregation.
Citation Information
Patent Citations
Network security situation assessment method based on random forest and Bayesian network
CN111191683A
Implementation method for literature identification and technical path evolution
CN117112784A