Stochastic Topic-Block Model for Textual Network Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for analyzing networks of nodes associated with textual data fail to effectively combine network interactions and textual content, leading to misleading cluster formations, especially in cases where textual content influences node connections, and are limited in handling large-scale and directed/undirected networks with high computational complexity.
Innovation Solution
The stochastic topic-block model (STBM) integrates network analysis with semantic analysis using a classification variational expectation-maximization (C-VEM) algorithm, allowing for the inference of topic-meaningful clusters that consider both network structure and textual content, and is applicable to both directed and undirected networks with reasonable computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If only network information (binary edges) is used for clustering, then the clustering process is computationally simple, but the clusters are not meaningful and fail to reflect textual content
Solution Approach 1:
The patent combines network structure information (binary edges from adjacency matrix) with textual content information (topics from LDA model) into a unified clustering framework. The stochastic block model integrates both data types to simultaneously capture structural patterns and semantic meaning, resolving the contradiction between simplicity and information completeness.
Solution Approach 2:
The invention creates a composite modeling approach that merges two different types of data representations: the structural adjacency matrix and the semantic topic distributions. This composite model allows the clustering algorithm to leverage both the topological structure of the network and the semantic content of interactions, producing clusters that are both structurally coherent and semantically meaningful.
2Measurement precision
If textual content analysis is integrated into network clustering, then cluster accuracy and interpretability improve, but computational complexity increases
Solution Approach 1:
The patent performs preliminary topic modeling using LDA to extract semantic themes from textual data before the clustering process. By pre-processing the text data into topic distributions, the method reduces the complexity of simultaneously analyzing both text and network structure, while still achieving high cluster accuracy through the integrated stochastic block model.
Solution Approach 2:
The invention transforms textual content into a different dimensional representation (topic distributions) that can be seamlessly integrated with network structure data. This dimensionality transformation allows the clustering algorithm to operate on a unified mathematical framework without the full computational burden of processing raw text during clustering, thus improving accuracy while managing complexity.
3Loss of information
If comprehensive text analysis is performed to determine clusters, then meaningful topic-based clusters are obtained, but processing time increases
Solution Approach 1:
The patent extracts only the essential semantic information from textual data in the form of topic distributions, rather than performing comprehensive text analysis during clustering. By extracting and retaining only the relevant semantic features (topic proportions), the method maintains semantic information while significantly reducing processing time compared to analyzing full text content.
4Adaptability or versatility
If the model handles both directed and undirected networks with large scale, then versatility is improved, but computational complexity increases
Solution Approach 1:
The stochastic block model implemented in the patent is designed to handle both directed and undirected networks within a single unified framework. The model's mathematical formulation naturally accommodates different network types without requiring separate algorithms, achieving versatility while managing computational complexity through efficient inference methods.
Data Source
AI summary
The invention relates to a method for clustering nodes of a network, the network comprising nodes associated with message edges of text data, the method comprising an initialization step of determination of a first initial clustering of the nodes, and a step of iterative inference of a generative model of text documents. Edges are modeled with a Stochastic Block Model (SBM) and the sets of documents between and within clusters are modeled according to a generative model of documents. The inference step comprises iteratively modelling the text documents and the underlying topics of their textual content, and updating the clustering as a function of the modelling, until a convergence criterion is fulfilled and an optimized clustering and corresponding optimized values of the parameters of the models are output.


