Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

63 results about "Topic model" patented technology

In machine learning and natural language processing, a topic model is a type of statistical model for discovering the abstract "topics" that occur in a collection of documents. Topic modeling is a frequently used text-mining tool for discovery of hidden semantic structures in a text body. Intuitively, given that a document is about a particular topic, one would expect particular words to appear in the document more or less frequently: "dog" and "bone" will appear more often in documents about dogs, "cat" and "meow" will appear in documents about cats, and "the" and "is" will appear equally in both. A document typically concerns multiple topics in different proportions; thus, in a document that is 10% about cats and 90% about dogs, there would probably be about 9 times more dog words than cat words. The "topics" produced by topic modeling techniques are clusters of similar words. A topic model captures this intuition in a mathematical framework, which allows examining a set of documents and discovering, based on the statistics of the words in each, what the topics might be and what each document's balance of topics is.

AI auxiliary reading method and system based on dynamic partitioning knowledge classification self-verification

PendingCN120508651ABiological modelsNatural language data processingKnowledge classificationText entry
The invention provides an AI auxiliary reading method and system based on dynamic partitioning knowledge classification self-verification, and relates to the technical field of artificial intelligence, the method comprises the following steps: carrying out dynamic semantic partitioning on a super-long text to obtain a plurality of initial text units with consistent internal semantics; performing topic model construction according to the initial text unit and a layered Dirichlet process to obtain a topic distribution tag; receiving a user query request, and extracting a retrieval text set according to the hierarchical retrieval strategy and a retrieval text corresponding to the user query request; inputting the constraint cue word extracted by the topic distribution tag corresponding to the retrieval text set, the retrieval text set and the retrieval text into a text generation model to output a to-be-calibrated text; if the error scoring function of the to-be-calibrated text and any initial text unit in the retrieval text set is smaller than a preset error threshold value, the to-be-calibrated text serves as a credible text to be output to a user. According to the invention, the credibility and accuracy of the answer can be enhanced.
Owner:WUHAN DINGSEN ELECTRONIC TECH CO LTD

Large language model privacy preservation system

Computer-implemented methods for a large language model privacy preservation system. Aspects include receiving prompt data from a user device. Aspects further include generating pre-processed prompt data using the prompt data from the user device. Aspects also include identifying a category for the pre-processed prompt data using topic modeling. Aspects include generating normalized prompt data using the pre-processed prompt data. Aspects further include storing the category and the normalized prompt data.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Earth and rockfill dam illness feature mining method and system based on historical text data

The invention discloses an earth and rockfill dam illness feature mining method and system based on historical text data, and the method comprises the steps: collecting earth and rockfill dam historical illness data, constructing an earth and rockfill dam illness diagnosis text corpus set, carrying out the structured preprocessing of the corpus set, and obtaining an illness feature and danger removal measure text corpus set; generating a structured word sequence subset through a word segmentation tool; processing the structured word sequence subset by adopting an LDA topic model, determining an optimal topic number through a confusion degree curve, and outputting a final topic and a corresponding topic word; constructing a visual network graph based on the subject term co-occurrence frequency; and carrying out centrality analysis on nodes in the visual network diagram, identifying key nodes in the visual network diagram, quantitatively analyzing association rules of the danger characteristics and danger removing measures, and completing feature mining of the earth and rockfill dam danger. The problems that traditional manual diagnosis is high in subjectivity and low in utilization rate of historical engineering data are solved, and intelligent auxiliary decision making of the earth and rockfill dam danger is achieved.
Owner:NANJING HYDRAULIC RES INST

Cross-platform theme decoupling Chinese aggressive language detection method based on causal relationship

The invention provides a cross-platform theme decoupling Chinese aggressive language detection method based on a causal relationship, and belongs to the field of natural language processing in an artificial intelligence technology.The method comprises the steps that firstly, a causal graph is constructed, and themes related to a platform and aggressive languages unrelated to the platform are decoupled; constructing a Chinese aggressive language detection model based on the causal graph; training the Chinese aggressive language detection model by using the Chinese aggressive language data set, performing topic modeling on the Chinese aggressive language data set by using a probability topic model, and constructing a loss function for optimizing the Chinese aggressive language detection model by combining results of the two; and detecting a to-be-detected text by using the Chinese aggressive language detection model combined with the loss function. According to the method, the causal relationship is introduced to effectively decouple the Chinese aggressive language data, so that more accurate Chinese aggressive language detection is realized in a cross-platform environment.
Owner:NANJING UNIV OF POSTS & TELECOMM

Topic model-based low-altitude economy teaching text classification method and system

The invention provides a low-altitude economy teaching text classification method and system based on a topic model, and the method comprises the steps: firstly obtaining a low-altitude economy teaching text set, constructing a low-altitude economy teaching term association network, enabling nodes to be teaching terms, and enabling edges to represent the term co-occurrence association strength, then, based on the low-altitude economy teaching term association network, mining an initial low-altitude economy teaching subject set which comprises a plurality of subject units formed by close co-occurrence terms, then, performing semantic vector representation on text units, optimizing the initial subject set, and based on the optimized subject set, constructing a low-altitude economy teaching text subject classification model; the method comprises the following steps of: inputting a to-be-classified text into a topic feature library and a topic classification rule set, finally inputting the to-be-classified text into a low-altitude economy teaching text topic classification model, and generating a topic classification result through topic feature matching and classification decision, so that the accuracy and efficiency of low-altitude economy teaching text classification are improved.
Owner:HUNAN INSTITUTE OF ENGINEERING

Trending topic discovery with keyword-based topic model

The present disclosure relates to a system, a method, and a product for topic discovery. The system includes a memory storing instructions; and a processor in communication with the memory. When the processor executes the instructions, the instructions are configured to cause the processor to: obtain text data, conduct pre-processing on the text data to obtain pre-processed text data, extract an entity list and a keyword list based on the pre-processed text data, generate an entity embedding list based on the entity list, clusterize the entity list based on the entity embedding list to obtain a plurality of entity clusters, each entity cluster comprising at least one entity, retrieve a co-occurring keyword list based on the plurality of entity clusters, the entity list, and the keyword list, and obtain a topic for each entity cluster of the plurality of entity clusters based on the co-occurring keyword list.
Owner:ACCENTURE GLOBAL SOLUTIONS LTD

Traffic transportation government affair hotline mining method based on natural language processing

The invention discloses a traffic transportation government affair hotline mining method based on natural language processing, and belongs to the technical field of artificial intelligence and machine learning. According to the method, for structured and unstructured data in a traffic transportation government affair hotline, firstly, a five-class word segmentation dictionary containing cleaning words, noise words, synonyms, additive words and stop words is constructed; converting the unstructured text into structured data through improved data cleaning, text word segmentation and feature representation; carrying out clustering analysis by utilizing an LDA topic model, and extracting public demand topics and high-frequency keywords; hot appeals and trend changes are mined in combination with space-time analysis and association analysis, and finally a visual analysis report is generated. According to the method, the problems of poor model interpretability and insufficient field adaptability in the prior art are solved, the accuracy of hotline data processing and the effectiveness of theme recognition are remarkably improved through multi-dictionary collaborative optimization and field knowledge fusion, and accurate decision support is provided for a traffic transportation management department. A real taxi field case in a certain city is used as an example for research, an experiment proves that the method has an accurate theme identification function, and the complaint and report work order amount in the traffic transportation government affair hotline taxi field in the city is reduced by 20% on year-on-year basis in 2024.
Owner:乌若愚

Updating support documentation for developer platforms with topic clustering of feedback

A method includes obtaining at least one feedback response from a developer regarding an answer and corresponding source documents of the answer in a discussion thread initiated by the developer. The method further includes converting the discussion thread into a topic model clustering input, responsive to the feedback response specifying an unsatisfactory category of feedback responses, to obtain a multitude of topic clustering model inputs. The method further includes periodically processing the multitude of topic clustering model inputs by a thread-topic clustering model to obtain a multitude of candidate topics. The method further includes processing, by an answer generation model, a first candidate topic of the multitude of candidate topics to obtain a multitude of corresponding documentation recommendations for the candidate topic. The method further includes presenting the first candidate topic and the multitude of corresponding documentation recommendations.
Owner:INTUIT INC

A topic model updating method and system, a storage medium and a server

The embodiment of the application discloses a kind of theme model updating method, system and storage medium and server, apply to the information processing technical field based on artificial intelligence.Theme model system will obtain the first label semantic feature and the second label semantic feature corresponding respectively to multiple old theme labels in the first theme model and multiple new theme models in the second theme model, and based on the first label semantic feature and the second label semantic label, mapping relationship is established between old theme label and new theme label, and then the old theme label in the first theme model is updated based on the mapping relationship.The updating of the first theme model existing in the system is realized automatically, the efficiency of the theme model is improved, and the theme model with larger dimensionality can also be updated, and the updating of the first theme model is not limited by the theme model acquisition method.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Text style recognition method and device, equipment and medium

The invention relates to the technical field of style recognition, and discloses a text style recognition method. The method comprises the steps of obtaining a to-be-recognized text, and determining a text publishing platform corresponding to the to-be-recognized text; performing topic recognition on the to-be-recognized text through an LDA topic model corresponding to the text publishing platform to obtain at least one text topic and topic keywords corresponding to the text topics; performing topic clustering on all the text topics and the topic keywords corresponding to the text topics through a visual clustering tool to obtain a text topic cluster; sending the text theme cluster to the client, and receiving an optimal theme quantity fed back by the client; and performing style mapping on the target theme in the optimal theme quantity to obtain a text style. According to the method, through the LDA topic model and the style mapping rule, the topic and the style in the text are recognized, the accuracy of style recognition is improved, and the efficiency of style recognition is improved.
Owner:SHENZHEN DONSON CLOUD TECHNOLOGY CO LTD

A Spatial Gene Identification and Extraction Method Based on Social Media Text Data

This invention discloses a method for spatial gene identification and extraction based on social media text data, comprising the following steps: collecting online text data about a city, then preprocessing the data to obtain dataset D1; constructing a dictionary and vector space in analysis software, introducing an LDA topic model, and classifying the obtained dataset D1 into topics; merging synonyms in each topic, and performing synonym replacement in dataset D1 to obtain dataset D2; counting the co-occurrence frequency of keywords in dataset D2 and constructing a co-occurrence matrix M; and using a hierarchical clustering model to cluster the semantic network analysis results to obtain spatial combination patterns, i.e., spatial genes. This invention collects online text data about a specific city from multiple social media platforms, providing a practical technical means for urban researchers to identify urban spatial genes by obtaining rich, non-invasive data.
Owner:SOUTHEAST UNIV +1

Topic model generation method and device, storage medium and program product

The embodiment of the invention provides a topic model generation method and device, a storage medium and a program product, and relates to the technical field of artificial intelligence. The method comprises the following steps: preprocessing an obtained user comment text and comment metadata, and correspondingly obtaining a word frequency vector and a text label; inputting the word frequency vector into a preset topic model, converting the word frequency vector into potential space representation, and performing normalization processing to obtain topic distribution corresponding to a user comment text output by the topic model; inputting the user comment text into a preset attention mechanism model to obtain a text embedding vector corresponding to the user comment text; and based on the topic distribution, the text embedding vector and the text label, carrying out joint training, and determining a trained target model. The attention mechanism and the comment metadata are combined in the modeling process, the trained target model is determined, and the problems that in the prior art, a topic model is insufficient in real scene data processing capacity and low in efficiency are solved.
Owner:CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1

Long text abstract generation method based on hierarchical graph comparison theme

PendingCN122021560ASemantic analysisText processingDocument representationInformation coverage
The invention discloses a long text abstract generation method based on hierarchical graph comparison themes, which comprises the following steps of: 1, preprocessing an original document, dividing sentence sequences, and obtaining global context-aware sentence and document representation through a hierarchical encoder network; 2, deducing document-level and sentence-level topic distribution by using a neural topic model; and 3, constructing a supervision graph based on a standard abstract to perform graph comparison learning so as to close the topic representation of a document and a key sentence and push redundant information. According to the method, the deep semantic structure of the long document can be effectively captured, so that the theme consistency and the information coverage degree of the abstract can be improved, and the redundancy is reduced.
Owner:ANHUI AGRICULTURAL UNIVERSITY

Multimodal context selection for large language model based resolutions addressing technical issues

A method for technical issue resolution. The method includes: receiving, from a user, a text query concerning a technical issue; obtaining query-related context relevant to the text query; and processing, through a large language model (LLM), the text query and the query-related context to produce a multimodal query response used by the user to address the technical issue. More specifically, embodiments described herein utilize text topic and zero shot classification models to translate multimodal technical documentation (e.g., including text and images) into topic relevant metadata; and process queries, pertaining to technical issues, using a multimodal LLM provided with query-related text and image context derived from said topic relevant metadata.
Owner:DELL PROD LP

Identifying and ranking potentially privileged documents using a machine learning topic model

A method for identifying and ranking potentially privileged documents using a machine learning topic model may include receiving a set of documents. The method may also include, for each of two or more documents in the set of documents, extracting a set of spans from the document, generating, using a machine learning topic model, a set of topics and a subset of legal topics for the set of spans, generating a vector of probabilities for each span with a probability being assigned to each topic in the set of topics for the span, assigning a score to one or more spans in the set of spans by summing the probabilities in the vector that are assigned to a topic in the subset of legal topics, and assigning a score to the document. The method may further include ranking the two or more documents by their assigned scores.
Owner:RELATIVITY ODA LLC

Ensemble of language models for improved user support

Certain aspects of the disclosure provide a method for providing user support by generating recommended response for a customer verbatim with an ensemble of machine learning models. The method includes processing a customer verbatim with a topic model trained to identify a topic associated with the customer verbatim. The method further includes processing the customer verbatim with a sentiment model trained to determine a sentiment of the customer verbatim. The method further includes processing the customer verbatim with an actionability model trained to assign an actionability score to the customer verbatim. The method includes processing the topic, the sentiment, and the actionability score with a recommendation model to generate the recommended response to the customer verbatim.
Owner:INTUIT INC

Short Text Topic Modeling Method Based on Dynamic Clustering and Word Embedding Enhancement

This invention discloses a short text topic modeling method based on dynamic clustering and word embedding enhancement. First, short text data is collected to obtain a short text stream. Then, the short text stream is clustered using the FastStream clustering method, and pseudo-documents are constructed based on the clustering results. Next, a word embedding matrix and a word co-occurrence matrix are formed using a word embedding model pre-trained on a large corpus. Subsequently, the topic distribution of the pseudo-documents and its word distribution are modeled using Dirichlet distribution to obtain a topic model. Then, the topic model is trained on the pseudo-documents using the Gibbs sampling method, updating the topic assignment of each word in the pseudo-documents and the topic distribution parameters of the pseudo-documents until the topic model converges. Finally, new short texts are acquired for topic inference to obtain the topic distribution. This invention utilizes the combination of FastStream clustering and word embedding techniques to effectively improve the topic modeling performance of short text data by creating pseudo-document views and enhancing word embeddings.
Owner:GUANGZHOU UNIVERSITY

Multi-modal co-situation prediction method based on supervision text assistance

The invention discloses a multi-modal co-situation prediction method based on supervision text assistance. The method comprises the following steps: 1, obtaining text, audio and video data and carrying out feature extraction; 2, calculating fused multi-modal features through an attention mechanism and a long-short-term memory network; 3, learning topic distribution of supervised texts by using a hidden Dirichlet topic model LDA; 4, network parameters are trained through a common situation level of a given training scene and theme distribution of a corresponding supervision text; and 5, calculating and predicting the common situation level of the multi-modal scene by using the trained network parameters. According to the method, the multi-modal data and the supervision text are comprehensively utilized as privilege information, so that the model prediction performance in a complex condition-sharing scene is enhanced, the condition-sharing level in a multi-modal scene can be predicted more meticulously, the accuracy and generalization of condition-sharing prediction are remarkably improved, and the mental health support effect is effectively improved.
Owner:UNIV OF SCI & TECH OF CHINA

Text classification method, electronic equipment, storage medium and program product

The embodiment of the invention provides a text classification method, electronic equipment, a storage medium and a program product. The method comprises the following steps: processing a to-be-classified text according to a lexical element division rule to obtain a lexical element sequence; according to the lexical element sequence, adopting a pre-trained unsupervised topic model to obtain a topic distribution data set corresponding to the to-be-classified text; adopting a pre-trained unsupervised clustering model to obtain a clustering distribution data set corresponding to the to-be-classified text; splicing the topic distribution data set and the clustering distribution data set to obtain a text hidden topic; obtaining a plurality of text clusters and current cluster hidden topics corresponding to the text clusters, and calculating the similarity between the text hidden topics and the current cluster hidden topics; and determining a target text cluster from the plurality of text clusters, and adding the to-be-classified text to the target text cluster. Through a cluster classification mechanism of subject distribution and cluster distribution conjoint analysis, the accuracy and result stability of text classification are improved.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD +1

Topic-enhanced and dialog-centric summarization of dialogues

This invention relates to a dialogue summarization method and system based on topic enhancement and dialogue centering. The method includes: extracting user dialogue and user dialogue summaries, labeling them, and constructing a training set; using the training set and an embedded topic model to train a deep learning network model M based on topic enhancement and dialogue centering. Model M obtains the utterance-level representation and utterance-level topic feature representation of the dialogue, and uses these as input. A feature-aware meta-network is used to remove noise, multi-head attention is used to capture the semantic relationships between features, and a gating mechanism is used for filtering and fusion to obtain a topic-enhanced dialogue context feature representation, which is then weighted to obtain the final representation; the final representation is used as input to generate a dialogue summary to train the model, thereby learning the semantic relationship between the user dialogue and the user dialogue summary; the user dialogue is input into the trained model M, and a summary of the user dialogue is output. This method and system are beneficial for improving the accuracy of dialogue summarization.
Owner:FUZHOU UNIV

Systems and methods for utilizing topic models to weight mixture-of-experts for improvement of language modeling

ActiveUS12718016B2EngineeringLanguage modelling
Systems and methods are disclosed for predicting a next text. A method may include receiving one or more documents, such as a document associated with a healthcare provider. The document is then processed to generate one or more tokens which are representative of the document. The document is then processed with a machine-learning model, such as a topic model, and a topic vector is output for the document. Based at least in a part on this topic vector, the document is then processed by one or more expert machine-learning models, which each output a probability vector. The various probability vectors are then further processed to calculate a total probability vector for the document. Based at least in part on the total probability vector for the document, a text output is selected.
Owner:UNITEDHEALTH GROUP INC

Big model-based long text official document key information extraction agent method

The application discloses a long-text official document key information extraction agent method based on a large model, relates to the technical field of artificial intelligence, and comprises the following steps: collecting original official document long-text data; performing structural analysis and hierarchical coding on the original official document text data; performing dynamic semantic segment division based on a topic model guide; constructing a long-text official document key information extraction model based on a bidirectional semantic encoder; performing model training and trainable parameter updating; performing long-text official document key information extraction; and constructing a long-text official document key information extraction agent based on the large model. The application adopts a multi-dimensional structural coding method which fuses official document hierarchies, formats and positions, converts domain prior knowledge into computable vectors, adopts a dynamic planning text segmentation algorithm based on topic consistency and semantic density scoring, guarantees the integrity of long-text semantic segments, and introduces a structure-guided cross-segment attention mechanism in a Transformer encoder, so that precise modeling of long-distance semantic dependence is realized through structural similarity constraints.
Owner:JILIN YOUYUN DIGITAL TECHNOLOGY CO LTD

News stance discrimination method and system based on heterogeneous graph neural network

The application discloses a news stand discrimination method and system based on a heterogeneous graph neural network, and the method comprises the following steps: step 1, using a named entity recognition technology and an LDA topic model to extract entity and topic information in news, and establishing a heterogeneous graph in association with a sentence; step 2, processing the constructed heterogeneous graph through a heterogeneous graph neural network to obtain feature vectors of all nodes in the heterogeneous graph; and step 3, fusing the feature vectors of all nodes output by the heterogeneous graph neural network to comprehensively judge the stand tendency of the news. The application can comprehensively judge the stand tendency of the news by combining important element information in the news and structural relationships between the element information, and has a high discrimination accuracy.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

Analysis Method,Apparatus,And Device For Investment Decision-Making,And Storage Medium

An analysis method for investment decision-making involves, firstly, acquiring news data, invoking a custom-trained topic model to extract entities from the news data, and creating a finite state mach
Owner:MIDAS ANALYTICS LTD

Determining adequacy of documentation using perplexity and probabilistic coherence

Technologies are provided for determining deficiencies in narrative textual data that may impact decision-making in a decisional context. A candidate text document and a reference corpus of text may be utilized to generate one or more topic models and document-term matrices, and then to determine a corresponding statistical perplexity and probabilistic coherence. Statistical determinations of a degree to which the candidate deviates from the reference normative corpus are determined, in terms of the statistical perplexity and probabilistic coherence of the candidate as compared to the reference. If the difference is statistically significant, a message may be reported to user, such as the author or an auditor of the candidate text document, so that the user has the opportunity to amend the candidate document so as to improve its adequacy for the decisional purposes in the context at hand.
Owner:CERNER INNOVATION INC

Measuring and visualizing topic model training convergence

A topic modeling system may include a stability monitor to obtain topic probability distributions for vocabulary items for multiple topics during training iterations of a topic model. For a training iteration and topic, the stability monitor may select a top number of vocabulary elements according to a probability distribution of the topic for the training iteration and a previous training iteration, where the selected vocabulary elements have higher probabilities than vocabulary elements not selected. Then, using a similarity function, top vocabulary elements of the training iteration are compared to top vocabulary elements of the previous training iteration to generate a stability metric indicating an amount of similarity between the probability distributions of the training iteration and the previous training iteration. Additional metrics may be derived and the cumulative metrics may be used to analyze or visualize the convergence or divergence of training of individual topics of the topic model.
Owner:ORACLE INT CORP

Automatic work summary generation method and device, medium and equipment

The invention discloses an automatic work summary generation method and device, a medium and equipment. The method comprises the steps that employee work activity data are collected through a plurality of data sources, and the data sources comprise a computer operation monitoring tool, a task management system and a version control system; preprocessing the collected work activity data, and extracting semantic information by using a natural language processing technology, including word segmentation, named entity recognition, part-of-speech tagging and dependency syntax analysis; based on the extracted semantic information, automatically classifying the work content by adopting a text clustering or topic model algorithm; inputting the classified structured data into a recurrent neural network-based generation model, and automatically generating a work summary text; and outputting the generated work summary, and collecting user feedback to optimize the generation model.
Owner:ANRUI DIGITAL INFORMATION TECH CO LTD

A city functional area identification method based on POI and improved topic model

The application discloses a city functional area identification method based on POI and an improved topic model, and belongs to the technical field of geographic information systems. The city functional area identification method based on POI and the improved topic model comprises the following steps: obtaining interest point data of a target functional area, wherein the interest point data comprises spatial position data of each interest point; dividing the target functional area into a plurality of functional sub-areas according to the spatial position data; determining sub-area semantic features of each functional sub-area; and obtaining regional spatial semantic features of the target functional area according to the sub-area semantic features, so that the problems of low recognition precision and poor accuracy existing in the prior art are solved.
Owner:QINGDAO UNIV OF TECH

Work order data processing method and system based on topic model, medium and equipment

The invention belongs to the technical field of data mining, and provides a work order data processing method and system based on a topic model, a medium and equipment, and the method comprises the steps: extracting a flow type work order data set through employing a trained representation extraction model, and obtaining a second corpus; constructing a second topic model based on the second corpus, and calculating a topic corresponding to the process type work order; counting the quantity of the flow-type work orders associated with each theme, and selecting the theme with the quantity of the associated flow-type work orders ranked as the first sixth preset rank as a second work order hotspot theme; obtaining an original process entry data set, and preprocessing the original process entry data set to obtain a third corpus and a third entry sequence; and based on the third entry sequence, utilizing the second topic model to calculate and obtain a topic corresponding to each process entry. The topic modeling is carried out on the corpus based on the topic model, the hot topic of the work order is intelligently extracted, and the topic is associated with the process entry data, so that management personnel at all levels can carry out optimization work in a targeted manner.
Owner:CHINA TOWER CO LTD

System and method for rapid initialization and transfer of topic models by a multi-stage approach

A system and method for a multi-stage approach for creating topic models is presented. The method includes applying a first stage topic model to textual data, wherein the first stage topic model is trained to discover a first plurality of topics and distributions of words in each topic of the first plurality of topics from the textual data; generating at least one seeded word for a subset of topics of the first plurality of topics, wherein the at least one seeded word is determined based on a plurality of selection rules and the distributions of words in the subset of topics discovered in the first stage topic model; and creating a second stage topic model by feeding the generated at least one seeded word to direct identification of the subset of topics of the first plurality of topics.
Owner:GONG IO INC