Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

51 results about "Topic model" patented technology

In machine learning and natural language processing, a topic model is a type of statistical model for discovering the abstract "topics" that occur in a collection of documents. Topic modeling is a frequently used text-mining tool for discovery of hidden semantic structures in a text body. Intuitively, given that a document is about a particular topic, one would expect particular words to appear in the document more or less frequently: "dog" and "bone" will appear more often in documents about dogs, "cat" and "meow" will appear in documents about cats, and "the" and "is" will appear equally in both. A document typically concerns multiple topics in different proportions; thus, in a document that is 10% about cats and 90% about dogs, there would probably be about 9 times more dog words than cat words. The "topics" produced by topic modeling techniques are clusters of similar words. A topic model captures this intuition in a mathematical framework, which allows examining a set of documents and discovering, based on the statistics of the words in each, what the topics might be and what each document's balance of topics is.

Large language model privacy preservation system

Computer-implemented methods for a large language model privacy preservation system. Aspects include receiving prompt data from a user device. Aspects further include generating pre-processed prompt data using the prompt data from the user device. Aspects also include identifying a category for the pre-processed prompt data using topic modeling. Aspects include generating normalized prompt data using the pre-processed prompt data. Aspects further include storing the category and the normalized prompt data.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Earth and rockfill dam illness feature mining method and system based on historical text data

The invention discloses an earth and rockfill dam illness feature mining method and system based on historical text data, and the method comprises the steps: collecting earth and rockfill dam historical illness data, constructing an earth and rockfill dam illness diagnosis text corpus set, carrying out the structured preprocessing of the corpus set, and obtaining an illness feature and danger removal measure text corpus set; generating a structured word sequence subset through a word segmentation tool; processing the structured word sequence subset by adopting an LDA topic model, determining an optimal topic number through a confusion degree curve, and outputting a final topic and a corresponding topic word; constructing a visual network graph based on the subject term co-occurrence frequency; and carrying out centrality analysis on nodes in the visual network diagram, identifying key nodes in the visual network diagram, quantitatively analyzing association rules of the danger characteristics and danger removing measures, and completing feature mining of the earth and rockfill dam danger. The problems that traditional manual diagnosis is high in subjectivity and low in utilization rate of historical engineering data are solved, and intelligent auxiliary decision making of the earth and rockfill dam danger is achieved.
Owner:NANJING HYDRAULIC RES INST

Trending topic discovery with keyword-based topic model

The present disclosure relates to a system, a method, and a product for topic discovery. The system includes a memory storing instructions; and a processor in communication with the memory. When the processor executes the instructions, the instructions are configured to cause the processor to: obtain text data, conduct pre-processing on the text data to obtain pre-processed text data, extract an entity list and a keyword list based on the pre-processed text data, generate an entity embedding list based on the entity list, clusterize the entity list based on the entity embedding list to obtain a plurality of entity clusters, each entity cluster comprising at least one entity, retrieve a co-occurring keyword list based on the plurality of entity clusters, the entity list, and the keyword list, and obtain a topic for each entity cluster of the plurality of entity clusters based on the co-occurring keyword list.
Owner:ACCENTURE GLOBAL SOLUTIONS LTD

Traffic transportation government affair hotline mining method based on natural language processing

The invention discloses a traffic transportation government affair hotline mining method based on natural language processing, and belongs to the technical field of artificial intelligence and machine learning. According to the method, for structured and unstructured data in a traffic transportation government affair hotline, firstly, a five-class word segmentation dictionary containing cleaning words, noise words, synonyms, additive words and stop words is constructed; converting the unstructured text into structured data through improved data cleaning, text word segmentation and feature representation; carrying out clustering analysis by utilizing an LDA topic model, and extracting public demand topics and high-frequency keywords; hot appeals and trend changes are mined in combination with space-time analysis and association analysis, and finally a visual analysis report is generated. According to the method, the problems of poor model interpretability and insufficient field adaptability in the prior art are solved, the accuracy of hotline data processing and the effectiveness of theme recognition are remarkably improved through multi-dictionary collaborative optimization and field knowledge fusion, and accurate decision support is provided for a traffic transportation management department. A real taxi field case in a certain city is used as an example for research, an experiment proves that the method has an accurate theme identification function, and the complaint and report work order amount in the traffic transportation government affair hotline taxi field in the city is reduced by 20% on year-on-year basis in 2024.
Owner:乌若愚

Updating support documentation for developer platforms with topic clustering of feedback

A method includes obtaining at least one feedback response from a developer regarding an answer and corresponding source documents of the answer in a discussion thread initiated by the developer. The method further includes converting the discussion thread into a topic model clustering input, responsive to the feedback response specifying an unsatisfactory category of feedback responses, to obtain a multitude of topic clustering model inputs. The method further includes periodically processing the multitude of topic clustering model inputs by a thread-topic clustering model to obtain a multitude of candidate topics. The method further includes processing, by an answer generation model, a first candidate topic of the multitude of candidate topics to obtain a multitude of corresponding documentation recommendations for the candidate topic. The method further includes presenting the first candidate topic and the multitude of corresponding documentation recommendations.
Owner:INTUIT INC

A topic model updating method and system, a storage medium and a server

The embodiment of the application discloses a kind of theme model updating method, system and storage medium and server, apply to the information processing technical field based on artificial intelligence.Theme model system will obtain the first label semantic feature and the second label semantic feature corresponding respectively to multiple old theme labels in the first theme model and multiple new theme models in the second theme model, and based on the first label semantic feature and the second label semantic label, mapping relationship is established between old theme label and new theme label, and then the old theme label in the first theme model is updated based on the mapping relationship.The updating of the first theme model existing in the system is realized automatically, the efficiency of the theme model is improved, and the theme model with larger dimensionality can also be updated, and the updating of the first theme model is not limited by the theme model acquisition method.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Text style recognition method and device, equipment and medium

The invention relates to the technical field of style recognition, and discloses a text style recognition method. The method comprises the steps of obtaining a to-be-recognized text, and determining a text publishing platform corresponding to the to-be-recognized text; performing topic recognition on the to-be-recognized text through an LDA topic model corresponding to the text publishing platform to obtain at least one text topic and topic keywords corresponding to the text topics; performing topic clustering on all the text topics and the topic keywords corresponding to the text topics through a visual clustering tool to obtain a text topic cluster; sending the text theme cluster to the client, and receiving an optimal theme quantity fed back by the client; and performing style mapping on the target theme in the optimal theme quantity to obtain a text style. According to the method, through the LDA topic model and the style mapping rule, the topic and the style in the text are recognized, the accuracy of style recognition is improved, and the efficiency of style recognition is improved.
Owner:SHENZHEN DONSON CLOUD TECHNOLOGY CO LTD

A Spatial Gene Identification and Extraction Method Based on Social Media Text Data

This invention discloses a method for spatial gene identification and extraction based on social media text data, comprising the following steps: collecting online text data about a city, then preprocessing the data to obtain dataset D1; constructing a dictionary and vector space in analysis software, introducing an LDA topic model, and classifying the obtained dataset D1 into topics; merging synonyms in each topic, and performing synonym replacement in dataset D1 to obtain dataset D2; counting the co-occurrence frequency of keywords in dataset D2 and constructing a co-occurrence matrix M; and using a hierarchical clustering model to cluster the semantic network analysis results to obtain spatial combination patterns, i.e., spatial genes. This invention collects online text data about a specific city from multiple social media platforms, providing a practical technical means for urban researchers to identify urban spatial genes by obtaining rich, non-invasive data.
Owner:SOUTHEAST UNIV +1

Topic model generation method and device, storage medium and program product

The embodiment of the invention provides a topic model generation method and device, a storage medium and a program product, and relates to the technical field of artificial intelligence. The method comprises the following steps: preprocessing an obtained user comment text and comment metadata, and correspondingly obtaining a word frequency vector and a text label; inputting the word frequency vector into a preset topic model, converting the word frequency vector into potential space representation, and performing normalization processing to obtain topic distribution corresponding to a user comment text output by the topic model; inputting the user comment text into a preset attention mechanism model to obtain a text embedding vector corresponding to the user comment text; and based on the topic distribution, the text embedding vector and the text label, carrying out joint training, and determining a trained target model. The attention mechanism and the comment metadata are combined in the modeling process, the trained target model is determined, and the problems that in the prior art, a topic model is insufficient in real scene data processing capacity and low in efficiency are solved.
Owner:CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1

Long text abstract generation method based on hierarchical graph comparison theme

PendingCN122021560ASemantic analysisText processingDocument representationInformation coverage
The invention discloses a long text abstract generation method based on hierarchical graph comparison themes, which comprises the following steps of: 1, preprocessing an original document, dividing sentence sequences, and obtaining global context-aware sentence and document representation through a hierarchical encoder network; 2, deducing document-level and sentence-level topic distribution by using a neural topic model; and 3, constructing a supervision graph based on a standard abstract to perform graph comparison learning so as to close the topic representation of a document and a key sentence and push redundant information. According to the method, the deep semantic structure of the long document can be effectively captured, so that the theme consistency and the information coverage degree of the abstract can be improved, and the redundancy is reduced.
Owner:ANHUI AGRICULTURAL UNIVERSITY

Multimodal context selection for large language model based resolutions addressing technical issues

A method for technical issue resolution. The method includes: receiving, from a user, a text query concerning a technical issue; obtaining query-related context relevant to the text query; and processing, through a large language model (LLM), the text query and the query-related context to produce a multimodal query response used by the user to address the technical issue. More specifically, embodiments described herein utilize text topic and zero shot classification models to translate multimodal technical documentation (e.g., including text and images) into topic relevant metadata; and process queries, pertaining to technical issues, using a multimodal LLM provided with query-related text and image context derived from said topic relevant metadata.
Owner:DELL PROD LP

Identifying and ranking potentially privileged documents using a machine learning topic model

A method for identifying and ranking potentially privileged documents using a machine learning topic model may include receiving a set of documents. The method may also include, for each of two or more documents in the set of documents, extracting a set of spans from the document, generating, using a machine learning topic model, a set of topics and a subset of legal topics for the set of spans, generating a vector of probabilities for each span with a probability being assigned to each topic in the set of topics for the span, assigning a score to one or more spans in the set of spans by summing the probabilities in the vector that are assigned to a topic in the subset of legal topics, and assigning a score to the document. The method may further include ranking the two or more documents by their assigned scores.
Owner:RELATIVITY ODA LLC

Ensemble of language models for improved user support

Certain aspects of the disclosure provide a method for providing user support by generating recommended response for a customer verbatim with an ensemble of machine learning models. The method includes processing a customer verbatim with a topic model trained to identify a topic associated with the customer verbatim. The method further includes processing the customer verbatim with a sentiment model trained to determine a sentiment of the customer verbatim. The method further includes processing the customer verbatim with an actionability model trained to assign an actionability score to the customer verbatim. The method includes processing the topic, the sentiment, and the actionability score with a recommendation model to generate the recommended response to the customer verbatim.
Owner:INTUIT INC

Short Text Topic Modeling Method Based on Dynamic Clustering and Word Embedding Enhancement

This invention discloses a short text topic modeling method based on dynamic clustering and word embedding enhancement. First, short text data is collected to obtain a short text stream. Then, the short text stream is clustered using the FastStream clustering method, and pseudo-documents are constructed based on the clustering results. Next, a word embedding matrix and a word co-occurrence matrix are formed using a word embedding model pre-trained on a large corpus. Subsequently, the topic distribution of the pseudo-documents and its word distribution are modeled using Dirichlet distribution to obtain a topic model. Then, the topic model is trained on the pseudo-documents using the Gibbs sampling method, updating the topic assignment of each word in the pseudo-documents and the topic distribution parameters of the pseudo-documents until the topic model converges. Finally, new short texts are acquired for topic inference to obtain the topic distribution. This invention utilizes the combination of FastStream clustering and word embedding techniques to effectively improve the topic modeling performance of short text data by creating pseudo-document views and enhancing word embeddings.
Owner:GUANGZHOU UNIVERSITY

Multi-modal co-situation prediction method based on supervision text assistance

The invention discloses a multi-modal co-situation prediction method based on supervision text assistance. The method comprises the following steps: 1, obtaining text, audio and video data and carrying out feature extraction; 2, calculating fused multi-modal features through an attention mechanism and a long-short-term memory network; 3, learning topic distribution of supervised texts by using a hidden Dirichlet topic model LDA; 4, network parameters are trained through a common situation level of a given training scene and theme distribution of a corresponding supervision text; and 5, calculating and predicting the common situation level of the multi-modal scene by using the trained network parameters. According to the method, the multi-modal data and the supervision text are comprehensively utilized as privilege information, so that the model prediction performance in a complex condition-sharing scene is enhanced, the condition-sharing level in a multi-modal scene can be predicted more meticulously, the accuracy and generalization of condition-sharing prediction are remarkably improved, and the mental health support effect is effectively improved.
Owner:UNIV OF SCI & TECH OF CHINA

Text classification method, electronic equipment, storage medium and program product

The embodiment of the invention provides a text classification method, electronic equipment, a storage medium and a program product. The method comprises the following steps: processing a to-be-classified text according to a lexical element division rule to obtain a lexical element sequence; according to the lexical element sequence, adopting a pre-trained unsupervised topic model to obtain a topic distribution data set corresponding to the to-be-classified text; adopting a pre-trained unsupervised clustering model to obtain a clustering distribution data set corresponding to the to-be-classified text; splicing the topic distribution data set and the clustering distribution data set to obtain a text hidden topic; obtaining a plurality of text clusters and current cluster hidden topics corresponding to the text clusters, and calculating the similarity between the text hidden topics and the current cluster hidden topics; and determining a target text cluster from the plurality of text clusters, and adding the to-be-classified text to the target text cluster. Through a cluster classification mechanism of subject distribution and cluster distribution conjoint analysis, the accuracy and result stability of text classification are improved.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD +1

Topic-enhanced and dialog-centric summarization of dialogues

This invention relates to a dialogue summarization method and system based on topic enhancement and dialogue centering. The method includes: extracting user dialogue and user dialogue summaries, labeling them, and constructing a training set; using the training set and an embedded topic model to train a deep learning network model M based on topic enhancement and dialogue centering. Model M obtains the utterance-level representation and utterance-level topic feature representation of the dialogue, and uses these as input. A feature-aware meta-network is used to remove noise, multi-head attention is used to capture the semantic relationships between features, and a gating mechanism is used for filtering and fusion to obtain a topic-enhanced dialogue context feature representation, which is then weighted to obtain the final representation; the final representation is used as input to generate a dialogue summary to train the model, thereby learning the semantic relationship between the user dialogue and the user dialogue summary; the user dialogue is input into the trained model M, and a summary of the user dialogue is output. This method and system are beneficial for improving the accuracy of dialogue summarization.
Owner:FUZHOU UNIV

Systems and methods for utilizing topic models to weight mixture-of-experts for improvement of language modeling

ActiveUS12718016B2EngineeringLanguage modelling
Systems and methods are disclosed for predicting a next text. A method may include receiving one or more documents, such as a document associated with a healthcare provider. The document is then processed to generate one or more tokens which are representative of the document. The document is then processed with a machine-learning model, such as a topic model, and a topic vector is output for the document. Based at least in a part on this topic vector, the document is then processed by one or more expert machine-learning models, which each output a probability vector. The various probability vectors are then further processed to calculate a total probability vector for the document. Based at least in part on the total probability vector for the document, a text output is selected.
Owner:UNITEDHEALTH GROUP INC

Big model-based long text official document key information extraction agent method

The application discloses a long-text official document key information extraction agent method based on a large model, relates to the technical field of artificial intelligence, and comprises the following steps: collecting original official document long-text data; performing structural analysis and hierarchical coding on the original official document text data; performing dynamic semantic segment division based on a topic model guide; constructing a long-text official document key information extraction model based on a bidirectional semantic encoder; performing model training and trainable parameter updating; performing long-text official document key information extraction; and constructing a long-text official document key information extraction agent based on the large model. The application adopts a multi-dimensional structural coding method which fuses official document hierarchies, formats and positions, converts domain prior knowledge into computable vectors, adopts a dynamic planning text segmentation algorithm based on topic consistency and semantic density scoring, guarantees the integrity of long-text semantic segments, and introduces a structure-guided cross-segment attention mechanism in a Transformer encoder, so that precise modeling of long-distance semantic dependence is realized through structural similarity constraints.
Owner:JILIN YOUYUN DIGITAL TECHNOLOGY CO LTD

News stance discrimination method and system based on heterogeneous graph neural network

The application discloses a news stand discrimination method and system based on a heterogeneous graph neural network, and the method comprises the following steps: step 1, using a named entity recognition technology and an LDA topic model to extract entity and topic information in news, and establishing a heterogeneous graph in association with a sentence; step 2, processing the constructed heterogeneous graph through a heterogeneous graph neural network to obtain feature vectors of all nodes in the heterogeneous graph; and step 3, fusing the feature vectors of all nodes output by the heterogeneous graph neural network to comprehensively judge the stand tendency of the news. The application can comprehensively judge the stand tendency of the news by combining important element information in the news and structural relationships between the element information, and has a high discrimination accuracy.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

Analysis Method,Apparatus,And Device For Investment Decision-Making,And Storage Medium

An analysis method for investment decision-making involves, firstly, acquiring news data, invoking a custom-trained topic model to extract entities from the news data, and creating a finite state mach
Owner:MIDAS ANALYTICS LTD

Determining adequacy of documentation using perplexity and probabilistic coherence

Technologies are provided for determining deficiencies in narrative textual data that may impact decision-making in a decisional context. A candidate text document and a reference corpus of text may be utilized to generate one or more topic models and document-term matrices, and then to determine a corresponding statistical perplexity and probabilistic coherence. Statistical determinations of a degree to which the candidate deviates from the reference normative corpus are determined, in terms of the statistical perplexity and probabilistic coherence of the candidate as compared to the reference. If the difference is statistically significant, a message may be reported to user, such as the author or an auditor of the candidate text document, so that the user has the opportunity to amend the candidate document so as to improve its adequacy for the decisional purposes in the context at hand.
Owner:CERNER INNOVATION INC

Automatic work summary generation method and device, medium and equipment

The invention discloses an automatic work summary generation method and device, a medium and equipment. The method comprises the steps that employee work activity data are collected through a plurality of data sources, and the data sources comprise a computer operation monitoring tool, a task management system and a version control system; preprocessing the collected work activity data, and extracting semantic information by using a natural language processing technology, including word segmentation, named entity recognition, part-of-speech tagging and dependency syntax analysis; based on the extracted semantic information, automatically classifying the work content by adopting a text clustering or topic model algorithm; inputting the classified structured data into a recurrent neural network-based generation model, and automatically generating a work summary text; and outputting the generated work summary, and collecting user feedback to optimize the generation model.
Owner:ANRUI DIGITAL INFORMATION TECH CO LTD

A city functional area identification method based on POI and improved topic model

The application discloses a city functional area identification method based on POI and an improved topic model, and belongs to the technical field of geographic information systems. The city functional area identification method based on POI and the improved topic model comprises the following steps: obtaining interest point data of a target functional area, wherein the interest point data comprises spatial position data of each interest point; dividing the target functional area into a plurality of functional sub-areas according to the spatial position data; determining sub-area semantic features of each functional sub-area; and obtaining regional spatial semantic features of the target functional area according to the sub-area semantic features, so that the problems of low recognition precision and poor accuracy existing in the prior art are solved.
Owner:QINGDAO UNIV OF TECH

Conference key information real-time extraction and knowledge pushing method based on large language model

The invention provides a conference key information real-time extraction and knowledge pushing method based on a large language model, and relates to the technical field of artificial intelligence and conference management, and the method comprises the steps: obtaining a voice data stream in real time, carrying out the transcription and semantic analysis, and dividing semantic segments based on semantic integrity constraints; identifying key information by using the information extraction template and the conference theme model; constructing a conference discussion evolution knowledge graph; and carrying out accurate knowledge pushing according to the map and the conference process. The conference efficiency can be improved, knowledge sharing is enhanced, and the decision process is optimized.
Owner:BEIJING YIZHUANG INTELLIGENT CITY RES INST GRP CO LTD

Method, apparatus, device and storage medium for constructing dataset

The disclosure provides a method, device and equipment for constructing a dataset, and a storage medium, relates to the technical field of computers, in particular to the fields of natural language processing, cloud computing, deep learning and the like. The specific implementation scheme is as follows: determining a target topic according to a topic-word distribution matrix output by a topic model based on a text set; determining a target text according to a text-topic distribution matrix output by the topic model based on the text set by using the target topic, wherein the target text is derived from the text set; and constructing a dataset based on the target text. According to the scheme of the disclosure, the target text required for constructing the dataset can be quickly and accurately screened from the text set.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Intention recognition method and device based on domain knowledge seed words and computer device

This application relates to an intent recognition method, apparatus, and computer device based on domain knowledge seed words. The method includes: acquiring and preprocessing a planning document corpus; calculating term importance weights to filter initial seed words; dividing the corpus into hierarchical seed words using a large language model; expanding the corpus using a dual-dimensional approach of general vocabulary similarity and domain context embedding similarity; filtering the corpus with a domain dictionary to obtain a hierarchical expanded seed word dictionary; increasing the Dirichlet prior weight of the corresponding level's expanded seed words in the topic model based on hierarchical matching relationships according to the expanded word dictionary corresponding to each level; performing topic modeling on the planning document corpus; outputting the vocabulary distribution of each topic and the topic distribution of each document; constructing a clustering tree based on topic semantic distance; filtering the optimal partition; and generating hierarchical intent results by combining the hierarchical structure. This method can effectively achieve automated and hierarchical intent recognition, ensuring that the recognition results align with domain business logic.
Owner:NAT UNIV OF DEFENSE TECH

Marketing topic key user based on multi-feature fusion and influence measurement method thereof

The invention relates to a marketing topic key user and influence measurement method based on multi-feature fusion, and belongs to the technical field of social network analysis. According to the method, modeling is performed on text data through an LDA topic model, and meanwhile, a marketing topic set is extracted by utilizing a clustering algorithm. The topic correlation is used as an index for measuring the correlation degree between the text and the topic, and the quality and correlation of the text and the topic are effectively evaluated. Secondly, introducing a fuzzy mathematical theory, and establishing a fuzzy comprehensive evaluation model based on information entropy; various characteristics of the user are fused as evaluation indexes, and the weight is determined by using an entropy weight method, so that the interference of subjective preference is reduced. And finally, constructing a multi-dimensional heterogeneous network based on marketing topics. By fusing the attribute characteristics of each node and utilizing the random walk strategy, the user behavior and interaction information are comprehensively considered, so that the influence of the user in the marketing topic is more comprehensively evaluated.
Owner:CHONGQING UNIV OF POSTS & TELECOMM +1

A method for extracting intelligence tags using heterogeneous BERT and semi-supervised SVM models

This invention discloses a method for extracting intelligence tags using heterogeneous BERT and semi-supervised SVM models, comprising: acquiring the original unstructured intelligence text dataset and performing text preprocessing and data augmentation operations on it; extracting topic words from the intelligence text based on BTM short text and LDA long text topic models to obtain topic word vectors of the intelligence text; representing the intelligence document with sentence vectors based on the Doc2Vec model; inputting the topic word vectors and sentence vectors into the BERT model for pre-training to obtain feature vectors with topic information; training and parameter tuning the semi-supervised SVM model using a small amount of labeled data and a large amount of unlabeled data; and inputting the feature vectors with topic information into the semi-supervised SVM binary classification algorithm model for training to obtain the data tag extraction results. This invention can effectively extract data tags from intelligence text, make full use of unlabeled data information, and improve the accuracy and generalization ability of the data tag extraction system.
Owner:THE QUARTERMASTER RES INST OF THE GENERAL LOGISTICS DEPT OF THE CPLA

Intention recognition method and device based on domain knowledge seed words and computer equipment

The invention relates to an intention recognition method and device based on field knowledge seed words and computer equipment. The method comprises the steps that a planning document corpus is obtained and preprocessed, term importance weights are calculated to screen initial seed words, the seed words of all levels are obtained through hierarchical granularity division of a large language model, and two-dimensional expansion of general vocabulary similarity and domain context embedding similarity is adopted to obtain a multi-level seed word; filtering through a domain dictionary to obtain a hierarchical extension seed word dictionary, improving the Dirichlet prior weight of the extension seed word of the corresponding hierarchy in the topic model according to the extension word dictionary corresponding to each hierarchy and a hierarchy matching relationship, carrying out topic modeling on the planning document corpus, and outputting the vocabulary distribution of each topic and the topic distribution of each document. And constructing a clustering tree based on the topic semantic distance, screening the optimal division, and then combining with the hierarchical structure to generate a hierarchical intention result. By adopting the method, the automatic and hierarchical intention recognition can be effectively realized, and the recognition result fits the domain business logic.
Owner:NAT UNIV OF DEFENSE TECH