Method and system for providing rare topic detection using hierarchical clustering

Through hierarchical topic modeling technology, the problem that existing technology is difficult to understand complex text data and real-time communication messages is solved, and the rapid identification and intelligent interpretation of rare topics are achieved, and the efficiency of information retrieval and text analysis is improved.

CN114424197BActive Publication Date: 2025-05-13INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202080066389.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-08
Filing Date
2020-09-29
Publication Date
2025-05-13
Estimated Expiration
2040-09-29

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and adaptively understand topics in communications/dialogues, providing intelligent explanations, overviews and understandings, especially in large-scale text-based data, especially when it comes to complex information access systems and real-time communication messages.

Method used

By using hierarchical topic modeling, hierarchical topic models can be learned from data sources, iteratively removes dominant words to identify rare topics, and provides multi-level generalization and interpretability by seeding hierarchical topic models.

Benefits of technology

It realizes a deep understanding of complex text data, can quickly identify rare topics, provide intelligent interpretation and multi-level generalization, and improves the efficiency of information retrieval and text analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114424197B_ABST
    Figure CN114424197B_ABST
Patent Text Reader

Abstract

A hierarchical topic model can be learned from one or more data sources. The hierarchical topic model can be used to iteratively remove one or more dominant words in a selected cluster. The dominant words can relate to one or more main topics of the cluster. The learned hierarchical topic model can be seeded with one or more words, n-grams, phrases, text fragments, or a combination thereof to evolve the hierarchical topic model, wherein when the seeding is complete, the removed domain words are restored.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates generally to computing systems and, more particularly, to various embodiments for providing rare topic detection using hierarchical clustering utilizing computing processors. Background Art

[0002] The advent of computers and network technology has made it possible to improve the quality of life while enhancing daily activities and simplifying information sharing. Due to recent developments in information technology and the increasing popularity of the Internet, a large amount of information is now available in digital form. The availability of this information provides many opportunities. In recent years, digital information and online information such as real-time communication messaging have become very popular. As the rapid progress of technology has achieved results, the need to make progress in these systems that are conducive to efficiency and improvement is greater. Summary of the invention

[0003] Embodiments are provided for providing rare topic detection using hierarchical topic modeling by a processor. The hierarchical topic model can be learned from one or more data sources. The hierarchical topic model can be used to iteratively remove one or more dominant words in a selected cluster. The dominant words can relate to one or more main topics of the cluster. The learned hierarchical topic model can be seeded with one or more words, n-grams, phrases, text snippets, or a combination thereof to evolve the hierarchical topic model, and when the seeding is complete, the removed domain words are restored. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] In order to easily understand the advantages of the present invention, a more particular description of the present invention briefly described above will be presented by reference to specific embodiments shown in the accompanying drawings. It should be understood that these drawings depict only typical embodiments of the present invention and are not therefore to be considered limiting of its scope. The present invention will be described and explained with additional specificity and detail through the use of the accompanying drawings, in which:

[0005] Figure 1 is a block diagram illustrating an exemplary cloud computing node according to an embodiment of the present invention;

[0006] Figure 2 is an additional block diagram depicting an exemplary cloud computing environment according to an embodiment of the present invention;

[0007] Figure 3 is an additional block diagram depicting the abstract model layers according to an embodiment of the present invention;

[0008] Figure 4 are additional graphs depicting inter-arrival times between analyzing real-time conversation data and recording messages in accordance with aspects of the present invention;

[0009] Figure 5 is a diagram depicting rare topic detection using hierarchical topic modeling in accordance with aspects of the present invention; and

[0010] Figure 6 is a flow chart depicting an exemplary method for providing rare topic detection using hierarchical topic modeling by a processor; again, in which aspects of the present invention may be implemented. DETAILED DESCRIPTION

[0011] As the amount of electronic information continues to increase, the demand for complex information access systems has also grown. Digital or "online" data has become increasingly accessible through real-time, global computer networks. The data can reflect many aspects of different organizations and groups or individuals, including science, politics, government, education, business, etc. As the use of collaborative and social communications increases, communications via text-based communications will also increase. For business and entertainment purposes, real-time communication messages (e.g., real-time chat sessions) are an important part of modern society. However, for each entity, regardless of size, using such collaborative and social communication methods may be an overwhelming experience, especially when a large amount of text-based data is generated by each application and service.

[0012] In addition, entities of various types (e.g., businesses, organizations, government agencies, educational institutions, etc.) often engage in corpus linguistics, which is the study of language expressed in a corpus (i.e., collection) of "real-use" texts. The core idea of ​​corpus linguistics is that analysis of expressions is best done within their natural use. By collecting writing samples, researchers are able to understand how individuals talk to each other. As such, the present invention employs different techniques that help understand and interpret message-based data.

[0013] In one aspect, topic modeling can be used to discover semantic structure within a corpus of text. Topic modeling can employ one or more operations to infer themes and meanings in text-based documents and / or conversations. Topic modeling and text mining can be used to gain insights into different communications. For example, if a business can mine customer feedback about a particular product or service, this information can prove valuable. When employing text mining / topic modeling techniques, one of the recommendations is that the more data available for analysis, the better the overall results. However, even with big data, practitioners may need to text mine a single conversation or a small text corpus to infer meaning.

[0014] Additionally, during a communication (e.g., a conversation between one or more users that may be in text form (e.g., a document, email, presentation, etc.) and / or audio / video form), it is necessary to quickly and adaptively understand the communication / conversation while providing intelligent interpretation, summary, and / or understanding related to the subject matter of such communication / conversation.

[0015] In some cases, for example, document clustering is to group similar documents together, thereby assigning them to the same implicit topic. Document clustering provides the ability to improve the effectiveness of information retrieval. Recently, latent semantic analysis operations and clustering hierarchical clustering have been adopted to group objects into clusters based on similarity. For example, latent semantic analysis, where n sentences are given, the framework lists the concepts cited in those sentences. That is, the theme is a "bag of words", where each document has multiple topics (with multinomial distribution) and each theme has multiple words (with Dirichlet distribution). However, the challenge of latent semantic analysis is that the communication / dialogue (e.g., dialogue / spoken English) words in the theme cannot satisfy the Dirichlet generation process and do not have the concept of hierarchical topics (e.g., data is a class of data plans and the data plan is a class of international data plans).

[0016] In a hierarchical clustering operation, documents are recursively merged from bottom to top, resulting in a decision tree of recursively partitioned clusters. The distance measures used to find similarities vary from single links to computationally more expensive links, but they are closely related to the nearest neighbor distance. The hierarchical clustering operation works by recursively merging a single best document or cluster pair, making the computational cost too high for document collections numbering in the tens of thousands. That is, documents are represented as vectors with distances (e.g., Euclidean) between them. However, when the "dominant" words are not removed from the vectors of the lower levels (e.g., the data is mainly at the highest level and 30% of the conversations occur, and "international" only occurs in 1%), the distance metric fails. As a result, there is still a challenge to provide an overview of the communication / conversation corpus to the topic (compared to just the documents).

[0017] Thus, various embodiments are shown herein to provide rare topic detection using hierarchical topic modeling by a processor. A hierarchical topic model can be learned from one or more data sources. The hierarchical topic model can be used to iteratively remove one or more dominant words in a selected cluster. The dominant words can relate to one or more main topics of the cluster. The learned hierarchical topic model can be seeded with one or more words, n-grams, phrases, text fragments, or a combination thereof to evolve the hierarchical topic model, and when the seeding is complete, the removed domain words are restored.

[0018] In one aspect, the present invention provides hierarchical topic modeling by providing a summarized version of a call (e.g., a speech-to-text transcription of a customer-agent interaction) clustered into multiple topics. That is, hierarchical topic modeling works on any type of text document, and long text documents can be converted into summaries, which are typically a collection of ngrams.

[0019] The overview of ngram words can be used to generate word vectors, and the word vectors can be weighted according to one or more assigned scores. A K-means clustering operation can be used in each iteration when aggregating the word vectors into K clusters, where "K" is a positive integer or a defined value. The K clusters may include one or more "king clusters". In one aspect, the king cluster is the largest cluster (e.g., the cluster containing the most documents or data sources) from a total of K clusters. The king cluster can be the largest cluster within multiple clusters.

[0020] For each cluster that is a king cluster, the hierarchical topic modeling operation is repeatedly performed by removing one or more "related" words from the previous run / execution (which no longer makes a difference for the next hierarchical topic modeling). In doing so, as the dominant words are removed, one or more rare topics are identified through a progressive drill-down operation (e.g., from iteratively performing the hierarchical topic modeling operation). Ngrams, fragments, and suggested topic names for each representative cluster can be identified. The one or more words that are removed / suppressed can be used for ngram / fragment identification to improve and provide enhanced readability / interpretability for one or more users.

[0021] For example, consider a hierarchical topic modeling operation that removes the word "access" in a first iteration (e.g., iteration "0"). In the next / subsequent iterative hierarchical topic modeling operation, the words "vpn" and "root" may be removed in one or more subsequent iterative hierarchical topic modeling operations (e.g., iteration "1" and / or iteration "N"). At the end of the iterative hierarchical topic modeling operation, the dominant words may be restored / unsuppressed while providing an interpretable explanation (e.g., user understandable) using one or more artificial intelligence ("AI") operations (such as, for example, "cannot access vpn" and / or "root access failed"). In addition, the present invention provides automatic configuration for iterative hierarchical topic modeling, such as, for example, configurable to select multiple iterations, synonyms to identify "similar" clusters. The operation for providing rare topic detection using hierarchical clustering also enables post-processing to combine or split one or more clusters, each of which can be understood / interpreted by one or more users.

[0022] In one aspect, one or more hierarchical topic models for incremental training and identification of differences can be learned. An existing hierarchical topic model (e.g., an existing tree structure) can be used to seed the learned hierarchical topic model (e.g., a new tree structure). Each clustering model in each tree node can be seeded based on the existing hierarchical topic model. It should be noted that the hierarchical topic model is in the form of a tree structure, in which each node represents a topic. The nodes corresponding to the king cluster are decomposed in each iteration. Incremental training represents a process in which the training process starts with an old model and then finds the best model with a new data set, rather than training the topic model from scratch. The learned existing hierarchical topic model can be retrained on a new data set, and the optimal solution is gradually explored for the clustering problem in the neighborhood of the previous solution. For further explanation, consider a topic model "v1" (e.g., an existing topic model) trained on data set 1 and a topic model "v2" (e.g., a new topic model trained on data set 2 with topic model v1 as a seed model). Data set 2 is a new data set. On data set 2, compared to learning the topic model from scratch, the present invention finds and / or identifies the best topic model close to topic model v1. Use the old topic model v1 to seed the base K-means clustering to get the new topic model v2. The seed model is the topic model trained for a specific time window, and the new model is trained with the new dataset on the next time window.

[0023] In one aspect, one or more hierarchical topic models can be used to identify / detect changes in clusters by using (a) cluster centers that have drifted the most to be identified as significant change candidates, (b) cluster weights that have significant differences, (c) cluster cohesion measures that have changed significantly, and (d) tree structures that have changed. That is, "change detection" refers to how a newly trained topic model is changed relative to a seed model and the change can be observed as described in (a)-(d).

[0024] It is understood in advance that although the present disclosure includes detailed descriptions about cloud computing, the implementation of the teachings cited herein is not limited to a cloud computing environment. Instead, embodiments of the present invention can be implemented in conjunction with any other type of computing environment now known or later developed.

[0025] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be quickly provisioned and released with minimal management effort or interaction with the provider of the service. The cloud model may include at least five characteristics, at least three service models, and at least four deployment models.

[0026] Features are as follows:

[0027] On-demand self-service: Cloud consumers can unilaterally and automatically provision computing capabilities, such as server time and network storage, as needed, without requiring human interaction with the provider of the service.

[0028] Broad network access: Capabilities are available over the network and accessed through standard mechanisms that facilitate the use of heterogeneous thin-client platforms or thick-client platforms (e.g., mobile phones, laptops, and PDAs).

[0029] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically assigned and reassigned as needed. There is a sense of location independence, as consumers typically do not have control or knowledge of the exact location of the provided resources, but may be able to specify the location at a higher level of abstraction (e.g., country, state, or data center).

[0030] Rapid elasticity: The ability to quickly and elastically provision capacity, in some cases automatically scaling down quickly and releasing quickly to scale up quickly. To the consumer, the capacity available for provisioning generally appears unlimited and can be purchased in any quantity at any time.

[0031] Measured services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of the utilized services.

[0032] The service model is as follows:

[0033] Software as a Service (SaaS): The capability provided to the consumer is to use the provider's applications running on a cloud infrastructure. The applications can be accessed from different client devices through a thin client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure including the network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.

[0034] Platform as a Service (PaaS): The capability provided to consumers is to deploy applications created or acquired by consumers using programming languages ​​and tools supported by the provider onto the cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure including networks, servers, operating systems or storage, but have control over the deployed applications and possible configuration of the application hosting environment.

[0035] Infrastructure as a Service (IaaS): The capabilities provided to consumers are the provision of processing, storage, networking, and other basic computing resources on which consumers can deploy and run arbitrary software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but have control over the operating system, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).

[0036] The deployment model is as follows:

[0037] Private Cloud: Cloud infrastructure is operated only for the organization. It can be managed by the organization or a third party and can exist on-premises or off-premises.

[0038] Community Cloud: Cloud infrastructure is shared by several organizations and supports a specific community with shared concerns (e.g., mission, security requirements, policies, and compliance considerations). It can be managed by the organization or a third party and can exist on-premises or off-premises.

[0039] Public cloud: Cloud infrastructure is made available to the public or large industry groups and is owned by an organization that sells cloud services.

[0040] Hybrid cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain unique entities but are bound together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).

[0041] Cloud computing environments are service-oriented and focus on statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is the infrastructure that includes a network of interconnected nodes.

[0042] See now Figure 1 , a schematic diagram of an example of a cloud computing node is shown. Cloud computing node 10 is merely one example of a suitable cloud computing node and is not intended to impose any limitation on the scope of use or functionality of the embodiments of the invention described herein. Regardless, cloud computing node 10 can be implemented and / or perform any of the functions set forth above.

[0043] In the cloud computing node 10, there is a computer system / server 12, which can operate with many other general or special computing system environments or configurations. Examples of well-known computing systems, environments and / or configurations that can be suitable for the computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, small computer systems, large computer systems, and distributed cloud computing environments including any of the above systems or devices, etc.

[0044] Computer system / server 12 may be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. Generally speaking, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform specific tasks or implement specific abstract data types. Computer system / server 12 may be practiced in a distributed cloud computing environment, where tasks are performed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules may be located in local and remote computer system storage media including memory storage devices.

[0045] like Figure 1 As shown, the computer system / server 12 in the cloud computing node 10 is shown in the form of a general-purpose computing device. The components of the computer system / server 12 may include, but are not limited to, one or more processors or processing units 16, a system memory 28, and a bus 18 that couples different system components including the system memory 28 to the processor 16.

[0046] Bus 18 represents any one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example and not limitation, such architectures include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0047] Computer system / server 12 typically includes a variety of computer system readable media. Such media can be any available media that can be accessed by computer system / server 12, and includes volatile and nonvolatile media, removable and non-removable media.

[0048] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer system / server 12 may also include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 34 may be provided for reading from and writing to a non-removable, non-volatile magnetic medium (not shown, and generally referred to as a "hard drive"). Although not shown, a disk drive for reading from or writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical drive for reading from or writing to a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM or other optical media) may be provided. In such a case, each may be connected to the bus 18 via one or more data media interfaces. As will be further depicted and described below, the system memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of an embodiment of the present invention.

[0049] By way of example and not limitation, a program / utility 40 having a set (at least one) of program modules 42, as well as an operating system, one or more applications, other program modules, and program data may be stored in system memory 28. Each or some combination of the operating system, one or more applications, other program modules, and program data may include an implementation of a network environment. Program modules 42 generally perform the functions and / or methods of embodiments of the present invention as described herein.

[0050] The computer system / server 12 may also communicate with one or more external devices 14, such as a keyboard, a pointing device, a display 24, etc.; one or more devices that enable a user to interact with the computer system / server 12; and / or any device that enables the computer system / server 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may occur via an input / output (I / O) interface 22. In addition, the computer system / server 12 may communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with other components of the computer system / server 12 via a bus 18. It should be understood that, although not shown, other hardware and / or software components may be used in conjunction with the computer system / server 12. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.

[0051] In the context of the present invention, and as will be understood by one of ordinary skill in the art, Figure 1 The various components depicted in the can be located in a moving vehicle. For example, some processing and data storage capabilities associated with the mechanisms of the illustrated embodiments may occur locally via local processing components, while the same components are connected via a network to remotely located distributed computing data processing and storage components to achieve different purposes of the present invention. Again, as will be appreciated by those of ordinary skill in the art, this description is intended to convey only a subset of the entire connected network of distributed computing components that may work together to accomplish different inventive aspects.

[0052] See now Figure 2 , depicts an illustrative cloud computing environment 50. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 with which a local computing device used by a cloud consumer can communicate, such as, for example, a personal digital assistant (PDA) or cellular phone 54A, a desktop computer 54B, a laptop computer 54C, and / or an automobile computer system 54N. The nodes 10 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as a private cloud, community cloud, public cloud, or hybrid cloud, or a combination thereof, as described above. This allows the cloud computing environment 50 to provide infrastructure, platform, and / or software as a service for which the cloud consumer does not need to maintain resources on a local computing device. It should be understood that Figure 2 The types of computing devices 54A-N shown in are intended to be illustrative only, and computing node 10 and cloud computing environment 50 may communicate with any type of computerized device over any type of network and / or network-addressable connection (eg, using a web browser).

[0053] See now Figure 3 , showing the cloud computing environment 50 ( Figure 2 ) provides a set of functional abstraction layers. It should be understood in advance that Figure 3 The components, layers, and functions shown in are intended to be illustrative only, and embodiments of the present invention are not limited thereto. As described, the following layers and corresponding functions are provided:

[0054] The device layer 55 includes physical and / or virtual devices that are embedded with and / or independent of electronics, sensors, actuators, and other objects to perform different tasks in the cloud computing environment 50. Each device in the device layer 55 combines networking capabilities with other functional abstraction layers so that information obtained from the device can be provided to other functional abstraction layers, and / or information from other abstraction layers can be provided to the device. In one embodiment, the different devices comprising the device layer 55 can be combined into a physical network collectively referred to as the "Internet of Things" (IoT). As will be understood by those of ordinary skill in the art, such a physical network allows the intercommunication, collection, and dissemination of data to achieve a variety of purposes.

[0055] The device layer 55 as shown includes sensors 52, actuators 53, a "learning" thermostat with integrated processing 56, sensor and network electronics, cameras 57, controllable household outlets / receptacles 58, and as shown a controllable electrical switch 59. Other possible devices may include, but are not limited to, various additional sensor devices, network devices, electronic devices (such as remote control devices), additional actuator devices, so-called "smart" appliances (such as refrigerators or washing machines / dryers), and a variety of other possible interconnected objects.

[0056] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include: mainframes 61; servers based on RISC (Reduced Instruction Set Computer) architecture 62; servers 63; blade servers 64; storage devices 65; and network and networking components 66. In some embodiments, software components include network application server software 67 and database software 68.

[0057] Virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers 71 ; virtual storage 72 ; virtual networks 73 , including virtual private networks; virtual applications and operating systems 74 ; and virtual clients 75 .

[0058] In one example, the management layer 80 may provide the functionality described below. Resource provisioning 81 provides dynamic procurement of computing resources and other resources for performing tasks within a cloud computing environment. Metering and pricing 82 provides cost tracking when resources are utilized within a cloud computing environment, and bills or invoices for the consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. User portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides cloud computing resource allocation and management so that the required service levels are met. Service level agreement (SLA) planning and fulfillment 85 provides pre-scheduling and procurement of cloud computing resources in anticipation of future requirements for the cloud computing resources based on the SLA.

[0059] The workload layer 90 provides examples of functionality that can take advantage of a cloud computing environment. Examples of workloads and functionality that can be provided from this layer include: mapping and navigation 91; software development and lifecycle management 92; virtual classroom education delivery 93; data analysis processing 94; transaction processing 95; and in the context of the illustrated embodiment of the present invention, different workloads and functionality 96 for providing rare topic detection using hierarchical clustering. In addition, the workloads and functionality 96 for providing rare topic detection using hierarchical clustering can include operations such as data analysis (including data collection and processing from organizational databases, online information, knowledge domains, data sources, and / or social networks / media, and other data storage systems, as well as prediction and data analysis functions). Those of ordinary skill in the art will recognize that the workloads and functionality 96 for providing rare topic detection using hierarchical clustering can also work in conjunction with other parts of different abstraction layers, such as those in hardware and software 60, virtualization 70, management 80, and other workloads 90 (e.g., such as data analysis and / or replaceability processing 94) to achieve different purposes of the illustrated embodiment of the present invention.

[0060] Now go to Figure 4 , block diagram 400 depicts a computing system for providing rare topic detection using hierarchical clustering. In one aspect, Figure 1-3 One or more of the components, modules, services, applications and / or functions described in the Figure 4 For example, in combination with the processing unit 16 Figure 1 The computer system / server 12 may be used to perform various computing, data processing and other functions according to various aspects of the present invention.

[0061] like Figure 4 As shown, system 400 may include server 402, one or more networks 404, and one or more data sources 406. Server 402 may include hierarchical topic modeling component 408, which may include learning component 410, hierarchical topic component 412, clustering component 414, identification component 415, enhancement component 416, and / or seeding component 418. Server 402 may also include or otherwise be associated with at least one memory 420. Server 402 may also include system bus 422, which may couple various components, including, but not limited to, hierarchical topic modeling component 408 and associated components, memory 420, and / or processor 424. Although in Figure 4 Server 402 is shown in FIG. 1 , but in other embodiments, any number of different types of devices may be associated with the server 402. Figure 4 The components shown are associated with or included in Figure 4The components shown in are as part of hierarchical topic modeling component 408. All such embodiments are contemplated.

[0062] Hierarchical topic modeling component 408 can use hierarchical topic modeling that can be learned from one or more data sources 406 to promote rare topic detection. Data source 406 may include structured and / or unstructured data. The term "unstructured data" may refer to data presented in an unrestricted natural language and intended for human consumption. Unstructured data may include, but is not limited to: session data associated with a computing system / application for communicating with one or more users, social media posts and / or comments, and associated metadata, news posts and / or comments, and associated metadata, and / or posts and / or comments, and associated metadata made by one or more users on one or more websites that promote discussion. Unstructured data may be generated by one or more entities (e.g., one or more users) and may include information contributed to a corpus (e.g., the Internet, a website, a network, etc.) in a non-digital language (e.g., a spoken language) intended for human consumption.

[0063] In various embodiments, one or more data sources 406 may include data accessible by server 402 directly or via one or more networks 404 (e.g., an intranet, the Internet, a communication system, and / or a combination thereof). For example, one or more data sources 406 may include a computer-readable storage device (e.g., a primary storage device, a secondary storage device, a tertiary storage device, or an offline storage device) that may store user-generated data. In another example, one or more data sources 406 may include a community host that includes a website and / or application that facilitates sharing of user-generated data via a network (e.g., the Internet).

[0064] One or more servers 402 including a hierarchical topic modeling component 408 and one or more data sources 406 may be connected directly or via one or more networks 404. Such networks 404 may include wired and wireless networks, including but not limited to cellular networks, wide area networks (WANs) (e.g., the Internet), or local area networks (LANs). For example, the server 402 may communicate with one or more data sources 406 (or vice versa) using virtually any desired wired or wireless technology (including, for example, cellular, WAN, wireless fidelity (Wi-Fi), Wi-Max, WLAN, etc.). Further, although in the illustrated embodiment, the hierarchical topic modeling component 408 is provided on the server device 402, it should be understood that the architecture of the system 400 is not limited thereto. For example, the hierarchical topic modeling component 408 or one or more components of the hierarchical topic modeling component 408 may be located at another device, such as another server device, a client device, etc.

[0065] In one aspect, the learning component 410 can learn the hierarchical topic model from one or more data sources 406. The learning component 410 can perform one or more machine learning operations, such as, for example, natural language processing (“NLP”). The topic model database 426 can store, maintain, and access each hierarchical topic model (including each newly learned hierarchical topic model) that can be maintained / stored in the memory 420 via the topic model database 426.

[0066] The clustering component 414 can generate one or more word vectors based on data obtained from one or more data sources 406, and score each word vector in the one or more word vectors. The clustering component 414 can also generate multiple clusters from one or more word vectors. The selected cluster can be identified from multiple clusters and identified / marked as a king cluster. That is, a K-means clustering operation can be used when aggregating word vectors into K clusters in each iteration, where "K" is a positive integer or a defined value. The K clusters may include one or more "king clusters". In one aspect, the king cluster is the largest cluster from a total of K clusters (e.g., the cluster containing the most documents or data sources). The king cluster and can be the largest cluster among multiple clusters.

[0067] The clustering component 414 can divide the selected cluster into multiple clusters at each iteration. The clustering component 414 associated with the recognition component 415 can identify an alternative selected cluster (e.g., a second or alternative king cluster) from multiple clusters, while iteratively removing one or more dominant words in the alternative selected cluster. That is, the clustering component 414 associated with the recognition component 415 can identify one or more differences between each cluster in the multiple clusters, while iteratively removing one or more dominant words in the selected cluster at each iteration. In one aspect, the alternative selected cluster can also be a king cluster, and the alternative king cluster is the largest cluster of subsequent cluster iterations from the multiple clusters.

[0068] The hierarchical topic component 412 can use the hierarchical topic model to iteratively remove one or more dominant words in the selected cluster. In one aspect, the dominant words are related to one or more main topics of the cluster.

[0069] Seeding component 418 can seed the learned hierarchical topic model with one or more words, n-grams, phrases, text fragments or combinations thereof to evolve the hierarchical topic model. Hierarchical topic component 412 associated with seeding component 418 and / or expansion component 416 can restore removed domain words after completing seeding. In one aspect, seeding component 418 can seed the hierarchical topic model with an existing topic model. In addition, seeding component 418 can seed each of a plurality of clusters according to one or more cluster models.

[0070] Thus, the hierarchical topic modeling component 408 provides interpretability and understandability, where topics can be explained by domain experts (e.g., descriptions of topics can be read by users). The hierarchical topic modeling component 408 provides multi-level summarization (e.g., word, ngram, fragment, document). In one aspect, word and ngram level representations can be used for machine learning, and ngram and fragment level representations are used for consumption by analysts who are domain experts. The hierarchical topic modeling component 408 provides scalability and real-time scoring (real-time), where training can occur from one or more corpora, and the hierarchical topic model can be trained in real-time.

[0071] Thus, as described herein, the hierarchical topic modeling component 408 provides a learned hierarchical topic model that progressively removes (e.g., suppresses or hides) one or more dominant words in a king cluster. A king cluster can be identified by (a) size (e.g., a king cluster is determined by the size of the cluster) and (b) lack of cohesion (e.g., large clusters tend to have low cohesion because they are sparser). The hierarchical topic modeling component 408 provides a hierarchical topic model that is enhanced with human-interpretable words, phrases, and fragments. The removed words (e.g., suppressed or hidden words) can be restored (e.g., not suppressed and / or not hidden) along the hierarchical structure to provide increased interpretability. The hierarchical topic modeling component 408 provides topic evolution by seeding the topic model for incremental training. A group of metrics can be used to capture differences (e.g., size, cohesion, centroid shift, tree structure changes). In one aspect, a group of metrics can be used to capture differences between new and old topic models, such as, for example: 1) size (e.g., how does the cluster (topic) size change, such as, for example, how does the number of documents falling under a topic change), 2) cohesion (e.g., do the clusters become sparser or tighter?) 3) centroid shift (e.g., how do the cluster centers move?) and / or 4) changes in tree structure (e.g., does the overall structure of the topic model change?).

[0072] Now turn to Figure 5 , Graph 500 depicts rare topic detection using hierarchical topic modeling. That is, Graph 500 depicts a plurality of clusters assuming that document feature vectors are in a two-dimensional ("2D") space. In one aspect, Figure 1-5 One or more of the components, modules, services, applications and / or functions described in the Figure 5 For the sake of brevity, repeated descriptions of similar elements, components, modules, services, applications, and / or functions used in other embodiments described herein are omitted.

[0073] For example, Figure 510 (e.g., original hierarchical topic model 510) depicts an original / existing topic model with clusters 1 to 4. Figure 520 (e.g., new hierarchical topic model 520) depicts the evolution of topic modeling by providing rare topic detection using hierarchical topic modeling. That is, the new hierarchical topic model 520 is obtained after seeding the original hierarchical topic model 510. As depicted, cluster 1 of the new hierarchical topic model 520 has increased in size. The center of cluster 2 has shifted, and the center of the new hierarchical topic model 520 has shrunk in size. Cluster 3 of the new hierarchical topic model 520 has disappeared (e.g., has been eliminated). The size of cluster 4 has decreased. It should be noted that the hierarchical topic model 520 is used only as an example and shows how the topic model evolves from the original seed model. Thus, as depicted, based on seeding and retraining the hierarchical topic model 520 on a new data set, one or more optimal solutions are incrementally identified for clustering, where the clusters evolve into one or more different shapes, sizes, and / or even existences.

[0074] Now turn to Figure 6 , depicts a method 600 for providing rare topic detection using hierarchical topic modeling by a processor, in which various aspects of the illustrated embodiment may be implemented. That is, Figure 6 6 is a flow chart of an additional example method 600 for providing rare topic detection using hierarchical topic modeling in a computing environment according to an example of the present invention. Function 600 may be implemented as a method executed as instructions on a machine, wherein the instructions are included on at least one computer readable medium or a non-transitory machine readable storage medium. Function 600 may start in block 602.

[0075] In block 604, a hierarchical topic model may be learned from one or more data sources. In block 606, the hierarchical topic model may be used to iteratively remove one or more dominant words in the selected cluster. The dominant words may relate to one or more main topics of the cluster. The learned hierarchical topic model may be seeded with one or more words, n-grams, phrases, text snippets, or a combination thereof to evolve the hierarchical topic model, and the domain words removed in block 608 are restored when the seeding is completed. Function 600 may end in block 610.

[0076] In one aspect, combining Figure 6 At least one box and / or as Figure 6As part of at least one box of , the operation of 600 may include one or more of each of the following. The operation of 600 may generate one or more word vectors and score each of the one or more word vectors, and may also generate multiple clusters from the one or more word vectors, wherein the selected cluster is identified from multiple clusters and is a king cluster, wherein the king cluster is the largest cluster from multiple clusters. The operation of 600 may split the selected cluster into multiple clusters at each iteration, and / or identify an alternative selected cluster from multiple clusters while iteratively removing one or more dominant words in the alternative selected cluster. The alternative selected cluster is the king cluster and the king cluster is the largest cluster from multiple clusters.

[0077] The operations of 600 may utilize an existing topic model to seed a hierarchical topic model, and / or seed each of a plurality of clusters according to one or more cluster models.

[0078] The operations of 600 may identify one or more differences between each of the plurality of clusters while iteratively removing one or more dominant terms in the selected cluster at each iteration.

[0079] The present invention may be a system, method, and / or computer program product. The computer program product may include computer-readable storage medium(s) having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention.

[0080] Computer readable storage medium can be a tangible device that can retain and store instructions for use by instruction execution devices. Computer readable storage medium can be, for example but not limited to, electronic storage device, magnetic storage device, optical storage device, electromagnetic storage device, semiconductor storage device or any suitable combination of the above. A non-exhaustive list of more specific examples of computer readable storage medium includes the following: portable computer disk, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanical encoding device such as punch card or a protruding structure in a groove with instructions recorded thereon, and any suitable combination of the above. Computer readable storage medium as used herein should not be interpreted as a temporary signal itself, such as radio waves or other free propagating electromagnetic waves, electromagnetic waves propagated by waveguides or other transmission media (e.g., light pulses passing through fiber optic cables) or electrical signals emitted by wires.

[0081] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network), or downloaded to an external computer or external storage device. The network can include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in a computer-readable storage medium in the corresponding computing / processing device.

[0082] The computer-readable program instructions for performing the operation of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as Smalltalk, C++, etc.) and conventional procedural programming languages ​​(such as "C" programming languages ​​or similar programming languages). The computer-readable program instructions can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer, partially on the remote computer, or completely on the remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network (including local area network (LAN) or wide area network (WAN)), or can be connected to an external computer (for example, using an Internet service provider through the Internet). In certain embodiments, the electronic circuit including, for example, a programmable logic circuit, a field programmable gate array (FPGA) or a programmable logic array (PLA) can be executed by using the state information of the computer-readable program instructions to personalize the electronic circuit to perform computer-readable program instructions, so as to perform various aspects of the present invention.

[0083] The present invention will be described below with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present invention. It should be understood that each box of the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer readable program instructions.

[0084] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device create a device for implementing the functions / actions specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which causes the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable storage medium in which the instructions are stored includes a manufactured product containing instructions for implementing aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0085] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable apparatus, or other device to produce computer-implemented processing, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0086] The flow charts and block diagrams in the accompanying drawings show the architecture, functions and operations of possible implementations of the system, method and computer program product according to different embodiments of the present invention. To this end, each box in the flow chart or block diagram may represent a module, segment or part of an instruction, which includes one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the box may not occur in the order marked in the figure. For example, depending on the functions involved, two boxes that can be shown sequentially can actually be executed substantially simultaneously, or these boxes can sometimes be executed in reverse order. It should also be noted that each box in the block diagram and / or flow chart, and the combination of boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or action or performs a combination of dedicated hardware and computer instructions.

Claims

1. A method for providing rare topic detection using hierarchical topic modeling by a processor, comprising: Learning a hierarchical topic model from one or more data sources; training the hierarchical topic model by iteratively removing one or more dominant terms in a selected cluster using the hierarchical topic model during progressive drill-down operations through a plurality of hierarchical topic modeling executions, wherein, in each iteration of the plurality of hierarchical topic modeling executions, the progressive drill-down operations remove those of the one or more dominant terms identified during a previous iteration that are no longer discriminatory for a next execution of the plurality of hierarchical topic modeling executions, and wherein the dominant terms are related to one or more primary topics of the cluster; and The learned hierarchical topic model is further trained by seeding the learned hierarchical topic model with one or more words, n-grams, phrases, text fragments, or a combination thereof to evolve the hierarchical topic model, wherein upon completion of the seeding, the removed dominant words are restored, and wherein each of the dominant words removed from each iteration and restored upon completion of the seeding are used together to form a natural language interpretation of each of the one or more main topics obtained from the hierarchical topic model within a corpus of the one or more data sources.

2. The method according to claim 1, further comprising: One or more word vectors are generated, and each of the one or more word vectors is scored.

3. The method according to claim 2, further comprising: A plurality of clusters are generated from the one or more word vectors, wherein a selected cluster is identified from the plurality of clusters and is a king cluster, wherein the king cluster is a largest cluster among the plurality of clusters.

4. The method according to claim 1, further comprising: Split the selected cluster into multiple clusters at each iteration; While iteratively removing one or more dominant words in the candidate selected cluster, an candidate selected cluster is identified from the plurality of clusters, wherein the candidate selected cluster is a king cluster and the king cluster is a largest cluster among the plurality of clusters.

5. The method according to claim 1, further comprising: The hierarchical topic model is seeded with an existing topic model.

6. The method according to claim 1, further comprising: Each of the plurality of clusters is seeded according to one or more cluster models.

7. The method according to claim 1, further comprising: While iteratively removing one or more dominant terms in the selected cluster at each iteration, one or more differences between each cluster in the plurality of clusters are identified.

8. A system for providing rare topic detection using hierarchical topic modeling in a computing environment, comprising: One or more computers having executable instructions that, when executed, cause the system to: Learning a hierarchical topic model from one or more data sources; training the hierarchical topic model by iteratively removing one or more dominant terms in a selected cluster using the hierarchical topic model during progressive drill-down operations through a plurality of hierarchical topic modeling executions, wherein, in each iteration of the plurality of hierarchical topic modeling executions, the progressive drill-down operations remove those of the one or more dominant terms identified during a previous iteration that are no longer discriminatory for a next execution of the plurality of hierarchical topic modeling executions, and wherein the dominant terms are related to one or more primary topics of the cluster; and The learned hierarchical topic model is further trained by seeding the learned hierarchical topic model with one or more words, n-grams, phrases, text fragments, or a combination thereof to evolve the hierarchical topic model, wherein upon completion of the seeding, the removed dominant words are restored, and wherein each of the dominant words removed from each iteration and restored upon completion of the seeding are used together to form a natural language interpretation of each of the one or more main topics obtained from the hierarchical topic model within a corpus of the one or more data sources.

9. The system according to claim 8, wherein: The executable instructions, when executed, cause the system to generate one or more word vectors and score each of the one or more word vectors.

10. The system according to claim 9, wherein: The executable instructions, when executed, cause the system to generate a plurality of clusters from the one or more word vectors, wherein a selected cluster is identified from the plurality of clusters and is a king cluster, wherein the king cluster is a largest cluster among the plurality of clusters.

11. The system according to claim 8, wherein: The executable instructions, when executed, cause the system to: Split the selected cluster into multiple clusters at each iteration; as well as While iteratively removing one or more dominant words in the candidate selected cluster, an candidate selected cluster is identified from the plurality of clusters, wherein the candidate selected cluster is a king cluster and the king cluster is a largest cluster among the plurality of clusters.

12. The system according to claim 8, wherein: The executable instructions, when executed, cause the system to seed the hierarchical topic model with an existing topic model.

13. The system according to claim 8, wherein: The executable instructions, when executed, cause the system to seed each of a plurality of clusters according to one or more cluster models.

14. The system according to claim 8, wherein: The executable instructions, when executed, cause the system to identify one or more differences between each of the plurality of clusters while iteratively removing one or more dominant terms in a selected cluster at each iteration.

15. A computer program product for providing rare topic detection using hierarchical topic modeling by a processor, the computer program product comprising a non-transitory computer readable storage medium having computer readable program code portions stored therein, the computer readable program code portions comprising: An executable part of learning a hierarchical topic model from one or more data sources; training an executable portion of the hierarchical topic model by iteratively removing one or more dominant terms in a selected cluster using the hierarchical topic model during progressive drill-down operations through a plurality of hierarchical topic modeling executions, wherein, in each iteration of the plurality of hierarchical topic modeling executions, the progressive drill-down operations remove those of the one or more dominant terms identified during a previous iteration that are no longer discriminatory for a next execution of the plurality of hierarchical topic modeling executions, and wherein the dominant terms are related to one or more primary topics of the cluster; and The learned hierarchical topic model is further trained to evolve an executable portion of the hierarchical topic model by seeding the learned hierarchical topic model with one or more words, n-grams, phrases, text fragments, or a combination thereof, wherein upon completion of the seeding, the removed dominant words are restored, and wherein each of the dominant words removed from each iteration and restored upon completion of the seeding are used together to form a natural language interpretation of each of the one or more main topics derived from the hierarchical topic model within a corpus of the one or more data sources.

16. The computer program product of claim 15, further comprising: An executable portion that generates one or more word vectors and scores each of the one or more word vectors.

17. The computer program product of claim 16, further comprising: An executable portion generates a plurality of clusters from the one or more word vectors, wherein a selected cluster is identified from the plurality of clusters and is a king cluster, wherein the king cluster is a largest cluster among the plurality of clusters.

18. The computer program product of claim 15, further comprising an executable portion for: Splitting the selected cluster into multiple clusters at each iteration; and While iteratively removing one or more dominant words in the candidate selected cluster, an candidate selected cluster is identified from the plurality of clusters, wherein the candidate selected cluster is a king cluster and the king cluster is a largest cluster among the plurality of clusters.

19. The computer program product of claim 15, further comprising an executable portion for: seeding the hierarchical topic model with an existing topic model; or Each of the plurality of clusters is seeded according to one or more cluster models.

20. The computer program product of claim 15, further comprising an executable portion for identifying one or more differences between each of the plurality of clusters while iteratively removing one or more dominant terms in a selected cluster at each iteration.

Citation Information

Patent Citations

  • Comparative web search system and method

    US20080222140A1

  • System And Method For Providing Multi-Core And Multi-Level Topical Organization In Social Indexes

    US20110270830A1

  • System and Method for Association Extraction for Surf-Shopping

    US20130212110A1

  • Data-dependent clustering of geospatial words

    US9697245B1