Context-Aware Data Mining
The method uses word embeddings to analyze log data, aggregating representations to retrieve similar segments, addressing the complexity of log data analysis and enhancing diagnostic efficiency.
Patent Information
- Application Number
- CN202080039160.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-11
- Filing Date
- 2020-05-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-05-27
AI Technical Summary
It is difficult for the prior art to effectively use the log data of the computing system for system diagnosis and troubleshooting, especially because the log data is large in size, many types and fast in speed, which leads to challenges for system administrators.
The context-aware data mining method of text documents is adopted, and the distributed embedding representation is calculated using the word embedding model to retrieve similar document fragments, return the ranking list, and use the neural network model to train the word embedding model and use the similarity metric to match.
Improve the ability to find meaningful events in log data, and improve the diagnostic and troubleshooting efficiency of log data through unsupervised learning and context-based matching.
Smart Images

Figure CN113906445B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to knowledge extraction, representation, retrieval, and reasoning, and more particularly to context-aware data mining of text documents. Background Art
[0002] Word embeddings are a class of techniques in which individual words are represented as real-valued vectors in a predefined vector space. Each word is associated with a point in the vector space. Each word is represented by a real-valued feature vector having dozens or hundreds of dimensions, where each dimension is associated with a feature representing some aspect of the word. This is in contrast to the thousands or millions of dimensions required for sparse word representations such as one-hot encoding, in which a word is represented by a single component in a vector whose size corresponds to the size of the vocabulary, and this representation is called a "bag of words". On the other hand, the number of features is much smaller than the size of the vocabulary. The distributed representation is learned based on the use of words, based on the idea that words with similar contexts will have similar meanings. This enables words used in a similar manner to produce similar representations, thus naturally capturing their meanings. This can be contrasted with the bag-of-words model in which different words with similar meanings can have very different representations. Using dense and low-dimensional vectors has computational advantages, because most neural network toolkits do not handle high-dimensional sparse vectors well. Another advantage of dense representations is the generalization ability: if certain features are considered to provide similar cues, it is worthwhile to provide a representation that can capture these similarities. In word embeddings for natural language processing (NLP), words or phrases from natural language are represented by vectors of real numbers. This representation can be based entirely on the way the word is used, i.e., its context.
[0003] Log data from computing systems is crucial for understanding and diagnosing system problems. Log data is very large in terms of volume, variety, velocity, etc., and using it for system diagnosis and troubleshooting is a challenge for system administrators. The log data of computing systems can be represented in the NLP format of word embeddings, although no specific representation is specified. For example, each word in the log can be used as a token, but the entire log line can also be regarded as a token. Summary of the Invention
[0004] According to one aspect of the present invention, there is provided a method for context-aware data mining of text documents, comprising: receiving a list of words parsed and preprocessed from an input query, using a word embedding model of the text document being queried to calculate relevant distributed embedding representations for each word in the list of words, aggregating the relevant distributed embedding representations of all words in the list of words to represent the input query with a single embedding, retrieving a ranked list of document fragments of N lines similar to the aggregated word embedding representation of the query, and returning the list of retrieved fragments to the user.
[0005] According to an embodiment, aggregating the relevant distributed embedding representations is performed using either the average of all relevant distributed embedding representations or the maximum of all relevant distributed embedding representations.
[0006] According to another embodiment, N is a positive integer provided by the user.
[0007] According to another embodiment, the method includes training a word embedding model of the text document in the following manner: parsing and preprocessing the text document and generating a list of tokenized words; defining a word dictionary from the list of tokenized words, where the word dictionary includes at least some tokens from the list of tokenized words; and training the word embedding model, where the word embedding model is a neural network model that represents each word or line in the dictionary with a vector.
[0008] According to another embodiment, parsing and preprocessing the text document includes: removing all punctuation and leading from each line in the text document, parsing numerical data, tokenizing the text document by words to form a list of tokenized words, where a token is either a single word, an N-gram of N consecutive words, or an entire line of the document, and returning the list of tokenized words.
[0009] According to another embodiment, the text document is a computer system log, and the numerical data includes decimal numbers and hexadecimal addresses.
[0010] According to another embodiment, the method includes parsing and preprocessing the input query by: removing all punctuation from the input query, parsing numerical data, tokenizing the input query by words to produce a list of tokenized words, where a token is either a single word, an N-gram of N consecutive words, or an entire line of the input query, and returning the list of tokenized words.
[0011] According to another embodiment, retrieving a ranked list of N lines of document fragments similar to the aggregated word embedding representation of a query includes: retrieving a ranked list of N lines of document fragments similar to the aggregated word embedding representation of a query includes: comparing the aggregated word embedding representation of the query with a word embedding model of a text document using a similarity metric, returning those fragments of the word embedding model of the text document whose similarity to the aggregated word embedding representation of the query is greater than a predetermined threshold, and ranking the retrieved document fragments by similarity.
[0012] According to another aspect of the present invention, there is provided a method for context-aware data mining of a text document, including: parsing and preprocessing the text document and generating a list of tokenized words; defining a word dictionary from the list of tokenized words, where the word dictionary includes at least some tokens from the list of tokenized words; and training a word embedding model, where the word embedding model is a neural network model that represents each word or line in the dictionary by a vector. Parsing and preprocessing the text document includes: removing all punctuation marks and leading characters from each line in the text document, parsing numerical data, tokenizing the text document by words to form a list of tokenized words, where a token is one of a single word, an N-gram of N consecutive words, or an entire line of the document, and returning the list of tokenized words.
[0013] According to an embodiment, the method includes: receiving a list of words parsed and preprocessed from an input query; using the word embedding model of the text document being queried to calculate the relevant distributed embedding representation of each word; aggregating the relevant distributed embedding representations of all words in the word list to represent the input query by a single embedding; retrieving a ranked list of N lines of document fragments similar to the aggregated word embedding representation of the query, and returning the list of the retrieved fragments to the user.
[0014] According to another embodiment, parsing and preprocessing the input query includes: removing all punctuation marks from the input query, parsing numerical data, tokenizing the input query by words to produce a list of tokenized words, where a token is one of a single word, an N-gram of N consecutive words, or an entire line of the input query, and returning the list of tokenized words.
[0015] According to another embodiment, retrieving a ranked list of N lines of document fragments similar to the aggregated word embedding representation of a query includes: comparing the aggregated word embedding representation of the query with a word embedding model of a text document using a similarity metric, returning those fragments of the word embedding model of the text document whose similarity to the aggregated word embedding representation of the query is greater than a predetermined threshold, and ranking the retrieved document fragments by similarity.
[0016] According to another embodiment, the text document is a computer system log, and the numerical data includes decimal numbers and hexadecimal addresses.
[0017] According to another embodiment, aggregating the relevant distributed embedding representations is performed using either the average of all the relevant distributed embedding representations or the maximum of all the relevant distributed embedding representations.
[0018] According to another embodiment, N is a positive integer provided by the user.
[0019] According to another aspect of the present invention, there is provided a computer-readable program storage device tangibly embodying an instruction program executable by a computer to perform method steps for context-aware data mining of text documents.
[0020] The exemplary embodiments described below are directed to a novel interface in which a user can express a query as any kind of text, such as words, lines, paragraphs, etc., and a dedicated NLP-based algorithm returns fragments of computer system log data having a word context similar to the query. The method according to an embodiment of the present disclosure is based on the context of the words in the query rather than on simple string matching. This improves the user's ability to find meaningful events in the logs. The method according to an embodiment is based on unsupervised learning. It relies on the text information already present in the logs and can be applied without any prior knowledge of the events, keywords, or the structure of the log text. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a block diagram of a method for processing a user query according to an embodiment of the present invention.
[0022] Figure 2 is a block diagram of a method for creating a model according to an embodiment of the present invention.
[0023] Figure 3 is a block diagram of a method according to another embodiment of the present invention.
[0024] Figure 4 is a schematic diagram of an exemplary cloud computing node implementing an embodiment of the present invention.
[0025] Figure 5 illustrates an exemplary cloud computing environment employed in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The exemplary embodiments described herein generally provide methods for NLP-based context-aware log mining. While the embodiments are susceptible to various modifications and alternative forms, specific embodiments thereof are shown by way of example in the drawings and will be described in detail herein. However, it should be understood that the present disclosure is not intended to be limited to the particular forms disclosed, but on the contrary, the present disclosure will cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure.
[0027] Figure 1 is a block diagram of a method for processing a user query according to an embodiment of the present disclosure. Figure 1 A use case is shown in which a user provides a query 110, which can be a single word 111, a line 112, or a paragraph 113, the number of lines 114 representing the size of the retrieved snippet, and a similarity threshold that defines how many snippets should be returned. For a method for retrieving 115 similar log snippets from a computer system log according to an embodiment, the input query 110 and the number of lines 114 are provided. The method according to the embodiment will return a set of snippets 120.1, 120.2, …, 120.M as output 120, and these snippets are sorted by similarity to the query text.
[0028] Figure 2 is a block diagram of a method for creating a model according to an embodiment of the present disclosure. Figure 2 The steps required to create a model according to an embodiment are shown. On the left is a flowchart of a method 210 for training a word embedding model. In the upper right is a flowchart of a method 220 for processing a system log file to obtain a list of tokenized words, and in the lower right is a block diagram 230 of a general word embedding structure.
[0029] Referring to the flowchart 210, the method for training the model includes parsing and preprocessing the log output of the raw data from the computer system log (211), defining a word dictionary (212), and training a word embedding model (213).
[0030] Method 220 is the steps involved in parsing and preprocessing the log output of the raw data (211), and includes removing all punctuation and leading characters from each line (222), parsing numbers and hexadecimal addresses (223), tokenizing the log by words (224), and returning a list of tokenized words (225). Specific tokens are used to parse numbers and hexadecimal addresses. According to an embodiment, a decimal number is represented by one token, a hexadecimal address is represented by another token, and the information of the number or address can be used as a placeholder together with the token, although the context is not limited to specific values. Although any NLP technique can be used, no text processing technique is used for words. Tokenizing the log means splitting the log into tokens, where a token can be defined as a single word or an N-gram of N consecutive words, or can also be defined as an entire line of the log. Once the log is tokenized, the dictionary is the set of all tokens, or a subset of the selected tokens, such as a subset of the most frequent tokens.
[0031] According to an embodiment, the dictionary is used to define and represent the words (or lines) considered in the word embedding model. Refer to Figure 2In step 231, the input word w[t] is represented as a one-hot vector with a number of elements equal to the size of the dictionary, where all 0s and 1s in the elements correspond to the word (or line). For this purpose, a dictionary needs to be defined to create these vectors. For example, the one-hot vector representations of "Rome", "Paris", "Italy", and "France" in a V-dimensional vector space are in the form of:
[0032] Rome = [1, 0, 0, 0, 0, …, 0],
[0033] Paris = [0, 1, 0, 0, 0, …, 0],
[0034] Italy = [0, 0, 1, 0, 0, …, 0],
[0035] France = [0, 0, 0, 1, 0, …, 0].
[0036] The word embedding model 230 uses a distance metric between N lines of log fragments, where N is a user-defined parameter. This distance metric is used to define the degree of similarity between the contexts of two log fragments. Specifically, this metric is used to retrieve the top N segments with the highest similarity to the user query. The word embedding model is a neural network model that represents words with embeddings (i.e., vectors). Then the word [t] 231 is projected 232 into the word embedding 233, which includes the words [t-2], [t-1], [t+1], and [t+2] corresponding to a window size of 5. The window size is one of the parameters of the model provided by the user. Once the word embedding model has been trained, the distance metric can be used between N-line segments of the log, where N is a user-defined parameter. A typical distance metric between word embeddings (i.e., vectors) is cosine similarity. Additionally, supervised learning methods can be used. This involves training a supervised model (such as a long short-term memory (LSTM)) to predict the similarity between documents.
[0037] Figure 3 is a block diagram of the search during this use case, including a flowchart showing the extraction 310 and query prediction 320 from the query use case. Referring to the query use case 310, the user 300 provides a query - such as a part of unstructured text of interest, and the number of lines of each segment of interest. The query use case 310 also shows training the model 313 by inputting 311 the system log in step 312 to create the model. The step of training the model in step 312 corresponds to the method according to the trained Figure 2 embodiment of the model 210.
[0038] In step 314, based on Figure 2Method 220 for parsing queries in a processing system log file, the method comprising the steps of: removing all punctuation and leading characters from each line, parsing numbers and hexadecimal addresses, and tokenizing the query by word: q = [w1, w2, …, w N .
[0039] Then, the output q is provided as input to step 315 of retrieving similar log segments, which corresponds to Figure 1 step 115, and outputs log segment 316. The model will retrieve an ordered list of segments similar to the query and output them to the user.
[0040] Box 320 is a flowchart of the steps involved in query prediction and begins by receiving, at step 321, a parsed and preprocessed list of words from the input. Then, at step 322, the Figure 2 word embedding model 230 is used to compute the relevant distributed embedding representation we i for each word w i ."Distributed embedding representation" is the representation of a word (or line) given by the word embedding model. At step 323, the relevant distributed embedding representations we i are aggregated with the mean (or maximum) of all the distributed word embeddings we i to represent the query by a single embedding qe. Fragments of each N lines in the log data are represented in the same way. The word embedding produces a representation with a vector for each word or line. Since the query contains more than one word or line, all the representations need to be aggregated into a single vector to represent the entire query as a vector, after which similarity metrics such as cosine similarity can be used.
[0041] At step 324, a ranked list of N-line log segments is retrieved where the aggregated word embeddings have a high similarity to the query representation. By comparing the aggregated word embedding representation of the query with the word embedding model of the log data using a similarity metric, those segments of the word embedding model of the log data whose similarity is greater than a predetermined threshold are returned and the retrieved segments are sorted by similarity value to retrieve the ranked list of log segments. The list of retrieved segments is returned at step 325 and output at step 316.
[0042] An actual example of the user query processing method according to an embodiment of the present invention is as follows. The method begins by receiving a user input query of word w q . Assume that the word w q frequently appears in patterns such as w1, w2, w q , w3, w4. By searching for w qIt is very likely to retrieve an exact match for the pattern. Now, assume there is another word pattern such as w1, w5, w6, w3, w4. Through lexical matching, this pattern will not be retrieved because it does not contain the query w q . However, using the proposed NLP method according to an embodiment of the present disclosure, the contexts of these two log segments are very similar, and the query pattern can be retrieved and presented to the user.
[0043] Although embodiments of the present disclosure have been described in the context of querying computer system logs, it will be apparent to those skilled in the art that the methods according to embodiments of the present disclosure can be applied to query any text document that is too large to be searched or understood by a single person.
[0044] System implementation
[0045] It should be understood that embodiments of the present disclosure can be implemented in different forms of hardware, software, firmware, special processes, or combinations thereof. In one embodiment, an embodiment of the present disclosure can be implemented in software form as an application tangibly embodied on a computer-readable program storage device. The application can be uploaded to and executed by a machine including any suitable architecture. Additionally, it should be understood in advance that although the present disclosure includes a detailed description of cloud computing, the implementation of the teachings recited herein is not limited to a cloud computing environment. Instead, embodiments of the present disclosure can be implemented in conjunction with any other type of computing environment now known or later developed. The automatic troubleshooting system according to an embodiment of the present disclosure is also applicable to cloud implementation.
[0046] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., a shared pool of configurable computing resources such as networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with the service provider. The cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0047] The characteristics are as follows:
[0048] On-demand self-service: Cloud consumers can unilaterally and automatically provision computing capabilities such as server time and network storage as needed, without human interaction with the service provider.
[0049] Broad network access: The capabilities are available over the network and accessed through standard mechanisms that facilitate the use of heterogeneous thin client platforms or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0050] Resource Pool: The provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically assigned and reassigned as needed. There is a sense of location independence, as consumers generally have no control or knowledge of the exact location of the resources provided, but may be able to specify a location at a higher level of abstraction (e.g., country, state, or data center).
[0051] Rapid Elasticity: The ability to provide capabilities quickly and elastically, with the ability to scale out rapidly in some cases and automatically scale in and release quickly to scale in. To the consumer, the capabilities available for provisioning generally appear unlimited and can be purchased in any quantity at any time.
[0052] Measured Service: The cloud system automatically controls and optimizes resource use by leveraging metering capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource use can be monitored, controlled, and reported, providing transparency for both the provider and consumer of the utilized service.
[0053] The service models are as follows:
[0054] Software as a Service (SaaS): The capabilities provided to the consumer are to use the provider's applications running on the cloud infrastructure. The applications can be accessed from different client devices through a thin client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.
[0055] Platform as a Service (PaaS): The capabilities provided to the consumer are to deploy applications created by the consumer or acquired and created using programming languages and tools supported by the provider onto the cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but has control over the deployed applications and possibly the application hosting environment configuration.
[0056] Infrastructure as a Service (IaaS): The capabilities provided to the consumer are to provide processing, storage, networking, and other basic computing resources where the consumer can deploy and run arbitrary software, which can include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but has control over the operating systems, storage, deployed applications, and possibly limited control over the selected networking components (e.g., host firewall).
[0057] The deployment models are as follows:
[0058] Private Cloud: The cloud infrastructure is for the exclusive use of an organization. It can be managed by the organization or a third party and can exist on-premises or off-premises.
[0059] Community cloud: The cloud infrastructure is shared by several organizations and supports a specific community that shares concerns (e.g., tasks, security requirements, policies, and compliance considerations). It can be managed by an organization or a third party and can exist on - premise or off - premise.
[0060] Public cloud: Makes the cloud infrastructure available to the public or a large industry group and is owned by an organization that sells cloud services.
[0061] Hybrid cloud: The cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain unique entities but are bound together by standardized or proprietary technologies (e.g., cloud bursting for load balancing between clouds) that enable data and application portability.
[0062] The cloud computing environment is service - oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. The core of cloud computing is an infrastructure that includes a network of interconnected nodes.
[0063] Now referring to Figure 4 , a schematic diagram of an example of a cloud computing node is shown. The cloud computing node 410 is merely an example of a suitable cloud computing node and is not intended to impose any limitation on the scope of use or functionality of the embodiments of the present disclosure described herein. In any case, the cloud computing node 410 is capable of implementing and / or performing any of the functions set forth above herein.
[0064] In the cloud computing node 410, there is a computer system / server 412, which can operate with many other general - purpose or special - purpose computing system environments or configurations. Examples of well - known computing systems, environments, and / or configurations suitable for use with the computer system / server 412 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor - based systems, set - top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems or devices, etc.
[0065] The computer system / server 412 can be described in the general context of computer system - executable instructions, such as program modules, executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform specific tasks or implement specific abstract data types. The computer system / server 412 can be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located in both local and remote computer system storage media including memory storage devices.
[0066] AsFigure 4 As shown, the computer system / server 412 in the cloud computing node 410 is shown in the form of a general-purpose computing device. The components of the computer system / server 412 may include, but are not limited to, one or more processors or processing units 416, a system memory 428, and a bus 418 that couples different system components including the system memory 428 to the processor 416.
[0067] The bus 418 represents one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example and not limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0068] The computer system / server 412 typically includes a variety of computer system readable media. Such media can be any available media accessible by the computer system / server 412, and it includes volatile and non-volatile media, removable and non-removable media.
[0069] The system memory 428 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 430 and / or cache memory 432. The computer system / server 412 may also include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 434 may be provided for reading from and writing to an non-removable non-volatile magnetic medium (not shown and typically referred to as a "hard disk drive"). Although not shown, a disk drive for reading from or writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a removable non-volatile optical disk such as a CD-ROM, DVD-ROM or other optical media may be provided. In such cases, each may be connected to the bus 418 by one or more data media interfaces. As will be further depicted and described below, the memory 428 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present disclosure.
[0070] A program / utilities 440 having a set (at least one) of program modules 442 can be stored in the memory 428, by way of example and not limitation, as well as an operating system, one or more application programs, other program modules, and program data. Each or some combination of the operating system, one or more application programs, other program modules, and program data can include an implementation of a network environment. The program modules 442 generally execute the functions and / or methods of embodiments of the present disclosure as described herein.
[0071] The computer system / server 412 can also communicate with one or more external devices 414 such as, for example, a keyboard, a pointing device, a display 424, and / or any device that enables a user to interact with the computer system / server 412; and / or any device (e.g., a network card, a modem, etc.) that enables the computer system / server 412 to communicate with one or more other computing devices. Such communication can occur via an input / output (I / O) interface 422. Additionally, the computer system / server 412 can communicate with one or more networks such as, for example, a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet) via a network adapter 420. As depicted, the network adapter 420 communicates with other components of the computer system / server 412 via a bus 418. It should be understood that although not shown, other hardware and / or software components can be used in conjunction with the computer system / server 412. Examples include, but are not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, among others.
[0072] Now referring to Figure 5 , an illustrative cloud computing environment 50 will be described. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 400 with which local computing devices used by cloud consumers can communicate, the local computing devices such as, for example, a personal digital assistant (PDA) or cellular phone 54A, a desktop computer 54B, a laptop computer 54C, and / or an in-vehicle computer system 54N. The nodes 400 can communicate with one another. They can be physically or virtually grouped (not shown) in one or more networks such as, for example, a private cloud, a community cloud, a public cloud, or a hybrid cloud, or a combination thereof. This allows the cloud computing environment 50 to provide infrastructure, platforms, and / or software as services for which cloud consumers do not need to maintain resources on local computing devices. It should be understood that Figure 5 the types of computing devices 54A-N shown in
[0073] Although embodiments of the present invention have been described in detail with reference to exemplary embodiments, those skilled in the art will recognize that various modifications and substitutions can be made thereto without departing from the scope of the invention set forth in the appended claims.
Claims
1. A computer - implemented method for context - aware data mining of text documents, comprising the following steps: Receiving a list of words parsed and pre - processed from an input query, wherein each word in the input query is represented as a one - hot vector, the number of elements in the one - hot vector being equal to the size of a dictionary created from the text document, where all 0s and one 1 in the elements correspond to each word; Using a word embedding model of the text document being queried to calculate a relevant distributed embedding representation for each word in the list of words, wherein the word embedding model is a neural network model that represents each word or row in the dictionary by a vector; Aggregating the relevant distributed embedding representations of all words in the list of words to represent the input query with a single embedding; Retrieving a ranked list of N - line document fragments similar to the aggregated word embedding representation of the query; and Returning the list of retrieved fragments to the user.
2. The method according to claim 1, wherein, The aggregation of the relevant distributed embedding representations is performed with either the average of all relevant distributed embedding representations or the maximum of all relevant distributed embedding representations.
3. The method according to claim 1, wherein, N is a positive integer provided by the user.
4. The method according to claim 1, further comprising training a word embedding model of the text document, comprising the following steps: Parsing and pre - processing the text document and generating a list of tokenized words; Defining a word dictionary from the list of tokenized words, wherein the word dictionary includes at least some tokens from the list of tokenized words; and Training the word embedding model.
5. The method according to claim 4, wherein Parsing and pre - processing the text document includes the following steps: Removing all punctuation and leading from each line in the text document; Parsing numerical data; Tokenizing the text document by words to form a list of tokenized words, wherein a token is either a single word, an N - gram of N consecutive words, or an entire line of the document; and Returning the list of tokenized words.
6. The method according to claim 5, wherein, The text document is a computer system log, and the numerical data includes decimal numbers and hexadecimal addresses.
7. The method according to claim 1, further comprising parsing and pre - processing the input query by the following steps: Removing all punctuation from the input query; Parsing numerical data; Tokenizing the input query by words to produce a list of tokenized words, wherein a token is either a single word, an N - gram of N consecutive words, or an entire line of the input query; and Returning the list of tokenized words.
8. The method according to claim 1, wherein Retrieving a ranked list of N - line document fragments similar to the aggregated word embedding representation of the query includes: comparing the aggregated word embedding representation of the query with the word embedding model of the text document using a similarity metric, returning those fragments of the word embedding model of the text document whose similarity to the aggregated word embedding representation of the query is greater than a predetermined threshold, and ranking the retrieved document fragments by similarity.
9. A computer - readable program storage device tangibly embodying a program of instructions executable by a computer to perform the method steps for context - aware data mining of text documents, the method comprising the following steps: Receive a list of words parsed and preprocessed from an input query, where each word in the input query is represented as a one-hot vector, the number of elements in the one-hot vector being equal to the size of a dictionary created from the text documents, where all 0s and one 1 in the elements correspond to each word; Use a word embedding model of the text document being queried to compute a relevant distributed embedding representation for each word in the list of words, where the word embedding model is a neural network model that represents each word or row in the dictionary by a vector; Aggregate the relevant distributed embedding representations of all words in the list of words by one of the average of all relevant distributed embedding representations or the maximum of all relevant distributed embedding representations to represent the input query by a single embedding; Retrieve a ranked list of N document fragments that are similar to the aggregated word embedding representation of the query, where N is a positive integer provided by the user; and Return the list of retrieved fragments to the user.
10. The computer-readable program storage device according to claim 9, wherein, The method further includes training a word embedding model of the text document, including the following steps: Parse and preprocess the text document and generate a list of tokenized words; Define a word dictionary from the list of tokenized words, where the word dictionary includes at least some tokens from the list of tokenized words; and Train the word embedding model, where parsing and preprocessing the text document includes the following steps: Remove all punctuation and leading from each line in the text document; Parse numeric data; Tokenize the text document by words to form a list of tokenized words, where a token is one of a single word, an N-gram of N consecutive words, or an entire line of the document; and Return the list of tokenized words.
11. The computer-readable program storage device according to claim 10, wherein, The text document is a computer system log, and the numeric data includes decimal numbers and hexadecimal addresses.
12. The computer-readable program storage device according to claim 9, wherein, The method further includes parsing and preprocessing the input query by the following steps: Remove all punctuation from the input query, Parse numeric data; Tokenize the input query by words to produce a list of tokenized words, where a token is one of a single word, an N-gram of N consecutive words, or an entire line of the input query; and Return the list of tokenized words.
13. The computer-readable program storage device according to claim 9, wherein, Retrieving a ranked list of N document fragments that are similar to the aggregated word embedding representation of the query includes: comparing the aggregated word embedding representation of the query with the word embedding model of the text document using a similarity metric, returning those fragments of the word embedding model of the text document whose similarity to the aggregated word embedding representation of the query is greater than a predetermined threshold, and ranking the retrieved document fragments by similarity.
Citation Information
Patent Citations
Retrieval mode generation method and device based on search engine
CN106777191A
Text similarity measure method combining word clustering and word combination semantic features
CN108399163A