Automatic summarization of the content of electronic messages
By implementing context-based automatic summary technology in an email server or independent automatic summary, the problems of dynamic and context dependence of email content are solved, and accurate and relevant automatic summary of email content is achieved, improving the user experience.
Patent Information
- Application Number
- CN202080009979.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-01-21
- Filing Date
- 2020-01-04
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-01-04
AI Technical Summary
It is difficult for prior art to effectively and automatically summarize content in emails, especially due to the dynamicity and contextual dependence of email content.
Implement context-based automatic summary techniques in an email server or standalone automatic summaryr, including using machine learning to detect templated messages, extract and abstract techniques to generate summary, and generate the most relevant summary through clustering and correlation score calculations.
Accurate and relevant automatic summary of email content, improve the user experience of automatic summary applications and services, and can capture the main topics in emails and adapt to their dynamic characteristics.
Smart Images

Figure CN113316775B_ABST
Abstract
Description
Technical Field
[0001] The present application relates generally to electronic message processing, and more particularly to methods and computing devices for automatic summarization of content in electronic messages. Background Art
[0002] Automatic summarization is the process of shortening the original text using software to create a summary with the main points of the original text. Techniques that are able to make a coherent summary take into account variables such as length, writing style, and grammar. Two techniques used for automatic summarization include extraction and abstraction. Extraction techniques select a subset of existing words, phrases, or sentences in the original text to form a summary. In contrast, abstraction techniques are able to build an internal semantic representation of the original text and then use natural language generation to create a summary that is closer to what a human might express. Summary of the invention
[0003] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0004] Although extraction and abstraction techniques may be sufficient to automatically summarize documents such as news reports, scientific papers, etc., such techniques may not be sufficient to summarize the content of electronic messages (e.g., emails) exchanged between users using computers, smart phones, or other suitable types of computing devices. One of the reasons for such deficiencies is that the state of the content in the email may be dynamic, i.e., changing over time. For example, statements for earlier events may not be as relevant now as they were then. Another reason may be that the relevance of the content in the email may depend on the context of the corresponding email communication. For example, a short statement from a manager in an organization may be more relevant than multiple long statements published by his / her subordinates.
[0005] Several embodiments of the disclosed technology can solve at least some of the above difficulties by realizing the context-based automatic summary of the content in the electronic message of email and / or other suitable types.In one implementation, the email server can be configured to receive emails sent to users or generated by users.After receiving the email, the automatic summarizer at the email server (or otherwise accessing the received email) can be configured to perform the first stage and the second stage of the summary processing of the received email, as described in more detail below.In another implementation, the automatic summarizer can be independent of the email server.In other implementations, the automatic summarizer can also be configured to perform text-to-speech conversion of the automatically generated summary of the email, insert the summary into the email or other suitable operations.
[0006] In certain embodiments, at the first stage of the summary process, the automatic summarizer can be configured to determine whether the received email is a templated message. Without being bound by theory, it is believed that the emails exchanged in enterprises, government offices, schools or other suitable types of organizations are often templated messages. Exemplary templated messages can include messages related to going out (OOF), personal leave, working from home (WFH), meeting invitations, automatic responses, status updates, welcome speeches, meeting notes, etc. In one implementation, the automatic summarizer can be configured to detect such templated messages using, for example, training data sets via machine learning. In this way, a machine learning model can be developed to include keywords / key phrases (e.g., "OOF", "WFH", etc.) indicating templated messages. In other implementations, the automatic summarizer can also be configured to use pre-configured message templates provided by an administrator or detect such templated messages via other suitable technologies.
[0007] The aforementioned templated messages can be effectively summarized using pre-configured summary templates. For example, a summary template for an OOF message can include "[sender] was OOF from [date / time 1] to [date / time 2]," where the parameters in parentheses (e.g., "sender") represent variables. Upon determining that the received email is an OOF message, the automatic summarizer can be configured to extract a value for [sender] by, for example, identifying a name in the "from" field in the title of the received email (e.g., "Anand"). The automatic summarizer can also be configured to identify a first date / time (e.g., "December 11, 2018") and a second date / time (e.g., "December 31, 2018") based on, for example, the format of the text in the received email. The automatic summarizer can be configured to subsequently write a summary by replacing the summary template with the identified sender and date / time, such as "Anand was OOF from December 11, 2018 to December 31, 2018."
[0008] When it is determined that the received email is not a templated message, the automatic summarizer can be configured to perform the second stage of summarization processing based on at least the content in the email body of the received email. In an exemplary implementation, the automatic summarizer can be configured to initially extract entity values (e.g., sender name, (one or more) recipient names, date / time of sending / receiving, etc.) and text or other suitable types of content from the email body. The automatic summarizer can be configured to subsequently decompose the content from the email body based on one or more machine learning models to classify individual sentences (or parts thereof) into different categories of statements. In some embodiments, exemplary categories of statements can include facts, reasoning, judgments, and decisions. For example, a statement of fact can be a statement of "our system crashed last night." A statement of reasoning can be "there must be an error in the code." A statement of judgment can be "our system is the worst," and a statement of decision can be "please contact the development team to fix it as soon as possible." In other embodiments, the automatic summarizer can also classify emails into categories of truth, evidence, reasoning, requests, or other suitable types.
[0009] The classification developer can be configured to generate one or more machine learning models by analyzing a set of emails of a user using a "neural network" or "artificial neural network" configured to "learn" or gradually improve task performance by learning from known examples. In some implementations, the neural network can include multiple layers of objects, generally referred to as "neurons" or "artificial neurons". Each neuron can be configured to perform a function, such as a nonlinear activation function, based on one or more inputs via corresponding connections. Artificial neurons and connections typically have contribution values that are adjusted as learning proceeds. The contribution value increases or decreases the strength of the input at the connection. Typically, artificial neurons are organized in layers. Different layers can perform different types of transformations on their respective inputs. Signals typically travel from an input layer to an output layer, possibly after traversing one or more intermediate layers. Therefore, by using a neural network, the classification developer can provide a set of classification models, and the automatic summarizer can use the set of classification models to classify sentences in received emails.
[0010] After the decomposition of the content in the email body of the received email is completed, the automatic summarizer can be configured to assign relevance scores to each classified statement based on the entity that made the statement, the recency of the statement, and / or other suitable criteria. For example, the automatic summarizer can be configured to determine the relevance scores of the statements made by the person based on his / her position in the organization by consulting an organizational chart, the person's position, etc. In this way, the statements made by the manager can have a higher relevance score than the statements made by his / her subordinates. In other examples, the automatic summarizer can be configured to assign a higher relevance score to the most recently made statements than to the previously made statements. In additional examples, the automatic summarizer can be configured to assign relevance scores based on the subject of the statement or other suitable criteria.
[0011] The automatic summarizer can also be configured to determine the context of statements of facts, reasoning, judgments, and decisions by clustering the statements based on their relative proximity, for example, according to a hierarchy of categories. For example, statements of facts, reasoning, and judgments close to a decision can be clustered around the decision, while other statements of facts, reasoning, and judgments close to another decision can be clustered around the other decision. In some embodiments, proximity can be based on a preset proximity threshold, such as the number of characters, words, sentences, etc. In other embodiments, the proximity threshold can be based on syntactic structures, such as punctuation, paragraphs, sections, etc. In other embodiments, the proximity threshold can be based on other suitable criteria.
[0012] In some scenarios, the received email may not contain any statements classified as decisions. In such scenarios, embodiments of the disclosed technology can include clustering the statements in the received email according to a hierarchy of categories from decisions, judgments, reasoning to facts. For example, when there are no statements of decisions in the received email, clustering can be performed around one or more statements of judgments. When there are no statements of decisions or judgments, clustering can be performed around one or more statements of reasoning. When the received email contains only statements of facts, automatic summarization can be based on individual facts.
[0013] Once the statements are clustered, the automatic summarizer can be configured to calculate a cluster score based on the assigned relevance scores of the individual statements in each cluster. In one example, the cluster score can be the sum of all relevance scores assigned to the statements belonging to the cluster. In another example, the cluster score can be the sum of all relevance scores assigned to the statements belonging to the cluster, and is biased based on the age of the statement, the number of recipients to which the statement is intended, or other suitable parameters of the individual statements. In any of the foregoing examples, the calculated cluster scores can be normalized based on a scale of, for example, zero to one hundred or other suitable range of values.
[0014] Based on the calculated clustering scores, the automatic summarizer can be configured to select multiple (e.g., one, two, three, etc.) clusters based on the calculated clustering scores, and apply extraction and / or abstraction techniques to generate the number of suggested summaries of the received emails. In some embodiments, the generated summaries can be output, for example, via a user interface, for the user to select as the subject or summary of the received email. In other embodiments, the generated summary with the highest clustering score can be automatically selected to be output to the user, for example, via a text-to-speech engine to convert the generated summary into a voice message. Then, the voice message can be played to the user via, for example, a smart phone or other suitable type of computing device.
[0015] Therefore, several embodiments of the disclosed technology can effectively perform automatic summarization of the content in emails and other types of electronic messages via the above-mentioned classification technology. Without being bound by theory, it is believed that clustering sentences according to the hierarchy of decisions, judgments, reasoning, and facts in the received emails can effectively capture the main topics contained in the received emails. In addition, by considering the dynamic characteristics of emails and other types of electronic messages and the sources of various sentences, the relevant topics contained in the received emails can be accurately captured and presented to the user. In this way, compared with other technologies, the user experience of automatic summary applications and / or services can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1A and Figure 1B is a schematic diagram illustrating a computing system that implements automatic summarization of content in an electronic message in accordance with an embodiment of the disclosed technology.
[0017] Figure 2 is a diagram illustrating an embodiment according to the disclosed technology Figure 1A and Figure 1B A schematic diagram of certain hardware / software components of a computing system.
[0018] Figures 3A-3C is a schematic diagram illustrating sentence clustering according to an embodiment of the disclosed technology.
[0019] Figures 4A-4C is a flow chart illustrating an exemplary process for automatic summarization of content in an electronic message according to an embodiment of the disclosed technology.
[0020] Figure 5 A computing device suitable for use with certain components of the computing system in FIG. 1 . DETAILED DESCRIPTION
[0021] The following describes certain embodiments of systems, devices, components, modules, routines, data structures, and processes for automatically summarizing the content of electronic messages in a computing system. In the following description, specific details of the components are included to provide a thorough understanding of certain embodiments of the disclosed technology. Those skilled in the relevant art will also understand that the technology is capable of additional embodiments. The technology can also be used without the following references. Figure 1A-5 The described embodiments are practiced without the need for several details.
[0022] As used herein, the term "email server" generally refers to a computer dedicated to running an application that is configured to receive incoming email from senders and forward outgoing email to recipients via a computer network (e.g., the Internet). Examples of such applications include Microsoft The email server can maintain and / or access one or more inboxes for the corresponding users. As used herein, an "inbox" is a folder configured to contain data representing incoming emails for a user. The email server can also maintain and / or access one or more draft folders and / or outboxes and / or other suitable mailboxes configured to store outgoing emails.
[0023] Also as used herein, "fact" generally refers to a statement containing information that is represented as having objective reality. For example, an exemplary fact could be the statement "Our system crashed last night." "Reasoning" generally refers to a statement containing a reasoned opinion formed due to known facts or evidence. An exemplary reasoning could be "There must be a bug in the code." "Judgment" generally refers to a statement containing an authoritative opinion. An exemplary judgment could be "Our system is the worst." "Decision" generally refers to a statement containing a call to an action and / or determination made during and / or after consideration. An exemplary decision could be "Please contact the development team to fix it as soon as possible."
[0024] Extraction and abstraction techniques may be sufficient to automatically summarize text in news reports, scientific papers, etc., but such techniques may not be sufficient to summarize the content of electronic messages (e.g., emails) exchanged between users. One reason for such inadequacies is that the state of the content in emails may change over time. For example, statements for earlier events may not be as relevant now as they were then. Another reason may be that the relevance of the content in an email may depend on the context of the corresponding email communication. For example, a short statement from a manager in an organization may be more relevant than multiple long statements published by his / her subordinates.
[0025] Several embodiments of the disclosed technology are directed to implementing context-based automatic summarization to efficiently perform automatic summarization of content in electronic messages. Specifically, aspects of the disclosed technology are directed to classifying individual sentences (or portions thereof) in the body of an email into different categories of statements, such as facts, reasoning, judgments, and decisions. Relevance scores can then be assigned to individual statements before the classified statements are clustered according to the hierarchy of decisions, judgments, reasoning, and facts. Cluster scores can then be calculated for the individual clusters. Based on the cluster scores, one or more clusters of statements can be selected to generate an appropriate summary and / or topic for the electronic message, as described below with reference to Figure 1A-5 Described in more detail.
[0026] Figure 1A is a schematic diagram illustrating a computing system 100 that implements automatic summarization of content in an electronic message according to an embodiment of the disclosed technology. Figure 1A As shown in , computing system 100 can include a computer network 104 that interconnects client devices 102 with one or more email servers 106 (referred to herein for simplicity as “email servers 106”). Email servers 106 are also interconnected with network storage 112 including one or more inboxes 114 and data store 108 including classification models 110. Computer network 104 can include an intranet, a wide area network, the Internet, or other suitable type of network. Although Figure 1A Specific components of computing system 100 are shown in FIG. 1 , but in other embodiments, computing system 100 can also include additional and / or different components or arrangements. For example, computing system 100 can also include additional network storage devices, additional hosts, and / or other suitable components (not shown). In other embodiments, network storage device 112 and / or data storage 108 can be integrated into email server 106.
[0027] The client devices 102 can each include a computing device that facilitates the corresponding user 101 to access computing services provided by the email server 106 via the computer network 104. For example, in the illustrated embodiment, the client devices 102 individually include smart phones and desktop computers. In other embodiments, the client devices 102 can also include laptops, tablet computers, game consoles, or other suitable computing devices. Although for the purpose of illustration, the client devices 102 are not shown in FIG. Figure 1A and Figure 1B A first user 101a and a second user 101b are shown in FIG. 1 , but in other embodiments, the computing system 100 can facilitate any suitable number of users 101 accessing suitable types of computing services provided by the email server 106 .
[0028] Email server 106 can be configured to facilitate email reception, storage, forwarding, and other related functions. Figure 1A As shown in , the first user 101a can generate an email 116 using the client device 102 directly or via other intermediate email servers (not shown) and send it to the email server 106. The email 116 is sent to the second user 101b and can include an email header 117, an email body 118, and one or more optional attachments (not shown). The email header 117 can include various fields, such as "From:", "To:", "Cc:", "Bcc:", etc. The email body 118 can include text arranged in sentences, paragraphs, sections, etc. and / or other suitable types of content. Upon receiving the email 116 from the first user 101a, the email server 106 can store a copy of the email 116 in the inbox 114 on the network storage device 112 corresponding to the second user 101b.
[0029] As in Figure 1A As shown in , the computing system 100 can include a classification developer 130 and an automatic summarizer 132, which are operably coupled to each other for automatically summarizing the content of the email 116 exchanged between the first user 101a and the second user 101b. Figure 1A In the illustrated example, the taxonomy developer 130 and the automatic summarizer 132 are components of the email server 106. In other examples, at least one of the taxonomy developer 130 and / or the automatic summarizer 132 can be a component hosted on one or more additional servers (not shown) that are separate from the email server 106 while still having access to the emails 116 in the inbox 114 at the network storage device 112.
[0030] According to an embodiment of the disclosed technology, the classification developer 130 can be configured to develop one or more classification models 110 that can be used to classify sentences in the email 116 via machine learning. For example, the classification developer 130 can be configured to generate one or more classification models 110 by analyzing a set of emails of the user 101 using a "neural network" or "artificial neural network" that is configured to "learn" or gradually improve the performance of a task by learning known examples. In some implementations, the neural network can include multiple layers of objects generally referred to as "neurons" or "artificial neurons". Each neuron can be configured to perform a function, such as a nonlinear activation function, based on one or more inputs via corresponding connections. Artificial neurons and connections typically have weight values that are adjusted as learning proceeds. The weight value increases or decreases the strength of the input at the connection. Typically, artificial neurons are organized in layers. Different layers can perform different types of transformations on their respective inputs. A signal typically travels from an input layer to an output layer, possibly after traversing one or more intermediate layers. Thus, by using neural networks, the classification developer 130 can provide one or more classification models that the automatic summarizer 132 can use to classify sentences in the received email 116 .
[0031] The automatic summarizer 132 can be configured to generate a suggested summary 119 of the content in the email 116 using the classification model 110. The automatic summarizer 132 can then provide the suggested summary 119 to the first user 101a for selection. The first user 101a can then select a summary 119' from the suggested summaries 119. In response to the selection, the automatic summarizer 132 (or other suitable component of the email server 106) can be configured to insert the selected summary 119' into the received email 116 before sending the email 116' to the second user 101b.
[0032] As in Figure 1AAs shown in , upon receiving the email 116', the email client 124 on the client device 102 can display the received email 116' as a message in the inbox of the second user 101b. For example, the exemplary email 116' can include a title 117 including the sender's name (i.e., "Jane Doe"), a subject line including a selected summary 119' (e.g., "Project Progress Summary"), and an email body 118 including exemplary text (such as "This week..."). Thus, the first user 101a can utilize the automatically generated summary 119' to effectively compose the email 116 to the second user 101b. In this way, the usability of the email service provided by the email server 106 can be improved even when the first user 101a is using speech-to-text conversion to compose the email 116, or otherwise does not have access to readily available typing tools.
[0033] In some embodiments, the automatic summarizer 132 can be configured to perform a first stage and a second stage of summarization processing on the received email 116. The first stage of summarization processing can include a template-based processing stage. The second stage of summarization processing can include: classifying sentences in the email body 118 based on the classification model 110, clustering the classified sentences according to categories, calculating cluster scores, and selecting clusters to automatically generate suggested summaries 119 for selection by the first user 101. Figure 2 Exemplary components and operations of automatic summarizer 132 are described in more detail.
[0034] In other embodiments, the automatic summarizer 132 can also be configured to perform other suitable operations. Figure 1B As shown in FIG. 1 , in response to a selection by the first user 101a, the automatic summarizer 132 can also be configured to convert the selected summary 119′ ( Figure 1A ) is converted into a voice message 120, and the voice message 120 is stored in the inbox 114 of the second user 101b in the network storage device 112. Based on the request of the second user 101b or in other suitable ways, the email server 106 can be configured to provide the generated voice message 120 to the client device 102 of the second user 101b. In turn, the client device 102 can be configured to play the voice message 120 containing the selected summary 119' to the second user 101b via, for example, the speaker 103.
[0035] Figure 2 is a schematic diagram illustrating certain hardware / software components of a computing system 100 according to an embodiment of the disclosed technology. Figure 2 For the sake of clarity, only Figure 1Aand Figure 1B Certain components of the computing system 100. Figure 2 In the Figures and in other Figures herein, individual software components, objects, classes, modules, and routines may be computer programs, procedures, or processes written as source code in C, C++, C#, Java, and / or other appropriate programming languages. Components may include, but are not limited to, one or more modules, objects, classes, routines, properties, processes, threads, executables, libraries, or other components. Components may be in source or binary form. Components may include aspects of source code prior to compilation (e.g., classes, properties, procedures, routines), compiled binary units (e.g., libraries, executables), or artifacts (e.g., objects, processes, threads) that are instantiated and used at runtime.
[0036] Components within a system can take different forms within the system. As an example, a system including a first component, a second component, and a third component can encompass, without limitation, a system in which the first component is a property in source code, the second component is a binary compiled library, and the third component is a thread created at runtime. A computer program, procedure, or process can be compiled into object, intermediate, or machine code and presented for execution by one or more processors of a personal computer, a network server, a laptop, a smart phone, and / or other suitable computing device.
[0037] Likewise, a component may include a hardware circuit. One of ordinary skill in the art will recognize that hardware may be considered to be petrochemical software, while software may be considered to be liquefied hardware. As just one example, the software instructions in a component may be burned into a programmable logic array circuit, or may be designed as a hardware circuit with an appropriate integrated circuit. Likewise, hardware may be simulated by software. Various implementations of source code, intermediate code, and / or object code and associated data may be stored in a computer memory, including a read-only memory, a random access memory, a disk storage medium, an optical storage medium, a flash memory device, and / or other suitable computer-readable storage medium, excluding propagation signals.
[0038] As in Figure 2 As shown in , the email server 106 can include a classification developer 130 and an automatic summarizer 132. Although the classification developer 130 and the automatic summarizer 132 are Figure 2 106, but in other embodiments, the classification developer 130 can be provided by one or more other online or offline servers (not shown) separate from the email server 106. In other embodiments, the email server 106 can be included in Figure 2 Additional and / or different components not shown.
[0039] The classification developer 130 can be configured to generate the classification model 110 via various machine learning techniques based on a data set including previous emails 116″ and associated sentence classes 122 and user input 115 regarding a summary of suggestions. The sentence classes 122 can be generated manually, automatically via unstructured learning, or via other suitable techniques. In one implementation, the classification developer 130 can be configured to use a neural network including multiple layers of objects, which are often referred to as “neurons” or “artificial neurons”, to perform machine learning based on the data set of emails 116″, as described above with reference to Figure 1A As described. By using a neural network, the classification developer 130 can provide a set of classification models 110, and the automatic summarizer 132 can use the set of classification models 110 to classify additional received emails 116. In one example, the classification model 110 can include various values of variables related to the email body 118. Exemplary variables can include keywords or key phrases (e.g., "maybe", "must be", etc.), grammar (e.g., verbs before nouns and adjectives), sentence structure (e.g., subjects followed by verbs and nouns), and other suitable content parameters. Thus, an exemplary classification model 110 can include an indication of a decision conditional on sentences having verbs before any nouns (e.g., "Please contact the development team to fix it as soon as possible"). In other examples, the classification model 110 can have other suitable conditions and indications. In the illustrated embodiment, the classification developer 130 provides the classification model 110 for storage in the data store 108. In other embodiments, the classification developer 130 can provide the classification model 110 directly to the automatic summarizer 132 or store the classification model 110 in other suitable locations.
[0040] As in Figure 2 As shown in , the automatic summarizer 132 can include a template processor 133, a classifier 134, a cluster generator 136, a summary generator 138, and a feedback processor 139 that are operatively coupled to each other. Figure 2 Specific components or modules of the automatic summarizer 132 are shown in FIG, but in other embodiments, the automatic summarizer 132 can also include an interface, a network, or other suitable types of components and / or modules. In other embodiments, at least one of the above components can be provided by an external application / server separate from the automatic summarizer 132.
[0041] In certain embodiments, at the first stage of the summary process, the template processor 133 of the automatic summarizer can be configured to determine whether the received email 116 from the first user 101a is a templated message. Without being bound by theory, it is believed that emails exchanged in businesses, government offices, schools, or other suitable types of organizations are often templated messages. Exemplary templated messages can include messages related to out of office (OOF), personal leave, work from home (WFH), meeting invitations, automatic responses, status updates, welcome speeches, meeting notes, etc. In one implementation, the template processor 133 can be configured to detect such templated messages via machine learning using, for example, a training data set with email 116". In other implementations, the template processor 133 can also be configured to detect such templated messages using pre-configured message templates provided by an administrator (not shown) or via other suitable techniques.
[0042] The template processor 133 can be configured to efficiently summarize templated messages using pre-configured summary templates. For example, a summary template for an OOF message can include "[Sender] was OOF from [Date / Time 1] to [Date / Time 2]," where the parameters in parentheses (e.g., "Sender") represent variables. Upon determining that the received email is an OOF message, the template processor 133 can be configured to extract the value of [Sender] by, for example, identifying a name in the "From" field in the header of the received email (e.g., "Anand"). The template processor 133 can also be configured to identify a first date / time (e.g., "December 11, 2018") and a second date / time (e.g., "December 31, 2018") based on, for example, the format of the text in the received email. The template processor 133 can be configured to then compose the summary 119 by substituting the identified sender and date / time into the summary template, such as "Anand was OOF from December 11, 2018 to December 31, 2018."
[0043] When it is determined that the received email 116 is not a templated message, the template processor 133 can be configured to forward the processing to the classifier 134 to perform the second stage of the summary processing based on at least the content in the email body 118 of the received email 116. In an exemplary implementation, the classifier 134 can be configured to initially extract entity values (e.g., sender name, (one or more) recipient names, send / receive date / time, etc.) and text or other suitable types of content from the email body 118. The classifier 134 can be configured to then decompose the content from the email body 118 based on one or more classification models 110 from the data store 108 to classify individual sentences (or parts thereof) in the email body 118 into different categories of statements. In some embodiments, exemplary categories of statements can include facts, reasoning, judgments, and decisions. For example, a factual statement can be a statement of "our system crashed last night." A reasoning statement can be "there must be an error in the code." A judgment statement could be "our system is the worst," and a decision statement could be "please contact the development team to fix it as soon as possible." In other embodiments, the automatic summarizer can also classify emails into categories of truth, evidence, reasoning, request, or other suitable types.
[0044] After completing the parsing of the content in the email body 118 of the received email 116, the classifier 134 can be configured to assign a relevance score to each of the classified statements based on the entity that made the statement, the recency of the statement, and / or other appropriate criteria. For example, the classifier 134 can be configured to determine the relevance score of statements made by a person based on his / her position in the organization by consulting an organizational chart, the person's position, etc. In this way, statements made by a manager can have a higher relevance score than those made by his / her subordinates. For example, in Figure 2 In the example shown in , statements by the second user 101b and other users 101n, shown as subordinates of the first user 101a, will have lower relevance scores than those made by the first user 101a. In other examples, the classifier 134 can be configured to assign a higher relevance score to a more recently made statement than to another previously made statement. In additional examples, the classifier 134 can be configured to assign the relevance score based on the subject matter of the statement or other suitable criteria.
[0045] The automatic summarizer 132 can also be configured to determine the context of statements of facts, reasoning, judgments, and decisions by clustering the statements using the cluster generator 136 based on the relative proximity of the statements, for example, according to the hierarchy of categories. For example, statements of facts, reasoning, and judgments close to a decision can be clustered around the decision, while other statements of facts, reasoning, and judgments close to another decision can be clustered around other decisions. In some embodiments, the proximity can be based on a preset proximity threshold, such as the number of characters, words, sentences, etc. In other embodiments, the proximity threshold can be based on syntactic structures, such as punctuation, paragraphs, sections, etc. In other embodiments, the proximity threshold can be based on other suitable criteria.
[0046] In some scenarios, the received email may not contain any statements classified as decisions. In such scenarios, the cluster generator 136 can be configured to cluster the statements in the received email 116 according to a hierarchy of categories from decision, judgment, reasoning to fact. For example, when there are no decision statements in the received email, clustering can be performed around one or more judgment statements. When there are no decision or judgment statements, clustering can be performed around one or more reasoning statements. When the received email contains only factual statements, automatic summarization can be based on individual facts. The following reference Figures 3A-3C An exemplary cluster 140 is described in more detail in Figures 3A-3C ).
[0047] Once the statements are clustered, the cluster generator 136 can be configured to calculate cluster scores for individual clusters 140 based on the assigned relevance scores of the individual statements in each cluster. In one example, the cluster score can be the sum of all relevance scores assigned to the statements belonging to the cluster. In another example, the cluster score can be the sum of all relevance scores assigned to the statements belonging to the cluster and be biased based on the age of the statements, the number of recipients to which the statements are directed, or other suitable parameters of the individual statements. In any of the foregoing examples, the calculated cluster scores can be normalized based on a scale of, for example, zero to one hundred or other suitable range of values.
[0048] Based on the calculated clustering scores, the cluster generator 136 can be configured to sort and select multiple (e.g., one, two, three, etc.) clusters 140 based on the calculated clustering scores (or other appropriate criteria), and forward the selected clusters 140 to the summary generator 138 for further processing. The summary generator 138 can be configured to apply extraction and / or abstraction techniques to generate multiple suggested summaries 119 for the received email 116. In some embodiments, the generated summaries 119 can be output, for example, via a user interface (not shown), for user selection as a subject or summary of the received email 116. In other embodiments, the generated summary 119 with the highest clustering score can be automatically selected for output to the first user 101a, for example, via a text-to-speech engine to convert the generated summary into a voice message 120 (in Figure 1B ). The voice message 120 can then be played to the user 101 via, for example, a smart phone or other suitable type of computing device. In further embodiments, after the user 101 selects one of the suggested summaries 119, the summary generator 138 can insert the selected summary 119 into an email 116 stored at an inbox 114 at the network storage device 112. In certain implementations, the feedback processor 139 can be configured to receive user input 115 regarding the relevance of the suggested summaries 119. In response to receiving the user input 115, the classifier 134 can be configured to reassign a relevance score to each classified statement; the summary generator 138 can regenerate the suggested summaries, or perform other suitable operations in the automatic summarizer 132.
[0049] Several embodiments of the disclosed technology are therefore able to effectively perform automatic summarization of the content in emails 116 and other types of electronic messages via the aforementioned classification technology. Without being bound by theory, it is believed that clustering statements according to the hierarchy of decisions, judgments, reasoning, and facts in the received emails 116 can effectively capture the main topics contained in the received emails. In addition, by considering the dynamic characteristics of emails 116 and other types of electronic messages and the sources of various statements, the relevant topics contained in the received emails can be accurately captured and displayed to the user. For example, the generated summary can be changed as new conversations are added and new users 101 are added or removed. The generated summary based on the same email can also change based on the person who finds the generated summary. For example, a manager's view of the generated summary may be different from that of any of the manager's subordinates. In this way, compared with other technologies, the user experience of automatic summary applications and / or services can be improved.
[0050] Figures 3A-3Cis a schematic diagram illustrating sentence clustering according to an embodiment of the disclosed technology. Figure 3A As shown in , cluster 140 can include decision 141 and one or more facts 142a and 142b, reasoning 144, and judgment 146 associated with decision 141, as represented by edge 143. As described above with reference to Figure 1A As described, when the email body 118 does not include any decision 141, clusters 140 can be generated based on decisions 146, such as in Figure 3B When the email body 118 does not include any decision 141 or judgment 146, clusters 140 can be generated based on reasoning 144, as shown in Figure 3C Although for illustrative purposes Figures 3A-3C A specific number of decisions, judgments, inferences, and facts are shown in FIG. 1 , but in other examples, each cluster 140 can include one of decisions, judgments, or inferences surrounded by any suitable number of statements of other categories.
[0051] Figures 4A-4C is a flow chart illustrating an exemplary process for automatically summarizing content in an electronic message according to an embodiment of the disclosed technology. Figure 1A and Figure 1B The process is described with reference to computing system 100 of FIG. 1 , but in other embodiments, the process can also be implemented in a computing system having additional and / or different components.
[0052] As in Figure 4A As shown in , process 200 can include receiving an email at stage 202. Process 200 can then include a decision stage 204 to determine whether the received email is a templated message. In one embodiment, the determination can be based on a template model developed using machine learning. In other embodiments, the determination can also be based on a message template provided by, for example, a manager or other suitable entity. In response to determining that the received email is a templated message, process 200 can continue at stage 206 to generate a summary of the received email based on a summary template. In some embodiments, generating a summary can include identifying an entity-specific summary template. For example, the sales department may have a different summary template than the finance department. Figure 4B In response to determining that the received email is not a templated message, process 200 can continue to perform classification-based summary processing at stage 208. Figure 4CThe exemplary operation of performing the classification-based summary process is described in more detail. Process 200 can also optionally include learning a new summary template at stage 211. The new summary template can be based on the generated summary from stage 210 or from other suitable sources. The new summary template can then be used to generate a summary based on the summary template in stage 206.
[0053] As in Figure 4B As shown in , an exemplary operation of generating a summary based on a summary template can include identifying a summary template corresponding to the received email at stage 212. In some embodiments, identifying the summary template can include determining whether there are any user or entity specific templates. In response to determining that there is a user or template entity. If there is a user or entity summary template, the operation can identify and / or select a user or entity specific summary template. Otherwise, the operation can include identifying or selecting a general summary template. Then, the operation can include extracting template values from the received email at stage 214. Exemplary extracted template values can include the sender's name, date / time, location, or other suitable information. Then, the operation can include inserting the extracted template values into the identified summary template to generate a summary of the email at stage 216. In some implementations, the summary template can be based on a user profile. For example, the email can be identified as corresponding to the template "leave application". If the user has a specific way of providing the subject line of the email, generating the summary can include generating the summary using the subject line structure used by the user in the email. The operation can also include receiving user feedback on the generated summary at stage 217. Based on the received user feedback, operations can include designating the identified summary template as a user- or entity-specific summary template at stage 212 , or performing other suitable operations to explore new user- or entity-specific summary templates.
[0054] As in Figure 4C As shown in , exemplary operations for performing classification-based summarization processing can include an optional stage of aggregating emails based on subject or other suitable attributes at stage 218. For example, vectorization or email / conversation fragments can be used to group emails with similar conversations together. The operations can also include classifying sentences in the body of the email at stage 220. Figure 2 An exemplary technique for classifying statements is described. Operations can then include assigning a relevance score to each classified statement at stage 222. Operations can then include clustering the classified statements at stage 224. Figures 3A-3CExample clusters are described. Operations can then include calculating a cluster score for each cluster at stage 226, and ranking the clusters based on one or more of the calculated scores, the recency of the emails in the cluster, or the user profiles of the authors of the emails in the organization. Such cluster rankings can be used to select the top five, top three, or other suitable number of clusters to be included in the final summary. Operations can also include generating a summary of one or more selected clusters at stage 228, as described above with reference to Figure 2 The operations described can also include collecting user feedback on the generated summary at stage 230. The collected user feedback can then be used to adjust the cluster scores and / or cluster rankings at stages 226 and 227, respectively.
[0055] Figure 5 It is applicable to Figure 1A and Figure 1B The computing device 300 may be adapted to include certain components of the computing system 100. For example, the computing device 300 may be adapted to Figure 1A The email server 106 or client device 102 of FIG. 302. In a very basic configuration 302, the computing device 300 can include one or more processors 304 and a system memory 306. A memory bus 308 can be used to communicate between the processor 304 and the system memory 306.
[0056] Depending on the desired configuration, the processor 304 can be of any type, including but not limited to: a microprocessor (μR), a microcontroller (μC), a digital signal processor (DSP), or any combination thereof. The processor 304 can include a multi-level cache (such as a level 1 cache 310 and a level 2 cache 312), a processor core 314, and registers 316. The exemplary processor core 314 can include an arithmetic logic unit (ALU), a floating point unit (FPU), a digital signal processing core (DSP core), or any combination thereof. An exemplary memory controller 318 can also be used with the processor 304, or in some implementations, the memory controller 318 can be an internal part of the processor 304.
[0057] Depending on the desired configuration, system memory 306 can be of any type including, but not limited to, volatile memory (such as RAM), non-volatile memory (such as ROM, flash memory, etc.), or any combination thereof. System memory 306 can include an operating system 320, one or more applications 322, and program data 324. The depicted basic configuration 302 is illustrated by those components within the inner dashed line.
[0058] The computing device 300 can have additional features or functions and additional interfaces to facilitate communication between the basic configuration 302 and any other devices and interfaces. For example, a bus / interface controller 330 can be used to facilitate communication between the basic configuration 302 and one or more data storage devices 332 via a storage interface bus 334. The data storage device 332 can be a removable storage device 336, a non-removable storage device 338, or a combination thereof. Examples of removable storage devices and non-removable storage devices include: magnetic disk devices, such as floppy disk drives and hard disk drives (HDDs), optical disk drives, such as compact disk (CD) drives or digital versatile disk (DVD) drives, solid state drives (SSDs) and tape drives, etc. Exemplary computer storage media can include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (e.g., computer readable instructions, data structures, program modules, or other data). The term "computer-readable storage medium" or "computer-readable storage device" does not include propagation signals and communication media.
[0059] System memory 306, removable storage device 336, and non-removable storage device 338 are examples of computer-readable storage media. Computer-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, digital versatile disks (DVD) or other optical storage devices, cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by computing device 300. Any such computer-readable storage media can be part of computing device 300. The term "computer-readable storage media" does not include propagating signals and communication media.
[0060] The computing device 300 can also include an interface bus 340 for facilitating communication from various interface devices (e.g., output devices 342, peripheral interfaces 344, and communication devices 346) to the basic configuration 302 via the bus / interface controller 330. Exemplary output devices 342 include a graphics processing unit 348 and an audio processing unit 350, which can be configured to communicate with various external devices, such as a display or speakers, via one or more A / V ports 352. Exemplary peripheral interfaces 344 include a serial interface controller 354 or a parallel interface controller 356, which can be configured to communicate with external devices such as input devices (e.g., keyboards, mice, pens, voice input devices, touch input devices, etc.) or other peripheral devices (e.g., printers, scanners, etc.) via one or more I / O ports 358. Exemplary communication devices 346 include a network controller 360, which can be arranged to facilitate communication with one or more other computing devices 362 over a network communication link via one or more communication ports 364.
[0061] A network communication link can be an example of a communication medium. Communication media can generally be embodied by computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transmission mechanism, and can include any information transfer medium. A "modulated data signal" can be a signal that sets or changes one or more characteristics in a manner that encodes information in a signal. By way of example and not limitation, communication media can include wired media such as a wired network or a direct wired connection, and wireless media such as acoustics, radio frequency (RF), microwaves, infrared (IR), and other wireless media. The term "computer-readable medium" as used herein can include both storage media and communication media.
[0062] The computing device 300 can be implemented as part of a small form factor portable (or mobile) electronic device, such as a cell phone, a personal data assistant (PDA), a personal media player device, a wireless network viewing device, a personal headset device, a specific application device, or a hybrid device including any of the above functions. The computing device 300 can also be implemented as a personal computer including both laptop computer and non-laptop computer configurations.
[0063] Based on the foregoing, it will be appreciated that the specific embodiments of the present disclosure have been described herein for illustrative purposes, but various modifications may be made without departing from the present disclosure. In addition, many elements of an embodiment may be combined with other embodiments to supplement or replace the elements of other embodiments. Therefore, the present technology is not subject to restrictions other than the appended claims.
Claims
1. A method for automatically summarizing the content of an electronic message, the method comprising: Upon receiving at an email server an incoming email containing an email body having one or more sentences, determining whether the incoming email is a templated message; as well as In response to determining that the incoming email is not a templated message, individually classifying the one or more sentences in the body of the email as statements of decision, judgment, reasoning, or fact; assigning a relevance score to each classified statement of decision, judgment, reasoning, or fact; clustering the classified sentences into one or more clusters according to a hierarchy from the categories of decision, judgment, reasoning, and fact; calculating cluster scores for individual clusters based on the assigned relevance scores of the statements in each cluster; selecting one or more of the clusters based on the calculated cluster scores to automatically generate a summary of the incoming email corresponding to each of the selected clusters; as well as Data representing at least one of the generated summaries is inserted into the incoming email before the incoming email is sent to a destination via a computer network.
2. The method according to claim 1, further comprising: assigning a relevance score to each classified statement of decision, judgment, inference, or fact, wherein the assigned relevance score is based on at least one of an identity of a sender of the incoming email or a recency of the corresponding statement; calculating cluster scores for individual clusters based on the assigned relevance scores of the statements in each cluster; and Wherein, selecting one or more of the clusters includes selecting one or more of the clusters based on the calculated cluster scores.
3. The method according to claim 1, wherein: Clustering the classified sentences includes: determining whether the classified statement includes at least one decision; and In response to determining that the classified statement includes at least one decision, a cluster is generated around the at least one decision.
4. The method according to claim 1, wherein: Clustering the classified sentences includes: determining whether the classified statement includes at least one decision; and In response to determining that the classified statement does not include at least one decision, determining whether the classified statement includes at least one predicate; and In response to determining that the classified statement includes at least one judgment, a cluster is generated around the at least one judgment.
5. The method according to claim 1, wherein: Clustering the classified sentences includes: determining whether the classified statement includes at least one decision; and In response to determining that the classified statement does not include at least one decision, determining whether the classified statement includes at least one judgment; In response to determining that the classified statement does not include at least one judgment, determining whether the classified statement includes at least one inference; and In response to determining that the classified statement includes at least one inference, clusters are generated around the at least one inference.
6. The method according to claim 1, further comprising: In response to determining that the incoming email is a templated message, identifying a summary template corresponding to the incoming email; as well as A summary is generated based on the identified summary template.
7. The method according to claim 1, wherein: Inserting data representing at least one of the generated summaries comprises: converting at least one of the generated summaries from text to a voice message; and The converted voice message is inserted into the incoming email for playback at the client device.
8. A computing device for processing electronic messages, the computing device comprising: processor; A memory comprising instructions executable by the processor to cause the computing device to perform the method according to one of claims 1-7.
Citation Information
Patent Citations
Method and system for summarizing emails and extracting tasks
US20170161372A1