Multiple compression structured content based on source content
Through multi-sentence compression technology based on graphs, candidate sentences are generated from the corpus and the final content is optimized, which solves the problems of low productivity and redundancy of content authors when generating new content, and achieves efficient and coherent content generation.
Patent Information
- Application Number
- CN201811203672.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-12-26
- Filing Date
- 2018-10-16
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2038-10-16
AI Technical Summary
Content authors need to manually manage pre-existing information when generating new content, resulting in low productivity, error-prone and difficult to meet the needs of different channels and audiences.
Through graph-based multi-sentence compression technology, the source content is identified and retrieved from the corpus, and candidate sentences are generated after parsing, mapping and weighting, and the coordinatedness and redundancy of the final content are optimized using a mixed integer program.
Simplifies the content generation process, improves productivity, reduces redundancy, and ensures content coherence and relevance to user input.
Smart Images

Figure CN109960721B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to data processing, and more particularly, to methods, systems, and computer-readable media for generating content using existing content. Background Art
[0002] As the number of channels for consuming content increases, content writers (e.g., authors) who are involved in writing textual content (e.g., articles) for various purposes need to ensure that the content they generate meets the requirements of their chosen content distribution channels and the needs of their desired target audiences. For example, while some channels (such as social media platforms) may require shorter content representations, other channels (such as newsletters, information brochures, newspapers, and websites) may allow for more elaborate content representations.
[0003] In order to meet the needs of specific selected channels and target audiences, content authors often search for pre-existing content that can be reused to generate new content or elaborate content. Usually, the additional information that the author is searching for already exists in various forms of expression, for example, on the Internet or in an enterprise setting (for example, a company's document database). In the absence of appropriate help, content authors manually manage (curate) such content from a corpus, thereby reducing its productivity. For example, searching for relevant information, analyzing relevant information to remove duplicate information and to ensure that various topics are covered and then preparing well-written content may be time intensive. In some cases, the tediousness of manual content management from pre-existing content causes the author to generate content from scratch, rather than spending time searching for pre-existing content that is difficult to locate to replace purposes. However, this manual content management may cause various errors and inconsistencies. Summary of the Invention
[0004] Embodiments of the present invention relate to methods, systems, and computer-readable media for generating content using existing content. In this regard, source content related to an input segment can be accessed and used as a basis for generating new content. When relevant source content is identified, the source content is typically compressed to generate new candidate content. The candidate content can then be evaluated to sort the content in a cohesive manner to form the final content. Advantageously, corpus-based automatic content generation optimizes relevance based on the input segment (e.g., keywords, phrases, or sentences), covers different information within the generated final content, minimizes content redundancy, and improves the coherence of the final content.
[0005] In order to generate new content, the embodiments described herein support extracting the user's intent from an input fragment. Thereafter, pre-existing source content (e.g., a fragment of text information) in the corpus can be identified and retrieved for use in generating candidate content. In particular, the input fragment is used to formulate a query that identifies pre-existing source content to be retrieved from the corpus. In addition, the retrieved source content from the corpus is compressed to form new candidate sentences that are included in the final content output. Specifically, a graph-based formulation and weighting system is utilized to support multi-sentence compression to generate new candidate content. Candidate content generation can be iteratively performed until the retrieved source content related to the fragment is exhausted. The generated new candidate content is sorted and ordered to form a coherent final content. It will be understood that the final content can meet the content length expected by the user.
[0006] This summary is provided to introduce some concepts in a simplified form, which will be further described in the detailed description below. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to help determine the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The present invention is described in detail below with reference to the accompanying drawings, in which:
[0008] Figure 1 is a schematic diagram of a system for facilitating content generation according to an embodiment of the present invention;
[0009] Figure 2 is a depiction of a content generation engine according to an embodiment of the present invention;
[0010] Figure 3 is a depiction of a flow chart illustrating a method of retrieving source content from a corpus according to an embodiment of the present invention;
[0011] Figure 4 is a depiction of a flowchart illustrating a method of compressing source content into candidate sentences according to an embodiment of the present invention;
[0012] Figure 5 is a depiction of a flow chart illustrating a method of ordering candidate sentences into a coherent final content according to an embodiment of the present invention;
[0013] Figure 6 is a depiction of a flow chart of an example content generation method according to an embodiment of the present invention; and
[0014] Figure 7 is a depiction of a block diagram of an exemplary computing environment suitable for implementing embodiments of the present invention. DETAILED DESCRIPTION
[0015] The subject matter of the invention is specifically described herein to meet statutory requirements. However, this description itself is not intended to limit the scope of this patent. On the contrary, the inventors have anticipated that the claimed subject matter may also be embodied in other ways, to include different steps or combinations of steps similar to the steps described in this document, as well as other existing or future technologies. In addition, although the terms "step" and / or "block" may be used herein to refer to different elements of the method employed, these terms should not be interpreted as implying any particular order among or between the steps disclosed herein unless the order of the steps is explicitly described.
[0016] Text content is usually prepared for different purposes. By way of example, text content can be an article that describes various rules and regulations in detail that the author wishes to construct. Text content can also include an article that describes the new and improved specifications of the recently released technology that the author wishes to generate in detail, for publishing on the company's website or in the instruction manual, user guide or quick start guide of the product. Text content can further comprise a longer article that describes risk factors, symptoms, treatment and the prevention of specific diseases in detail that will be published in professional journals (such as medical journals). When preparing such text content, the author can consider the various lengths of the desired content, the channels for content distribution and / or the expected target audience.
[0017] Often, the information an author seeks to cover in a new article already exists. For example, the desired information may exist in some form within a corporate corpus or on the internet, including past articles, fact sheets, technical specifications, and other documents that can be repurposed to suit the author's current needs. However, this information can be distributed across many documents and systems. Therefore, locating and identifying the desired information is often difficult and time-consuming.
[0018] Furthermore, when searching for and identifying relevant information, content authors often manually compose new content, which is also time-consuming and error-prone. In this regard, upon identifying information of interest, authors analyze data and, based on the information obtained, identify how to compose new content. For example, authors may identify which information to use and how to organize it. Unfortunately, in addition to being time-consuming, this manual content generation can result in duplicate information, grammatical errors, incoherence, and the like. Additionally, in the absence of a source retrieval mechanism suitable for identifying the author's desired information, authors often manually create content from scratch, further reducing productivity.
[0019] To avoid manually searching for additional source information to create new content, one conventional approach involves identifying key concepts in the initial document and linking these concepts to corresponding Wikipedia pages. The author can then navigate to the linked Wikipedia pages and examine them to determine whether they contain useful information that the author wishes to manually repurpose to generate the content. However, while this solution identifies relevant Wikipedia pages that may contain useful information, it merely reduces the amount of manual searching required. Furthermore, after manually reviewing the Wikipedia pages to identify any useful information, the author still needs to manually generate the content.
[0020] There are also some efforts towards providing content in the context of knowledge management. One such solution aims to enhance question answering by automatically extending a given text corpus with relevant content from a large external source (such as the Internet) using an extension algorithm. In particular, the solution uses Internet resources to build answers based on "paragraph blocks" to extend the "seed document". However, this solution is not intended for human consumption, but ideally only applies to the question answering space. Additionally, the solution does not take into account lexical and semantic redundancy in the answers, which is unnecessary for content authors. Recent text generation relies on training a neural network that can learn the generation process. However, this neural network training relies on a wide range of training corpora that include both content fragments and expected content generation, which requires meaningful annotations.
[0021] Accordingly, embodiments described herein relate to automatically generating content using existing content. In this regard, embodiments described herein automatically generate content (also referred to herein as final content) using available source content without the need for training data. In operation, based on a fragment input by a user, relevant source content can be identified. In embodiments, content is generated from a pre-existing collection of source content covering various aspects of the target information to diversify or expand the coverage of the technical topics covered in the content generation. Using such relevant source content, candidate content (such as candidate sentences) can be generated. As discussed herein, candidate content generation identifies content relevant to the user input and compresses the content to minimize redundancy in the content. To select appropriate content for construction, graphical representations of various sentences can be used. The most "beneficial" portions of the graphical representations are identified, and the corresponding sentences can be compressed to generate candidate content with minimal content redundancy. Content compression reduces syntactic and lexical redundancy caused by multiple representations of the same information in different parts of the content corpus. Furthermore, content compression supports generating new content by observing sentence structure in the corpus, rather than simply selecting sentences from the corpus. When generating candidate content from the most "interesting" portion of the graphical representation, the graph is adjusted to take into account the information in the generated candidate content, thereby improving the information coverage in subsequently generated content. It will be appreciated that candidate content can be iteratively generated until the overall information related to the user input in the graphical representation is exhausted (at least to some extent).
[0022] When generating candidate content, final content can be selected. In an embodiment, a mixed integer program (MIP) is used to select content that maximizes relevance to user input and is coherent. In this regard, the candidate content can be sorted and / or combined in a coherent manner. In addition, when optimizing the final content, the final content can be constructed according to a specific budget or desired content length. In this way, final content with compressed sentences and sequences can be output to produce coherent content of the desired length.
[0023] Now go to Figure 1 , a schematic diagram is provided to illustrate an exemplary system 100 to which some embodiments of the present invention may be applied. In addition to other components not shown, the environment may include a content generation engine 102, a user device 104, and a data store 106. It should be understood that Figure 1 The system 100 shown in FIG. 1 is an example of a suitable computing system. Figure 1 Any component shown in the can be accessed via any type of computing device (such as a Figure 7The components can communicate with each other via one or more networks 108, which may include, but are not limited to, one or more local area networks (LANs) and / or wide area networks (WANs). Such networking environments are common in offices, enterprise-wide computer networks, intranets, and the Internet.
[0024] Should be understood that this and other arrangements described herein are only set forth as examples.Other arrangements and elements (such as, machines, interfaces, functions, sequences, functional groups etc.) can be used as those additional or alternatives shown, and some elements can be omitted completely.In addition, many elements described herein are functional entities, which can be implemented as discrete or distributed components or implemented together with other components, and can be implemented with any suitable combination and position.The various functions performed by one or more entities described herein can be performed by hardware, firmware and / or software.For example, the various functions can be performed by the processor of the instruction stored in the memory.
[0025] Typically, system 100 supports content generation using existing content. As used herein, content generally refers to electronic text content, such as documents, web pages, articles, etc. Content generated using pre-existing source content is generally referred to as final content in this article. Source content generally describes pre-existing text content (e.g., in a corpus). In this regard, source content can include, for example, various documents located on the Internet or in a data store.
[0026] At a high level, the final content is generated using pre-existing source content from a corpus, which, once retrieved, is parsed, mapped, weighted, and compressed (typically in the form of sentences) to form candidate content. As used herein, candidate content generally refers to newly generated content that can be used in the final content. Candidate content is generally described herein as candidate sentences constructed using graph-based compression of similar sentences from the source content. However, candidate content can be a variety of other content fragments and is not intended to be limited to sentences. Furthermore, candidate content can be generated in a variety of ways. As described herein, candidate content can cover different aspects of the fragments input by the author. The candidate content can then be weighted and sorted to generate a coherent final content, which the author can then utilize and / or edit to obtain an acceptable final version of the content.
[0027] As an example, let's assume that an author (such as an employee of a corporate company) wishes to learn more about specific company regulations regarding activities or tasks that the employee must complete and wishes to generate a single piece of content containing the specific regulatory information. Furthermore, let's assume that the information that the employee is attempting to repurpose to generate the final regulatory content exists as source content in various documents within the corporate company's corpus. In this scenario, the employee enters a snippet (e.g., a sentence, phrase, or keyword) related to the regulation that the employee wishes to learn about. After obtaining the snippet, the author needs to extract the query. This query can then be used to identify and retrieve source content from the corporate company's corpus that is related to the entered snippet (e.g., a sentence, phrase, or keyword related to the specific regulation that the employee wishes to learn about). The retrieved source content containing the regulatory information is then parsed into sentences, mapped to a selected graph, and weighted. Thereafter, the sentences are further parsed into word tokens, mapped to a compressed graph, and weighted. The mapped word tokens are then compressed into candidate content (such as candidate sentences) to be included in the final content output. The generated candidate content can each be different and contain different information from the source content related to the desired regulation. Such candidate content is then weighted and ranked to output final coherent content relevant to the regulations.
[0028] return Figure 1 In operation, the user device 104 can access the content generation engine 102 via a network 108 (e.g., a LAN or the Internet). For example, the user device 104 can provide and / or receive data from the content generation engine 102 via the network 108. The network 108 may include multiple networks or multiple networks of networks, but it is shown in a simplified form so as not to obscure various aspects of the present disclosure. By way of example, the network 108 may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the Internet, and / or one or more private networks. Network environments are common in offices, enterprise-wide computer networks, intranets, and the Internet. Accordingly, the network 108 is not described in detail.
[0029] A user device (such as user device 104) can be any computing device that can support a user to provide a snippet. A snippet as used herein generally refers to a text item entered by an author, which is an indicator of the author's intent. In this regard, a snippet can take many forms, for example, a word, a phrase, a sentence, a paragraph, a set of keywords, etc. The snippet can be analyzed to formulate a query to identify and retrieve source content from a corpus. For example, a user can provide a snippet to the content generation engine 102 via a browser or application installed on the user device 104. In addition, any type of user interface can be used to input such a snippet. In some cases, a user can input a snippet, for example, by typing the snippet or copying / pasting the snippet.
[0030] In response to providing the fragment, the user device 104 can obtain and present the final content or a portion thereof. In this regard, the final content generated in response to the user-provided fragment can be provided to the user device for display to the user (e.g., via a browser or application installed on the user device 104).
[0031] In some cases, the user device 104 accesses the content generation engine 102 via a web browser, terminal, or standalone PC application operable on the user device. The user device 104 may be operated by an administrator, who may be an individual who manages content associated with a document, website, application, or the like. For example, a user may be any individual associated with an entity that publishes content (e.g., via the Internet), such as an author or publisher. Although only one user device 104 may be in operation, the administrator may be an individual who manages content associated with a document, website, application, or the like. Figure 1 , but multiple user devices associated with any number of users can be utilized to perform the embodiments described herein. User device 104 can take various forms, such as a personal computer (PC), a laptop computer, a mobile phone, a tablet computer, a wearable computer, a personal digital assistant (PDA), an MP3 player, a global positioning system (GPS) device, a video player, a digital video recorder (DVR), a cable box, a set-top box, a handheld communication device, a smart phone, a smart watch, a workstation, any combination of these depicted devices, or any other suitable device. In addition, user device 104 may include one or more processors, and one or more computer-readable media. The computer-readable medium may include computer-readable instructions that can be executed by one or more processors.
[0032] In addition to other data, the data store 106 (e.g., a corpus) includes source content data that can contain information desired by the author to support candidate content generation and final content generation. As described in more detail below, the data store 106 can include source content data that includes electronic text content, such as documents, web pages, articles, etc., and / or metadata associated therewith. Such source content data can be stored in the data store 106 and can be accessed by any component of the system 100. The data store can also be updated at any time, including an increase or decrease in the amount of source content data, or an increase or decrease in the amount of content in the data store that is not related to the fragments entered by the author. In addition, the information covered in the various documents in the corpus can be changed or updated at any time.
[0033] The content generation engine 102 is typically configured to generate candidate content (e.g., sentences) from existing source content and thereafter utilize such candidate content to construct a coherent final content. In particular, the content generation engine 102 can identify and retrieve source content, parse and compress the source content into candidate sentences, and combine the candidate sentences into a coherent final content. In an embodiment, and at a high level, the content generation engine 102 formulates a query to identify and retrieve source content from a corpus. In particular, the content generation engine 102 can extract author intent from a fragment of author input and use the author intent to formulate a query. The query then identifies and retrieves source content from the corpus. The retrieved source content is then parsed and compressed to generate candidate sentences. The candidate sentences are then sorted to generate a coherent final content output.
[0034] An exemplary content generation engine Figure 2 is provided. Figure 2 As shown, the content generation engine 200 includes a source content retrieval manager 202, a candidate content generation manager 204, and a final content generation manager 206. The source content retrieval manager 202 generally facilitates retrieving pre-existing source content from a corpus. The candidate content generation manager 204 utilizes the pre-existing source content (e.g., via graph-based sentence compression) to generate candidate content. The final content generation manager 206 generally utilizes the candidate content to generate final content, e.g., for output to a user device. Advantageously, the generated final content conveys relevant information in a coherent and non-redundant manner.
[0035] Although shown as the separate component of content generation engine 200, any number of components can be used to perform the functions described herein. In addition, although shown as a part of content generation engine, these components can be distributed via any number of devices. For example, the source content retrieval manager can be provided via a device, server or server cluster, and the candidate content generation manager can be provided via another device, server or server cluster. The components identified herein are listed only as examples to simplify or illustrate the discussion of function. As additional or alternatives to those shown, other arrangements and elements (for example, machines, interfaces, functions, sequences and functional groupings, etc.) can be used, and some elements can be omitted completely. In addition, many elements described herein are functional entities, which can be implemented as discrete or distributed components or implemented together with other components, and can be implemented in any suitable combination and position. The various functions performed by one or more components described herein can be performed by hardware, firmware and / or software. For example, the various functions can be performed by a processor that performs an instruction stored in a memory.
[0036] As described, the source content retrieval manager 202 is generally configured to collect source content retrieved from a corpus for use in generating candidate content. The source content retrieval manager 202 may include a snippet collector 212, a query formulator 214, and a source content acquirer 216. Although shown as separate components of the source content retrieval manager 202, any number of components may be used to perform the functions described herein.
[0037] The snippet collector 212 is configured to collect or obtain snippets, for example, via user (e.g., author) input. As described, snippets can be keywords, phrases, and sentences, but are not limited to these text arrangements. Snippets can be collected or obtained in any manner. In some cases, snippets are provided by users of the content generation engine (such as enterprise content authors). In this regard, the enterprise content author or authors can enter or input snippets, for example, via a graphical user interface accessible through an application on the user's device. As an example, a user can enter snippets via Figure 1 The snippet may be entered by a user device 104 connected to the network 108. For example, an enterprise content author may provide keywords, sentences, or phrases.
[0038] The query formulator 214 is configured to identify or extract the user's intent (i.e., the author's need) from the input segment. Based on the user's intent, the query formulator 214 can formulate a query to identify and obtain source content related to the segment from the corpus. In order to identify the user's intent, a set of keywords can be extracted from the input segment or identified within the input segment. In an embodiment, a set of the top "K" keywords in the segment can be extracted using the inverse document frequency (IDF) of the words in the corpus. This approach can capture the most important keywords in the segment with respect to the corpus. After obtaining the set of keywords, a query can be generated by concatenating these keywords or a portion thereof.
[0039] It will be appreciated that the number of keywords selected can affect the relevance and quantity of source content that is available and retrieved. For example, a low value of K keywords can result in a query that is not sufficiently representative of the source content in the corpus and retrieves a large amount of source content that may not be very relevant to the input segment. On the other hand, a higher value of K keywords can result in a more specific query that may not retrieve as much source content from the corpus. Furthermore, in an embodiment, the frequency of terms in a segment sentence, phrase, etc. is not considered in the weighting process for the selection of K, because most terms appear only once in the segment, and multiple occurrences may not be intentional, but rather accidental or erroneous, and therefore misleading.
[0040] The source content acquirer 216 is typically configured to obtain source content or content fragments relevant to the generated query. That is, the source content acquirer 216 utilizes the query to obtain relevant source content (e.g., from a corpus). Such source content can be obtained from any number of sources (such as various distributed sources). As described herein, the choice of K affects the relevant amount of source content identified in the corpus by the source content acquirer 216 and retrieved from the corpus by the source content. For example, returning an example of an employee author who expects to generate content about a specific company's rules and regulations, a higher K value may result in a more specific query that may retrieve less and more tightly tailored source content from the corpus that only relates to the rules and regulations specified by the author. However, a lower K value may result in a query that is insufficiently representative of the source content in the corpus and retrieves source content that covers additional information in addition to the rules and regulations desired by the author, such as other company rules and regulations that the author does not intend to understand or generate content for.
[0041] As described, the candidate content generation manager 204 is generally configured to generate candidate content by compressing the identified and retrieved source content from the corpus. The candidate content generation manager 204 may include a source content parser 220, a sentence mapper 222, a sentence parser 224, a word token mapper 226, a reward allocator 228, a content compressor 230, and a content collector 232. Although shown as separate components of the candidate content generation manager 204, any number of components may be used to perform the functions described herein.
[0042] The source content parser 220 is configured to parse the source content or content fragments obtained by the source content retriever 218. In an embodiment, the obtained source content can be parsed into sentences. In this regard, the source content parser 220 breaks down the retrieved source content into individual sentences so as to construct sentences in a form suitable for mapping sentences. Although generally discussed as parsing content into sentences, it is understood that other content fragments can be used to parse the source content.
[0043] The sentence mapper 222 typically maps sentences to a first graph, which is referred to herein as a selection graph. In particular, sentences can be represented graphically using a knot notation, where each sentence is represented by a node on the selection graph.
[0044] The reward allocator 228 assigns an initial reward (i.e., a node weight) to each node on the selection graph. In an embodiment, the initial reward or weight assigned to a node may be based on the similarity of the node to the query. Additionally or alternatively, the initial reward or weight assigned to a node may be based on the amount of information present in the sentence associated with the node. For example, a higher reward may indicate that the sentence contains multiple concepts rather than a single concept or topic. Edge weights may also be provided for edges between pairs of nodes based on the information overlap between corresponding nodes.
[0045] Sentence parser 224 is usually configured to parse sentences into word tags (token). In this regard, sentence parser 224 decomposes sentence tags into separate word tags so that words are in a form suitable for word tag mapper 226 to map word tags to the second graph (commonly referred to as compression graph). In particular, sentence parser iteratively selects a subgraph (i.e., part) of the selection graph for parsing. The part selected for parsing and compression in the selection graph can be a part or sentence set that is identified as the most valuable part in the selection graph. That is, a part of the nodes that are closely related to each other in the selection graph can be identified for sentence compression. In one embodiment, the selected subgraph can include a node and its 1-hop or 2-hop neighbor. In this case, the sentences corresponding to these nodes can be parsed into word tags. As an example only, the node with the maximum gain, and the sentence corresponding to the node and the sentence in the 1-hop or 2-hop neighbor of the node are selected and put into set "S", which can then be used for the first iteration of multi-sentence compression. In other words, the node on the selection graph, its corresponding sentence and its corresponding 1-hop or 2-hop neighbor are selected and parsed into word tags. For example, iterative sentence node selection can be performed by first selecting the node with the largest gain on the selection graph, which can be expressed as:
[0046]
[0047] The selection graph is given by G(v,e), and each node on the selection graph representing a sentence from the source content is represented by v in the graph. i ∈V, the initial reward is given by r o i Given, for v i , N i refers to node v i The edge weight between each pair of sentence nodes is w ij Given by G vi Given, and where l is the maximum gain G vi The initial sentence node v * i When parsing the sentences into words, the word token mapper 226 maps the tokenized words to a second graph (generally referred to herein as a compressed graph). In an embodiment, the words are represented by knot notations, where each word is represented by a node on the compressed graph, and each sentence represents a directed path in the compressed graph. Special nodes are used as sentence start nodes and sentence end nodes. A single node mapped to the compressed graph can represent each occurrence of a word within the same part of speech (POS) tag.
[0048] The reward allocator 228 can assign edge weights between each pair of node nodes. Such weights can represent the relationship of the words to each other. For example, the edge weights can represent the number of times the ordered combination of these node words appears in all sentences in the set S. The shortest path (normalized by path length) can be identified, and the top K sentences generated are used for further processing as described below.
[0049] Content compressor 230 is usually configured to generate candidate content from compression graph. Accordingly, content compressor 230 can identify the path of the shortest path (by path length standardization) on the compression graph, and compress these paths into candidate content (such as candidate sentences). The candidate content generated by content compressor 230 is normally the content that covers the information contained in the sentence set from which candidate content is generated. In one embodiment, the shortest path is identified, and the first K sentences generated are identified for compression. For example, in one embodiment, the minimum number of words of each generated sentence can be limited to 10 words, wherein each iteration selects a sentence. This traversal of the path causes the most appropriate sentence to be generated based on the co-occurrence in the corpus. Usually, the shortest path is utilized to produce content (such as a sentence) in compressed form, which captures the information from multiple sentences.
[0050] Based on performing content compression, the reward allocator 228 can reallocate rewards or weights. In this regard, rewards (i.e., weights) can be assigned to compressed candidate sentences based on their similarity to the query. In order to take into account the information captured by each sentence compressed into a candidate sentence for subsequent subgraph selection in the selection graph, the reward allocator 228 can update the rewards of the following vertices in the selection graph whose information overlaps significantly (or exceeds a threshold) with the candidate content that has already been generated. This reduces the rewards for sentences covered by the current set of generated candidate sentences, thereby reducing the chance that the same information already included in the generated candidate content is included in the subsequently generated candidates. In particular, this ensures that the information covered by the subsequent candidate sentence generation is different from the information that has already been generated, while ensuring that the information coverage of the subsequently generated candidate sentences is still relevant to the input query.
[0051] The content collector 232 is typically configured to collect a set of candidate content. In this regard, the content collector 232 can collect the generated candidate content after each compression iteration. Advantageously, the generated candidate content (such as sentences) generally covers the information space related to the input segment. It will be appreciated that any number of candidate content can be generated and collected. For example, candidate content can be generated and collected until a threshold number of candidate sentences are generated. As another example, candidate content generation can continue until no nodes with significant rewards remain, which results in the same subgraph being selected for compression in consecutive iterations.
[0052] Since the compressed candidate content may not be initially grammatically correct and / or ordered, the final content generation manager 206 may generate coherent final content. The final content generation manager 206 may include a sentence orderer 234 and a final content generator 236.
[0053] The sentence sorter 234 is generally configured to sort (i.e., combine, organize) the candidate content into a coherent final content. In an embodiment, the sentence sorter 234 can sort the candidate content by selecting an appropriate set of compressed content to be sorted (e.g., sentences) and their order and using an integer linear program formula (mixed integer program (MIP)). The integer linear program objective can be given by:
[0054]
[0055] So that:
[0056] y i,j =0if coh ij <σ or i=j (4)
[0057]
[0058]
[0059]
[0060]
[0061] Among them, the binary variable x i indicates the selection / non-selection of the compressed sentence i, and the binary variable y i,j Indicates from sentence x i and x j The conversion of y i,j Traversing the sentence path will produce the final generated sentence. For each compressed sentence, w iIndicates a combination of the relevance of the sentence to the fragment and its overall linguistic quality. This ensures that the selected compressed sentences are noise-free and relevant to the fragment. The second term in Equation 2 maximizes the coherence of the selected sentence flow. The constraint in Equation 3 takes into account the author's requirement to limit the generated content to the length of the target budget B. Equation 4 prohibits the flow of content between smaller coherent sentences and avoids loops in arcs. An arc between two nodes exists if both sentences are selected for the final content and they are continuous. Equation 6 limits the number of starting and ending sentences to 1. Equations 7 and 8 limit the number of input and output arcs from the selected node to 1 each, forcing the formation of paths via arcs that can indicate the flow in the selected content.
[0062] The final content generator 236 is typically configured to output a sorted set of candidate content, for example, to a user device. By way of example, the final content generator 236 is configured to output the sorted content to the author of the input segment in some cases. The embodiments described herein enable the generation of final content that reduces the redundancy that would be included in the absence of the content, increases the overall coherence of the content, and increases the overall information coverage of the content. In particular, content compression (multi-sentence compression) is used to reduce redundant information in the candidate content by compressing information from multiple sentences into a single sentence. Graph-based candidate selection for compression enables coverage of various aspects of the final content output. Finally, the coherence of the final content output is enhanced using integer linear program formulas.
[0063] An example content generation algorithm based on an input fragment can be expressed as follows:
[0064]
[0065] Return now Figure 3 , a flow chart showing an exemplary method 300 for retrieving source content from a corpus according to an embodiment of the present invention is illustrated. In an embodiment, the method 300 is executed by a content generation engine such as Figure 2 The content generation engine 200 of is executed. Initially, as shown at box 302, author-entered snippets are collected. The author-entered snippets may include keywords, sentences, phrases, etc. At box 304, author needs associated with the author-entered snippets are extracted. In some cases, the author needs indicate the author's intention. Thereafter, at box 306, a query is formulated based on the extracted author needs to identify pre-existing source content in the corpus that is related to the author-entered snippets. At box 308, the pre-existing source content identified in the corpus is retrieved for further processing. As described, source content refers to electronic text content, such as documents, web pages, articles, etc.
[0066] Now refer to Figure 4, a flowchart illustrating an exemplary method 400 of generating candidate sentences using multi-sentence compression according to an embodiment of the present invention is illustrated. In an embodiment, the method 400 is performed by a content generation engine (such as Figure 2 ) is executed by the content generation engine 200 of . Initially, and as shown in box 402, source content from the corpus is retrieved. At box 404, the source content is parsed into sentences. Referring to box 406, the sentences are then mapped to the selection graph. As described, the mapped sentence tokens are mapped in knot notations, where each node represents a single sentence. At box 408, the mapped sentences are assigned initial rewards (i.e., weights) and edge weights. In an embodiment, the nodes mapped to the selection graph are weighted based on their similarity to the query, and edge weights are assigned to the edges between them based on the information overlap of each node pair. At box 410, the sentences from the selection graph are parsed into word tokens. At box 412, the word tokens are mapped to a compressed graph. In an embodiment, the mapped tokenized words are represented using knot notations, where each word is represented by a node on the compressed graph, and each sentence represents a directed path in the compressed graph. Special nodes are used as sentence start nodes and sentence end nodes. A single node mapped to the compression graph can represent all occurrences of a word within the same POS tag. Referring to box 414, edge weights are assigned between each word node pair. The edge weights can represent the number of times that ordered combination of those node words occurs within all sentences in the set S. The shortest path (normalized by path length) is identified, and the top K sentences generated are used for further processing. At box 416, the mapped word tags are compressed into candidate sentences to be used for the final content. This candidate sentence generation can be iteratively repeated until all relevant information mapped to the selection graph is exhausted.
[0067] Return now Figure 5 , a flowchart showing an exemplary method 500 of sorting candidate sentences into final content according to an embodiment of the present invention is illustrated. In an embodiment, the method 500 is performed by a content generation engine (such as Figure 2 The content generation engine 200 of is executed. Initially, and as shown at block 502, the compressed generated candidate sentences are collected. At block 504, the candidate sentences are assigned weights. As described, the reward (i.e., weight) assigned to each compressed candidate sentence can be based on its similarity to the query. Referring to block 506, the weighted candidate sentences are sorted to form the final content. The final content typically includes increased information coverage, reduced redundancy, and a coherent overall flow.
[0068] Return now Figure 6 , a flowchart showing an exemplary method 600 of sorting candidate sentences into final content is illustrated. In an embodiment, the method 600 is performed by a content generation engine (such as Figure 2) is executed by the content generation engine 200. Initially, and as shown in box 602, a fragment entered by the author is received. The fragment entered by the author may include keywords, sentences, phrases, etc. At box 604, the author needs associated with the fragment entered by the author are extracted to formulate a query. The query is formulated based on the extracted author needs to identify pre-existing source content in the corpus that is related to the fragment entered by the author. At box 606, the pre-existing source content identified in the corpus is retrieved for further processing. As described, source content refers to electronic text content, such as documents, web pages, articles, etc. Referring to box 608, and as described herein, the retrieved source content is parsed, mapped and weighted to a selection graph and a compression graph for further processing and candidate sentence generation. In particular, the source content is parsed into sentence tokens, which are then mapped to the selection graph. The mapped sentence tokens can be mapped with node notations, where each node represents a single sentence. The mapped sentence tokens are then assigned initial rewards (i.e., weights) and edge weights, where they are weighted based on the similarity of the nodes mapped to the selection graph to the query, and edge weights are assigned to the edges between each pair of nodes based on the information overlap of each pair of nodes. In addition, the sentences from the selection graph are parsed into word tokens, which are then mapped to the compression graph. The mapped tokenized words can be represented using knot notations, where each word is represented by a node on the compression graph and each sentence is represented by a directed path in the compression graph. A single node mapped to the compression graph can represent all occurrences of words within the same POS tag. Edge weights are assigned between each pair of word nodes, and the edge weights can represent the number of times an ordered combination of these node words appears in all sentences in the set S. The shortest path (normalized by the path length) is identified, and the top K sentences generated are used for further processing. At box 610, candidate sentences are generated and weighted, where the multi-sentence compression that enables candidate sentence generation is iteratively repeated until all relevant source content mapped to the selection graph is exhausted. At block 612 , the candidate sentences are ranked into final content to be output to the author.
[0069] Having described embodiments of the present invention, an exemplary operating environment in which embodiments of the present invention may be implemented will be described below to provide a general context for various aspects of the present invention. Figure 7 , an exemplary operating environment for implementing embodiments of the present invention is shown and generally designated as computing device 700. Computing device 700 is only one example of a suitable computing environment and is not intended to suggest any limitation as to the scope of use or functionality of the invention. Neither should computing device 700 be interpreted as having any dependency or requirement relating to any one or combination of components illustrated.
[0070] The present invention may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions (such as program modules) executed by a computer or other machine (such as a personal data assistant or other handheld device). Generally, program modules include routines, programs, objects, components, data structures, etc., which refer to code that performs specific tasks or implements specific abstract data types. The present invention can be implemented in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present invention can also be implemented in distributed computing environments, where tasks are performed by remote processing devices linked through a communications network.
[0071] refer to Figure 7 , computing device 700 includes a bus 710 that directly or indirectly couples the following devices: memory 712, one or more processors 714, one or more presentation components 716, input / output (I / O) ports 718, I / O components 720, and an illustrative power supply 722. Bus 710 represents what may be one or more buses (such as an address bus, a data bus, or a combination thereof). Although for clarity, Figure 7 The various boxes of FIG are shown with lines, but in reality, the depiction of the various components is not so clear, and metaphorically, the lines would more accurately be gray and fuzzy. For example, a presentation component such as a display device can be considered an I / O component. Furthermore, a processor has memory. The inventors recognize that this is the nature of the art and reiterate that Figure 7 The figures are merely illustrations of exemplary computing devices that may be used in conjunction with one or more embodiments of the present invention. No distinction is made between categories such as "workstation," "server," "laptop," "handheld device," etc., as all of these are in the Figure 7 is contemplated within the scope of and references to “computing devices.”
[0072] Computing device 700 typically includes a variety of computer-readable media. Computer-readable media can be any available media accessible by computing device 700 and includes both volatile and non-volatile media, as well as removable and non-removable media. By way of example, and not limitation, computer-readable media can include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVDs) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and is accessible by computing device 700. Computer storage media do not themselves include signals. Communication media typically embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and include any information delivery media. The term "modulated data signal" refers to a signal that has one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared and other wireless media. Combinations of any of the above should also be included within the scope of computer-readable media.
[0073] Memory 712 includes computer storage media in the form of volatile and / or nonvolatile memory. Memory can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, and the like. Computing device 700 includes one or more processors that access data from various entities, such as memory 712 or I / O components 720. Presentation component 716 presents data indications to a user or other device. Exemplary presentation components include a display device, a speaker, a printing component, a vibration component, and the like.
[0074] I / O ports 718 allow computing device 700 to be logically coupled to other devices including I / O components 720, some of which may be built-in. Illustrative components include microphones, joysticks, game controllers, satellite dishes, scanners, printers, wireless devices, and the like. I / O components 720 can provide a natural user interface (NUI) that processes in-air gestures, voice, or other physiological input generated by the user. In some cases, the input can be transmitted to an appropriate network element for further processing. The NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometrics, gesture recognition on and near the screen, in-air gestures, head and eye tracking, and touch recognition (as described in more detail below) associated with the display of computing device 700. Computing device 700 can be equipped with a depth camera for gesture detection and recognition, such as a stereo camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations thereof. Additionally, computing device 700 can be equipped with an accelerometer or gyroscope capable of detecting motion. The output of the accelerometer or gyroscope can be provided to the display of computing device 700 to present immersive augmented reality or virtual reality.
[0075] The present invention has been described with respect to specific embodiments, which are intended in all respects to be illustrative rather than restrictive. Alternative embodiments will become apparent to those skilled in the art to which the present invention pertains without departing from the scope of the present invention.
[0076] From the foregoing it will be seen that the present invention is well adapted to attain all of the objects and aims set forth above, together with other advantages which are obvious and inherent to the systems and methods. It will be understood that certain features and subcombinations are useful and may be employed without reference to other features and subcombinations. This is contemplated by and is within the scope of the claims.
Claims
1. A computer storage medium storing computer-usable instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations comprising: Obtaining source content associated with an input segment received via a user interface; identifying, using a graphical representation of a plurality of sentences from the source content, a set of sentences in related source content having overlapping information, wherein the overlapping information is determined based on a set of keywords associated with the input segment received via the user interface; generating candidate sentences by compressing content in the sentence set having overlapping information; generating a final content including a candidate sentence set, wherein the candidate sentence set includes the candidate sentence; as well as The final content is provided as content automatically created in response to the input fragment. 2 . The computer storage medium of claim 1 , further comprising parsing the source content associated with the input segment into a plurality of sentences.
3. The computer storage medium of claim 2, wherein each of the plurality of sentences is mapped to a first graph and assigned a weight, wherein the weight is related to a relevance between the corresponding sentence and the input segment received via the user interface.
4. The computer storage medium of claim 3, wherein at least a portion of the weighted sentence is parsed into a plurality of word tokens.
5. The computer storage medium of claim 4 , wherein each of the plurality of word tokens is mapped to a second graph and assigned a weight, wherein the weight for the word token is related to a relevance between the corresponding word token and the input segment received via the user interface.
6. The computer storage medium of claim 5, wherein the weighted word tokens are compressed to generate the candidate sentences.
7. The computer storage medium of claim 1, wherein the generating of the final content comprises sorting at least a portion of the set of candidate sentences based on a candidate sentence ranking, the final content being combined to reduce information coverage redundancy and optimize overall coherence.
8. A computer-implemented method for generating content based on graph-based sentence compression using source content present in a retrieved corpus, the method comprising: Retrieve source content associated with an input segment received via a user interface; identifying, using a graphical representation of a plurality of sentences from the source content, a set of sentences in the related source content having overlapping information, wherein the overlapping information is determined based on a set of keywords associated with the input segment; generating candidate sentences by compressing content in the sentence set having overlapping information; generating a final content including a candidate sentence set, wherein the candidate sentence set includes the candidate sentence; as well as The final content is provided as content automatically created in response to the input fragment. 9 . The method of claim 8 , further comprising parsing the source content associated with the input segment into a plurality of sentences.
10. The method of claim 9, wherein each of the plurality of sentences is mapped to a first graph and assigned a weight, wherein the weight is related to a relevance between the corresponding sentence and the input segment received via the user interface. The method of claim 10 , wherein at least a portion of the weighted sentence is parsed into a plurality of word tokens.
12. The method of claim 11, wherein each of the plurality of word tokens is mapped to a second graph and assigned a weight, wherein the weight for the word token is related to a correlation between the corresponding word token and the input segment. The method of claim 12 , wherein the weighted word tokens are compressed to generate the candidate sentences.
14. The method of claim 8, wherein the generating of the final content comprises sorting at least a portion of the set of candidate sentences based on a candidate sentence ranking, the final content being combined to reduce information coverage redundancy and optimize overall coherence.
15. A computer system comprising: one or more processors; as well as a non-transitory computer-readable storage medium coupled to the one or more processors, the non-transitory computer-readable storage medium having instructions stored thereon that, when executed by the one or more processors, cause the computer system to provide: means for identifying a set of sentences in source content having overlapping information associated with an input segment received via a user interface, wherein the overlapping information is determined based on a set of keywords associated with the input segment received via the user interface; means for generating candidate sentences by compressing content in said sentence set having overlapping information; as well as Means for generating final content comprising a set of candidate sentences, the set of candidate sentences comprising the candidate sentence.
16. The system of claim 15, further comprising means for parsing the source content associated with the input segment into a plurality of sentences. 17 . The system of claim 16 , wherein each of the plurality of sentences is mapped to a first graph and assigned a weight, wherein the weight is related to a relevance between the corresponding sentence and the input segment. The system of claim 17 , wherein the weighted sentence is parsed into a plurality of word tokens.
19. The system of claim 18, wherein each of the word tokens is mapped to a second graph and assigned a weight, the weight for the word token being related to a correlation between the corresponding word token and the input segment.
20. The system of claim 19, wherein the weighted word tokens are compressed to generate the candidate sentences.