Generate presentation slides with extracted content.

By automatically extracting the content of data files through machine learning models and rules of a collaborative document system, presentation slides with extracted content are generated, solving the problems of low efficiency and consistency in slide generation in existing technologies, and improving processing speed and user experience.

CN115203399BActive Publication Date: 2026-04-03GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2017-11-10
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies require users to manually process a large amount of content when generating presentation slides, which leads to reduced computing device processing speed, significant impact on network bandwidth, difficulty in maintaining design consistency, and a monotonous user experience.

Method used

By using a collaborative document system, machine learning models and rules are used to automatically extract the content of data files and generate presentation slides with the extracted content, including layout templates and presentation visualizations, reducing content redundancy and improving processing efficiency.

Benefits of technology

It improves the processing speed and network bandwidth utilization for generating presentation slides, reduces file size, enhances slide design consistency and user experience, and reduces the user's operational burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115203399B_ABST
    Figure CN115203399B_ABST
Patent Text Reader

Abstract

This disclosure relates to generating presentation slides with extracted content. One method includes: receiving one or more data files as source material for slide generation; obtaining content from the one or more data files for use as slides in a slide presentation; identifying a layout template for the slides based on the content; and extracting the content into extracted content to generate presentation visualizations based on the extracted content. The extracted content includes a subset of the content. The method further includes generating the slides based on the presentation visualizations and the layout template.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Case Analysis

[0002] This application is a divisional application of Chinese Invention Patent Application No. 201711108164.6, filed on November 10, 2017. Technical Field

[0003] Various aspects and embodiments of this disclosure relate to electronic documents, and more specifically to generating presentation slides with extracted content. Background Technology

[0004] A presentation slide set can include a collection of slides that can be shown to one or more people to provide visual clarification during a presentation. Presentation slides can refer to display pages that include text, images, video, and / or audio for presentation to one or more people. For example, presentation slides may include short sentences or key points describing the information that the presenter will convey. To prepare a presentation slide set, the presenter may conduct research and collect one or more documents related to the presentation topic. The presenter may summarize the extensive textual content from these documents into brief descriptions or clarifications in the presentation slides. Furthermore, the presenter may design and arrange the layout and formatting of each slide, such as font size, color, background color, key point alignment, and / or animation configuration. Summary of the Invention

[0005] Various aspects and embodiments of this disclosure relate to generating slides with extracted content for use in PowerPoint presentations. One or more data files can be obtained as supporting material for slide generation. Content from one or more data files can be extracted automatically, or the user can select desired content to be extracted for slide generation. Layout templates can be identified and selected based on the type of extracted content. The extracted content can be refined into extractable content to generate presentation visualizations such as lists (e.g., key points), data charts, data tables, images, etc. Presentation slides can then be generated based on the presentation visualizations and layout templates. Attached Figure Description

[0006] The various aspects and embodiments of this disclosure will be more fully understood from the detailed description given below and the accompanying drawings, which are provided for explanation and understanding only.

[0007] Figure 1 An example of a system architecture used in embodiments of this disclosure is shown.

[0008] Figure 2The illustration is a flowchart of an aspect of a method for generating slides including refined content, according to one embodiment of the present disclosure.

[0009] Figure 3 An example slide presentation, according to an illustrative implementation, includes a set of slides generated from a data file.

[0010] Figure 4 The illustration is a flowchart of a method for representing summary text in a list of slides according to one embodiment of the present disclosure.

[0011] Figure 5 A more detailed example of slides in a slide presentation generated from a data file according to an illustrative implementation is shown.

[0012] Figure 6 This illustrates an example of sending selected portions of content from a data file to a slide presentation, according to an illustrative implementation.

[0013] Figure 7 An example is shown of receiving selected content and dividing the selected content into separate slides, according to an illustrative implementation.

[0014] Figure 8 The illustration is a flowchart of an aspect of a method for representing an image extracted from source material in a slide and text associated with that image, according to one embodiment of the present disclosure.

[0015] Figure 9 Examples of representations according to an illustrative embodiment, including images extracted from source material in a slide and text associated with those images, are shown.

[0016] Figure 10 The illustration is a flowchart of an aspect of a method for representing extracted data in a data chart on a slide, according to one embodiment of the present disclosure.

[0017] Figure 11 An example is shown where the extracted range of data is represented in a data chart on a slide, according to an illustrative embodiment.

[0018] Figure 12 Examples are shown of different extracted ranges of data represented in data tables on different slides according to illustrative embodiments.

[0019] Figure 13 The figure shows a block diagram of an example computing system operating according to one or more aspects of this disclosure. Detailed Implementation

[0020] Typically, when parsing source content and creating a PowerPoint presentation, users can perform many actions. For example, a user might need to locate and open every relevant data file on a specific topic. A user might need to parse a large portion of the source content and select key points from it to include in the PowerPoint presentation. A user might need to copy many sections of the source content to various slides of the PowerPoint presentation. In some cases, users may select a larger portion than required to be included in the slides to fully represent key points, facts, statistics, opinions, etc.

[0021] In some cases, generating large-format PowerPoint presentations can slow down computing devices and / or negatively impact network bandwidth when transmitting them over a network to user devices. Furthermore, users must define the presentation structure and create logical breakpoints for the slides. In some situations, users may create more slides than are needed to adequately represent the presentation topic. Larger presentations result in increased file sizes, negatively impacting processing speed and network bandwidth. Users must also choose a consistent and visually appealing design and apply it to every slide in the presentation. Therefore, it should be recognized that these actions can be monotonous for users and are not optimal for the performance of computing devices and / or networks.

[0022] The aspects and embodiments disclosed herein relate to a collaborative document system that addresses at least these deficiencies by generating presentation slides with refined content for PowerPoint presentations. The embodiments disclosed herein can be applied to any suitable data file including any appropriate content (e.g., text, tables, images, audio, video, etc.) to generate slides for PowerPoint presentations. For example, such a data file may include an electronic document uploaded by a user device or created using a collaborative document system.

[0023] Electronic documents refer to media content used in electronic form. Media content can include text, data tables, videos, images, graphics, slides, charts, software programming code, designs, lists, plans, blueprints, maps, etc. Electronic documents can be stored in a cloud-based environment. Electronic documents that users already have the right to access and / or edit can be referred to herein as collaborative documents. With collaborative documents, users are able to see content changes (e.g., character by character) as other collaborators edit the document. Although the collaborative document system is described in the remainder of this disclosure as implementing the disclosed techniques, it should be noted that any suitable system or application (e.g., a local application installed on a user's device) can generate slides for a slide presentation based on the content of one or more data files.

[0024] Collaborative document systems allow document owners to invite other users to join as collaborators on documents stored in a cloud-based environment. The collaborative document can be provided to the collaborators' user devices by one or more servers within the cloud-based environment. Each collaborator can be associated with a user type (e.g., editor, reviewer, viewer, etc.). Different views and capabilities can be provided to collaborators based on their user type to edit, comment on, review, or simply view the collaborative document. Once granted access to the collaborative document, collaborators can access it to perform actions permitted for their user type.

[0025] Using a collaborative document system, users can create or open collaborative documents (e.g., in a web browser) and share them with one or more collaborators. In some embodiments, the collaborative document can be a slide presentation automatically generated by a slide generation module based on the content of one or more data files. The slide generation module can receive one or more data files or selected content from one or more data files as input and extract certain content from the data files or selected content. The slide generation module can select one or more layout templates for one or more slides based on the type of content. For example, a layout template including a title can be selected for content that serves as a title in a data file, a layout template including a block header can be selected for content that serves as a block header in a data file, and / or a layout template including both a title and body text can be selected for content in a data file that includes text, data tables, images, etc. Therefore, in some embodiments, the format and style of the input data files can be maintained for the slides generated by the slide generation module, as further described below.

[0026] In some implementations, the extracted content can be refined into refined content using a slide generation module. Refining can refer to reducing the extracted content from a first amount of content to a second amount of content. For example, refining can refer to summarizing the text included in the extracted content from a first number of sentences to a second number of sentences that is less than the first number of sentences. Refining and summarizing are used interchangeably herein. In another example, refining can refer to reducing a data table in the extracted content to a selected range of data that is a subset of the entire data table. Furthermore, refining can refer to identifying an image in the content and extracting that image from the content. Presentation visualizations that include the refined content can be generated based on the type of the refined content. For example, presentation visualizations that include lists (e.g., key points) can be generated for refined content that has a text type, presentation visualizations that include data tables, data charts, or data graphs can be generated for refined content that includes images, and / or presentation visualizations that include images can be generated for refined content that includes images.

[0027] An example of how these technologies can be used could be a scenario where an employee wants to create a set of slides for a "sales overview for the fiscal year." The employee might have market reports, spreadsheets of sales data, documents about new products, etc. The employee can upload these documents to a slide generation module for automatic content extraction and slide generation. The module can extract sentences from a market report to include in a list of key points in the first slide. It can also extract sales data from a spreadsheet of sales data and select data charts to visually represent the data, or select tables to plot the sales data in the second slide. Furthermore, the module can extract product images from documents about new products and extract text associated with those images to generate a product introduction slide in the third slide that includes both images and relevant text.

[0028] The disclosed techniques improve processing speed by presenting content from data files more efficiently by refining it into a reduced format. For example, in some implementations, text in a data file can be summarized into a reduced set of sentences before being presented as a list of key points in a slide. Refining the content reduces the file size of the slide presentation. Simultaneously, these techniques can determine logical breakpoints in the content based on headers, formatting, content size, etc., to create an effective number of slides to adequately represent the presentation topic. Network bandwidth can be improved by sending slide presentations with reduced file sizes or a more efficient number of slides over a network. These techniques can also structure the slide presentation based on the formatting of one or more data files and apply a consistent theme / design to the slides to enhance the graphical appearance of the presentation and maintain a shared look between the data file and the slide presentation. Furthermore, by automatically reducing and refining content and automatically creating slides for the presentation topic, the disclosed techniques increase the reliability of collaborative document systems and reduce or eliminate the need for manual review of the results of these techniques.

[0029] Figure 1This is an example of a system architecture 100 for implementing the present disclosure. System architecture 100 includes a cloud-based environment 110 connected to user devices 120A-120Z via a network 130. Although system architecture 100 is described in the context of cloud-based environment 110, which enables communication between servers 112A-112Z in cloud-based environment 110 and communication with user devices 120A-120Z on network 130 for storing and sharing data, it should be understood that the embodiments described herein can also be applied to locally interconnected systems. Cloud-based environment 110 refers to a collection of physical machines hosting applications (e.g., word processing applications, spreadsheet applications, PowerPoint presentation applications) that provide one or more services (e.g., word processing, spreadsheet processing, PowerPoint presentations) to multiple user devices 120A-120Z via network 130. Network 130 may be a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or a wide area network (WAN)) or a combination thereof. Network 130 may include wireless infrastructure that can be provided by one or more wireless communication systems, such as Wi-Fi hotspots connected to network 130 and / or wireless carrier systems that can be implemented using various data processing devices, communication towers, etc. Alternatively or additionally, network 130 may include wired infrastructure (e.g., Ethernet).

[0030] The cloud-based environment 110 may include one or more servers 112A-112Z, a training engine 115, and / or a data storage 114. The training engine 115 and / or data storage 114 may be decoupled from and communicatively coupled to servers 112A-112Z, or the training engine 115 and / or data storage 114 may be a part of one or more of servers 112A-112Z. Data storage 114 may store data files 116 that include content such as text, data tables, images, videos, audio, etc. In one embodiment, data file 116 may be any suitable data file, including content uploaded by user devices 120A-120Z to the cloud-based environment 110 or from a server within or outside the cloud-based environment 110. In another embodiment, data file 116 may be a collaborative document shared with one or more users. Collaborative documents can be word processing documents, spreadsheet documents, or any suitable electronic documents that can be shared with users (e.g., electronic documents containing content such as text, data tables, videos, images, graphics, slides, charts, software programming code, designs, lists, plans, blueprints, maps, etc.).

[0031] Collaborative documents can be created by an author, who can then share them with other users (such as collaborators). Sharing a collaborative document can mean authorizing other users to access (view and / or edit) the document. Sharing a collaborative document can include informing other users about the document via a message (e.g., email, text message, etc.) that includes a link to the document. The level of permissions granted to each user can be based on their specific user type. For example, a user with the "Editor" user type can open the collaborative document and make direct changes. Similarly, multiple collaborators can make changes to the content presented in the collaborative document.

[0032] Training engine 115 may include one or more processing devices, such as a computer, microprocessor, logic device, or other device or processor configured with hardware, firmware, and software to perform some of the embodiments described herein. Training engine 115 may include, or have access to, a set of training data files and a corresponding overview for each training data file, which is used by training engine 115 as training data to train machine learning model 113 to perform extraction-based generalization. Machine learning model 113 may refer to a model artifact created by training engine 115 using training inputs and corresponding target outputs. Training inputs may include the set of training data files, and corresponding target outputs may include an overview of each training input. In some embodiments, training data files and corresponding target outputs may include a specific format (e.g., a list of points). Machine learning model 113 may use the training inputs and target outputs to learn features of words, phrases, or sentences in the text such that they are good candidates to be included in the overview (extracted content). These features may include position in the text (e.g., the first sentence may be the topic sentence and provide a good overview of the paragraph, the preceding sentences may be relevant, and the last sentence may be a conclusion and relevant), frequent words or phrases, the number of words in a sentence, etc. Once trained, machine learning model 113 can be applied to a new data file 116 to obtain an overview (extracted content) for that new data file 116. In some implementations, the extracted content can be used to generate presentation visualizations for inclusion in a layout template of a new slide. In some implementations, machine learning model 113 can learn the format of text to output extracted content using specific presentation visualizations (e.g., a list of points).

[0033] Servers 112A-112Z can be physical machines (e.g., server machines, desktop computers, etc.), each physical machine including one or more processing devices communicatively coupled to memory devices and input / output (I / O) devices. The processing devices may include computers, microprocessors, logic devices, or other devices or processors configured with hardware, firmware, and software to perform some of the embodiments described herein. Each of servers 112A-112Z can host slide generation modules (118A-118Z). Slide generation modules 118A-118Z can be implemented as computer instructions executable by one or more processing devices on each of servers 112A-112Z. Slide generation modules 118A-118Z can generate slide presentations 117 with slides having content extracted from one or more data files 116 (e.g., collaborative documents). Users can manually identify one or more data files 116 as supporting material for the slide generation modules 118A-118Z used to generate the slide presentation 117, or users can identify specific content portions of one or more data files 116 as supporting material for the slide generation modules 118A-118Z used to generate the slide presentation 117. The slide presentation 117 can be shared with one or more users and can be a collaborative document.

[0034] In some implementations, the slide generation modules 118A-118Z can identify a layout template for each slide in the slide presentation 117 based on content. Various layout templates may include a “title” layout template for the presentation’s title, a “section header” layout template for intermediate headers (e.g., headers not associated with body content), a “heading plus body” layout template for a parent heading with associated body content (e.g., text, data, images, etc.), a “heading plus body” layout template for body content without an associated parent heading, and so on. It should be understood that any suitable layout template can be used.

[0035] The slide generation modules 118A-118Z can refine content into extracted content to generate presentation visualizations based on the extracted content. As described above, in one embodiment, the slide generation modules 118A-118Z can apply content as input to a machine learning model 113, which is trained to produce the extracted content as the target output. In one embodiment, the slide generation modules 118A-118Z can use one or more rules 119 that define heuristics for extracting content. Rules 119 can be predefined by the developer. Rules 119 can be applied to content included in one or more data files 116 to generate a slide presentation 117 to extract the content.

[0036] For example, if the content is text, Rule 119 could define that the text to be included in a slide does not overflow the slide. In this case, the text can be refined into different subsets, and each subset can be included in a different slide, so that the subsets of text fit appropriately into different slides. Rule 119 could define that complete sentences or individual bullet points are not broken down when text is separated between slides. Another Rule 119 could define the sentences used for bullet points based on the sentence's position in the paragraph (e.g., the first sentence in the paragraph of text is used as the bullet point of the slide because the first sentence may be the topic sentence; or the last sentence in the paragraph is used as the bullet point because the last sentence may include the conclusion). Another Rule 119 could define that words or phrases that frequently appear in the body of the text will be refined, and individual sentences will be included as bullet points with frequently occurring words or phrases, while some other sentences with less frequently occurring words or phrases will be ignored. Another Rule 119 could define the maximum number of sentences refined for representation in the slide.

[0037] If the content is data from a data table, another rule 119 can define that certain column headers in the data table are identified, and a range of data associated with these column headers is extracted (e.g., for sales data, rule 119 can define that the data range associated with column headers such as "sales," "region," etc., is extracted from the data table), while ignoring data from other data tables associated with column headers. Additionally, rule 119 can define that data ranges of a specific size will be selected to be represented in the data table to fit correctly within the slide so that the data table does not overflow the slide. In this case, the rule can define creating another slide for the remaining data ranges not used in the first slide, and can define reusing column headers from the data table in the first slide in the data table of the second slide.

[0038] If the content is an image, another rule 119 can define extracting the image from the rest of the content and using that image as the refined content. Another rule 119 can use a caption associated with the image to define the title of the slide that includes the image. Alternatively, if the image does not have a caption, rule 119 can use the text closest to the image to define the title of the slide that includes the image. Another rule 119 can define including text surrounding the image in the notes section associated with the slide, but not actually including text within the slide itself.

[0039] One or more rules 119 may also define which presentation visualization to generate based at least on the type of content in one or more data files 116. For example, if the content is text, rule 119 may define generating a list (points) as a presentation visualization; if the content is data, rule 119 may define generating a data table or data chart as a presentation visualization; if the content is an image, rule 119 may define generating a data chart as a presentation visualization, and so on. Rule 119 may define including presentation visualizations in the body portion of the identified layout template used for the content.

[0040] The slide generation modules 118A-118Z can generate presentation visualizations (e.g., list of points, data charts, data tables, images, etc.) based on extracted content. The slide generation modules 118A-118Z can also generate slide presentations 117 with or without presentation visualizations, based on layout templates selected for specific content. For example, a "Title" layout template may not include presentation visualizations, but a "Title plus Body" layout template may include them.

[0041] One or more servers 112A-112Z can provide a collaborative document environment 122A-122Z to user devices 120A-120Z. The servers 112A-112Z selected to provide the collaborative document environment 122A-122Z can be based on certain load balancing technologies, service level agreements, performance metrics, etc. The collaborative document environment 122A-122Z can provide a user interface 124A-124Z for displaying a slide presentation 117 generated based on the content of one or more data files 116. The collaborative document environment 122A-122Z allows users using different user devices 120A-120Z to simultaneously access the slide presentation 117 to review, edit, view, and / or suggest changes to the slide presentation 117 in their respective user interfaces 124A-124Z. In one embodiment, the user interface 124A-124Z can be a webpage rendered by a web browser and displayed in a web browser window on the user device 120A-120Z. In another embodiment, the user interface 124A-124Z may be included in a standalone application downloaded to the user device 120A-120Z and natively run on the user device 120A-120Z.

[0042] User devices 120A-120Z may include one or more processing devices communicatively coupled to memory devices and I / O devices. User devices 120A-120Z may be desktop computers, laptop computers, tablet computers, mobile phones (e.g., smartphones), or any suitable computing device. User devices 120A-120Z may include components such as input devices and output devices. Users can be authenticated via server 112A-112Z using a username and password (or other identification information) provided by the user through user interface 124A-124Z, allowing the same user device 120A-120Z to be used by different users at different times. In some embodiments, slide generation modules 118A-118Z may be part of user devices 120A-120Z. For example, in some embodiments, user devices 120A-120Z may have locally installed applications, including slide generation modules 118A-118Z for users to access, view, edit, and / or automatically generate slide presentations 117 with extracted content.

[0043] Figure 2A flowchart depicts aspects of a method 200 for generating slides including refined content according to one embodiment of the present disclosure. Each of the method 200 and its various functions, routines, subroutines, or operations can be executed by one or more processing devices of a computer device performing the method. In some embodiments, method 200 can be executed by a single processing thread. Alternatively, method 200 can be executed by two or more processing threads, each thread executing one or more individual functions, routines, subroutines, or operations of the method. In illustrative examples, the processing threads implementing method 200 can be synchronized (e.g., using semaphores, critical blocks, and / or other thread synchronization mechanisms). Alternatively, the processes implementing method 200 can be executed asynchronously relative to each other.

[0044] For simplicity, the methods of this disclosure are depicted as a series of actions. However, the actions according to this disclosure may occur in various orders and / or simultaneously, and may occur with other actions not presented and described herein. Furthermore, not all illustrated actions may require implementation of the methods according to the disclosed subject matter. Moreover, those skilled in the art will understand and recognize that these methods may alternatively be represented as a series of interrelated states via state diagrams or events. Furthermore, it should be understood that the methods disclosed in this specification can be stored on an article to facilitate the transmission and delivery of such methods to a computing device. The term "article of manufacture" as used herein is intended to include a computer program accessible from any computer-readable device or storage medium. In one embodiment, method 200 may be executed by one or more slide generation modules 118A-118Z executed by one or more processing devices of servers 112A-112Z in a cloud-based environment 110. In some embodiments, method 200 may be executed by one or more processing devices of user devices 120A-120Z executing slide generation modules 118A-118Z.

[0045] Method 200 may begin at box 202. At box 202, the processing device may receive one or more data files 116 as source material for slide generation. In one embodiment, the one or more data files 116 may include collaborative documents (e.g., text documents, spreadsheet documents, etc.), non-collaborative documents (e.g., text documents, spreadsheet documents), saved web pages, database files, image files, video files, audio files, animated content, or any suitable media files. Data files 116 may include content such as text, data tables, images, etc. Data files 116 may be uploaded to a cloud-based environment 110 or created and stored in data storage 114 using a collaborative document environment 122A.

[0046] At box 204, the processing device can obtain the content of the slides for the slide presentation 117 from one or more data files 116. The processing device can parse the data file 116 to identify the content and automatically extract the content, including any applicable formatting (e.g., titles, block headers, parent headers with body content, etc.) and styles, as shown in the following reference. Figure 4-5 As shown. In some embodiments, the user can select content (e.g., a text paragraph, a data table from a spreadsheet, one or more images), and the processing device can obtain the user-selected content from one or more data files 116, as referenced below. Figure 6 As shown.

[0047] At box 206, the processing device can identify the layout template of the slides based on the content. As described above, the layout templates can include a "heading" layout template for describing the presentation topic based on the headings or top-level headers in the content of data file 116, a "block header" layout template for describing intermediate headers (e.g., headers unrelated to the body content), and a "heading plus body" layout template for displaying body content (e.g., text, data, images, etc.) in the body portion of the slides and displaying a parent header associated with the body content as the slide title. In some cases, the slide generation modules 118A-118Z can recognize the format of the content to identify the appropriate layout template. Therefore, the formatting of data file 116 can be maintained for the slide presentation 117.

[0048] For example, headings, headers, body content, etc., can be identified in the content by parsing the document and / or using the metadata of data file 116. If the content includes headings, a “Heading” layout template can be identified and the text of the heading can be selected. If the content includes intermediate headers (e.g., block headers unrelated to the body content), a “Block Header” layout template can be identified and the text of a specific intermediate header can be selected. If the content includes body content with an associated parent header, a “Heading Plus Body” layout template can be identified and selected, and once extracted, the body content can be included in the body of the slide, and the text of the parent header may be included in the title of the slide. For example, if the content includes a parent header and associated text (e.g., key points), the text of the parent header can be extracted and set as the title of the slide, and the body of the slide can be specified after the text is extracted. If the content includes body content without an associated parent header, headings can be automatically selected based on the body content. For example, rule 119 can define certain keywords to search within content that can be selected and used as headings.

[0049] At block 208, the processing device can refine the content into refined content to generate presentation visualizations based on the refined content. The processing device can apply machine learning model 113 or rule 119 to obtain the refined content. The refined content includes a subset of the original content. The generated presentation visualizations can be based on the application of one of the rules in rule 119. For example, when the layout template is “Title and Body” and the content includes text with an associated parent header, rule 119 can define the presentation visualizations that select the list of points to be generated, and the text representing the refined content in the list of points in the body of the layout template, and the text of the parent header can be set as the title of the layout template for that slide.

[0050] In some embodiments, the original content can be refined by applying a machine learning model 113 to the original content. As described above, the machine learning model 113 can be trained to select certain features from the content (e.g., sentences in certain positions within a paragraph, frequently used words or phrases, etc.) and output refined content. In some embodiments, the machine learning model 113 can use specific presentation visualizations (e.g., a list of key points) to output refined content. In some embodiments, one or more rules 119 can define which presentation visualization is used for content refinement. For example, a list of key points can be generated for refined content that includes text.

[0051] In some embodiments, the original content can be refined according to Rule 119. For example, Rule 119 may define a subset (e.g., a maximum number) of sentences to be selected from the content to refine the content for representation in a presentation visualization (e.g., a list of points). Rule 119 may also define which sentences to select based on their position in a paragraph (e.g., the first sentence in a paragraph, the first two or three sentences (any number) in a paragraph, the last sentence in a paragraph), based on frequently used words or phrases, etc. Rule 119 may also define a range of data to be selected from a data table when the content includes a data table, and may define the presentation visualization (e.g., a data chart, a data table) to be selected to represent that range of data. For example, Rule 119 may define which column headers to search when selecting a range of data, and when a column header is found, Rule 119 may define a range of data associated with the column header to be selected. Furthermore, when certain column headers are found, there may be a mapping from the column header to a specific data chart. For example, the column header for "Sales" may be mapped to a bar chart. Rule 119 can also define how to extract images when the content includes images (e.g., extracting an image as a single object without cropping it).

[0052] At box 210, the processing device can generate slides based on presentation visualizations and layout templates. For example, slides with a "title and body" layout template can be generated, and presentation visualizations including a list of key points derived from the original text can be included in the body of the layout template, while text associated with the parent header of the derived sentences can be included in the title of the layout template used for the slides. A default theme can be applied to each slide in the slide presentation 117 to provide consistency and an enhanced look and feel to the slide presentation 117. Once created, the user can configure the default theme and / or modify the theme of the slide presentation 117.

[0053] In some embodiments, the original content (e.g., text paragraphs or original data tables) may be stored in a notes section of a slide that includes the corresponding extracted content. The notes section can provide further context to the presenter during the presentation. Additionally, in some embodiments, the processing device may receive user interactions with the generated slides. The processing device can use user interactions (e.g., any editing, modification, reformatting, etc.) to update or create new rules 119 for defining heuristics about how to generate subsequent slides for a particular user. For example, the user's preferred font type, font size or color, paragraph formatting, point style, etc., may be captured for a specific layout template. As another example, if the user rewrites automatically extracted point sentences in a slide, similar language may be stored and applied to other slides. Alternatively, the processing device may update (retrain) a machine learning model 113 based on user interactions.

[0054] In some implementations, participant feedback can be obtained to improve the slide presentation 117. For example, participant feedback can be provided via a rating system used for the presentation. In some implementations, participant engagement and response data can be recorded (e.g., where / which slides users spent the most time on, which slides users commented or messaged the most on, etc.). The processing device can generate subsequent slides based on the participant engagement and response data, including similar layout templates, refined content, design / theme, style, etc.

[0055] In some embodiments, the processing device may obtain second content (e.g., text, data, images, etc.) from one or more data files 116 for use as a second slide in a slide presentation 117. The processing device may identify a second layout template for the second slide based on the second content. Depending on the type of the second content, the second layout template may differ from the layout template selected for the first slide. The processing device may refine the second content into second refined content to generate a second presentation visualization based on the second refined content. A machine learning model 113 or rule 119 may be applied to the second content to refine it. The second refined content may include a subset of the second content. The processing device may generate a second slide based on the second presentation visualization and the second layout template. It should be understood that, where appropriate, the process may continue to generate as many slides as possible until the content of one or more data files 116 is included in the corresponding slide of the slide presentation 117.

[0056] Figure 3 An example of a slide presentation 117, comprising a set of slides generated from a data file 116, is shown according to an illustrative embodiment. As depicted, a collaborative document environment 122A is provided from a server 112A and displayed via a user interface 124A. The data file 116 is opened in a collaborative word processing application provided by the collaborative document environment 122A in a first browser window 300. It should be understood that the collaborative document environment 122A can be displayed in the user interface 124A of the native application on the first browser window 300 without using a browser. The depicted data file 116 includes text (e.g., titles, section headers, parent headers with associated text, etc.).

[0057] Users can access the file menu options (“Tools”) using the opened data file 116 and select option (e.g., link) 302 (“Generate Slide Presentation”) to generate a slide presentation 117. After selecting option 302, the data file 116 can be received by the slide generation module 118A as source material for slide generation. The slide generation module 118A can obtain content from the data file 116 by recognizing and extracting content. In some cases, the format of the content can be obtained. For example, the slide generation module 118A can determine the formatting information associated with the content (e.g., titles, block headings, parent headings associated with the content, content, etc.). The slide generation module 118A can recognize the layout templates used for various parts of the text. For example, the various parts can be determined based on the formatting information.

[0058] Different layout templates can be selected for different sections. For example, for a section of text with formatted information indicating that the text is a title, the "Title" layout template can be selected; for a section of text with formatted information indicating that the text is a block heading (e.g., a title not associated with the body content), the "Block Header" layout template can be selected; for a section of text with formatted information indicating that the text includes both the parent heading and the body text associated with the parent heading, the "Title Plus Body" layout template can be selected, and so on. The slide generation module 118A can refine the content into refined content to generate presentation visualizations based on the refined content.

[0059] The slide generation module 118A can generate one or more slides to be included in the slide presentation 117 based on layout templates and / or presentation visualizations. As depicted, the slide presentation 117 is displayed by a collaborative slideshow application in a second browser window 304, separate from the collaborative word processing application displayed in the first browser window 300. Slides in the slide presentation 117 with a “title plus body” layout template may include presentation visualizations containing lists of refined text, as described in more detail below.

[0060] Figure 4 A flowchart illustrating various aspects of a method 400 for representing summary text in a list in a slide according to one embodiment of this disclosure is provided. Method 400 may be performed in the same or similar manner as described above with respect to method 200. In one embodiment, method 400 may be performed by one or more slide generation modules 118A-118Z executed by one or more processing devices of servers 112A-112Z in a cloud-based environment 110. In some embodiments, method 400 may be performed by one or more processing devices of user devices 120A-120Z executing slide generation modules 118A-118Z.

[0061] Before method 400 begins, the processing device may have already received one or more data files 116 and extracted content from one or more data files 116. Alternatively, the processing device may have received a selection of content from one or more data files 116 before method 400 begins. The content may include text (e.g., one or more parent headers, paragraphs of text, or including sentence summaries, etc.).

[0062] Method 400 can begin at box 402. At box 402, the processing device can extract text from the content, including a first set of sentences. The first set of sentences can be in paragraph form or can be represented in a list (e.g., a list of points). At box 404, the processing device can summarize the first set of sentences into a second set of sentences as refined content. The second set of sentences can include fewer sentences than the first set. Summarization can be performed by applying machine learning model 113 to the first set of sentences or by applying one or more rules 119 to the first set of sentences. At box 406, the processing device can generate a presentation visualization item including a list (e.g., points) based on the second set of sentences. For example, each sentence in the second set of sentences can be represented as a separate entry in the list (e.g., a point).

[0063] Figure 5 A more detailed example of slides 500, 502, and 504 in a slide presentation 117 generated from data file 116 according to an illustrative implementation is shown. As depicted, the data file includes text content formatted with a title (“Marketing Plan”), a block header unrelated to the body content (“Goals”), and a parent header associated with the original body text 505 (“Personal Goals (Marketing Director)”). The original body text 505 includes two sentences represented in a list of points (“Devote 20 hours per month to monthly marketing theme” and “Speak at 20 events in FY 2013”).

[0064] The slide generation module 118A can receive data file 116, extract content, and generate slides 500, 502, and 504. The slide generation module 118A can recognize the layout templates for different parts of the text in data file 116. For example, for the title section (“Marketing Plan”), the “Title” layout template is selected, and the text of the title section is set to the title in the layout template described in slide 500. For the block header section (“Goals”), the “Block Header” layout template is selected, and the text of the block header is set to the title in the layout template described in 502. For the parent header (“PersonalGoals (Marketing Director)”) and associated text sections, the “Title plus Body” layout template is recognized, and the text of the parent header is set as the title of slide 504, and the associated text is set as the body text of slide 504. Therefore, each slide 500, 502, and 504 includes different parts of the text from data file 116 and different layout templates. In this way, the structure of the slide presentation 117 can be mapped to the format of the original data file 116.

[0065] As shown in slide 504, the extracted content 508 is used to generate the presentation visualization 506. In this example, the presentation visualization 506 is a list of points, but it should be understood that any suitable list or presentation visualization can be used. The presentation visualization 506 is included in the body text of the layout template. The extracted text 508 can be generated by applying machine learning model 113 or rule 119 to the original body text 505. The extracted text 508 contains fewer sentences than the original body text 505.

[0066] Figure 6The illustration shows an example of how a portion of content, according to an illustrative embodiment, can be selected to be sent from data file 116 to a slide presentation 117. As depicted, data file 116 is a collaborative word processing document displayed in a first browser window 300 of user interface 124A by a collaborative word processing application of collaborative document environment 122A. In some cases, a user may want to create slides only for some content in data file 116. Therefore, the user can select (e.g., highlight) text portion 600. In the depicted example, text portion 600 includes a first parent heading (“Tactical Goals”) and associated body text (three sentences represented in the list of points), and a second parent heading (“Strategic Goals”) and associated body text (three sentences). An options menu 602 may appear when text portion 600 is selected or when input is received from an input peripheral device (e.g., selecting a mouse button). From options menu 602, the user can select an option (e.g., a link) 604 (“Send to Slides”), and another options menu 606 may appear, including the available slide presentation 117. Options menu 606 also allows the user to create a new slide presentation 117 using the selected text portion 600. From options menu 606, the user can select a link 608 to send the selected text portion 600 to the desired slide presentation 117 (“Marketing Plan”).

[0067] The slide generation module 118A can receive a selected text portion 600 and can identify the layout template used for the text portion 600. In some embodiments, the slide generation module 118A can determine that the formatting information of the selected text portion 600 indicates the existence of two different parent headers and two corresponding text bodies. Therefore, the slide generation module 118A can divide a single selection of the text portion 600 into two slides.

[0068] For example, Figure 7The illustration shows an example of receiving selected content (text portion 600) and dividing the selected content into separate slides 700 and 702 according to an illustrative implementation. A parent header can be used as a logical breakpoint to divide the selected content into separate slides 700 and 702. The layout template for slide 700 can be "title plus body text," where the parent header ("Tactical Goals") is set as the title of slide 700, and the body text associated with the parent header is set as the body text in slide 700. Similarly, the layout template for slide 702 can be "title plus body text," where the parent header ("Strategic Goals") is set as the title of slide 702, and the body text associated with the parent header is set as the body text in slide 702. More specifically, the body text can be refined into refined text 704 and 706 by applying machine learning model 113 or rule 119, and presentation visualizations 708 and 710 (e.g., a list of points) can be generated based on refined text 704 and 706. It should be understood that both texts 704 and 706 contain fewer sentences than their respective texts. Figure 6 The original text corresponding to data file 116 in the data file.

[0069] Figure 8 A flowchart depicts aspects of a method for representing images and text associated with images extracted from source material in a slide, according to one embodiment of this disclosure. Method 800 can be performed in the same or similar manner as described above with respect to method 200. In one embodiment, method 800 can be performed by one or more slide generation modules 118A-118Z executed by one or more processing devices of servers 112A-112Z in a cloud-based environment 110. In some embodiments, method 800 can be performed by one or more processing devices of user devices 120A-120Z executing slide generation modules 118A-118Z.

[0070] Before method 800 begins, the processing device may have received one or more data files 116 and content extracted from them. Alternatively, before method 800 begins, the processing device may have received a selection of content from one or more data files 116. The content may include images and text associated with the images.

[0071] Method 800 can begin at box 802. At box 802, the processing device can extract an image from the content as refined content. The image can be recognized by the processing device while it is parsing the data file 116. One or more rules 119 defining how the image is extracted can be applied to the image. For example, rule 119 can define that the image will be extracted as a single object and should not be cropped. Rule 119 can also define how the image is resized to fit appropriately within the body of a layout template (e.g., "heading plus body"). At box 804, the processing device can generate a presentation visualization that includes the image.

[0072] Figure 9 An example is shown representing an image 900 extracted from a data file 116 in slide 904 and text 902 associated with the image 900, according to an illustrative embodiment. The depicted data file 116 includes various texts 906 describing the image 900 and descriptive text 902 associated with the image 900 (“Product XYZ”). The slide generation module 118A can identify and extract the image 900 from the data file 116 as refined content, and identify a layout template (“title plus body text”) for the refined content. The slide generation module 118A can generate a presentation visualization 908 that includes the image 900 extracted and resized according to one or more rules 119. The presentation visualization 908 can be included in the body text of the layout template.

[0073] In some embodiments, certain text in data file 116 may be extracted and set as the title 910 of slide 904, which includes image 902. For example, as depicted, explanatory text 902 may be extracted and set as the title 910 of slide 904. If explanatory text is not present in data file 116, one or more words, phrases, or sentences of various texts 906 describing the product may be extracted and set as the title 910 of slide 904.

[0074] Figure 10 A flowchart illustrating aspects of a method 1000 for representing extracted data in a data chart in a slide according to one embodiment of the present disclosure is provided. Method 1000 may be performed in the same or similar manner as described above with respect to method 200. In one embodiment, method 1000 may be performed by one or more slide generation modules 118A-118Z executed by one or more processing devices of servers 112A-112Z in a cloud-based environment 110. In some embodiments, method 1000 may be performed by one or more processing devices of user devices 120A-120Z executing slide generation modules 118A-118Z.

[0075] Before method 1000 begins, the processing device may have received one or more data files 116 and content extracted from them. Alternatively, the processing device may receive a selection of content from one or more data files 116 before method 400 begins. The content may include a data table containing data.

[0076] Method 1000 can begin at box 1002. At box 1002, the processing device can extract a data table from the content. At box 1004, the processing device can select a range of data from the data table as the refined content. One or more rules 119 can be applied to the data table of the content to select this data range. One or more rules 119 can define which column headers to search in the data table, and, if found, the range of data to extract. Rule 119 can also define which data chart to use based on the mapping between the data chart and the identified column headers, such as... Figure 11 As shown. One or more rules 119 may also define the maximum number of rows that can be selected to fit appropriately within a data table in a slide, and correspondingly, a range of data can be selected. Any additional rows may be included in a separate data table in one or more other slides, such as Figure 12 As shown. In box 1006, the processing device can generate presentation visualizations, including data charts, based on the data within this range.

[0077] Figure 11An example of data 1100 representing an extracted range in data chart 1102 in slide 1104, according to an illustrative embodiment, is shown. The depicted data file 116 includes data table 1106. Slide generation module 118A can identify and extract data table 1106 and apply one or more rules 119 to it. These rules 119 can define a range of data 1100 to be extracted, associated with certain column headers (e.g., “Sales” and “Region”). The defined column headers may relate to one or more criteria presented regarding an entity (e.g., sales, finance, inventor, product, etc.). Description text 1108 may also be associated with data table 1106 in data file 116. Slide generation module 118A can select data range 1100 from the extracted data table 1106 as refined content to generate a presentation visualization including data chart 1102 based on the data range 1100. As depicted, the data chart includes regional and sales data associated with each region. In one example, the bar chart is selected based on rule 119, which defines the mapping between the bar chart and column headers related to sales information (such as "Sales" and "Region"). The presentation visualizations, including data chart 1102, can be included in the body of the layout template ("title plus body") used to generate slide 1104.

[0078] In some embodiments, certain text in data file 116 may be extracted and set as the title 1110 of slide 1104, which includes data chart 1102. For example, as depicted, explanatory text 1108 may be extracted and set as the title 1110 of slide 1104. If explanatory text is not present in data file 116, one or more words, phrases, or sentences near the text of data table 1106 in data file 116 may be extracted and set as the title 1110 of slide 904.

[0079] Figure 12Examples of data 1202 and 1204 representing different extracted ranges in data tables 1206 and 1208 in different slides 1210 and 1212, according to an illustrative embodiment, are shown. The depicted data file 116 includes data table 1214. A slide generation module 118A can identify and extract data table 1214 and apply one or more rules 119 to data table 1214. One or more rules 119 can define the maximum number of rows selected to fit within a single slide. In the described example, the maximum number of rows is two, but it should be understood that any suitable number can be used. A first range of data 1202 is selected to have two rows, and a presentation visualization including data table 1206 is generated based on the first range of data 1202. A second range of data 1204 is selected to have two rows, and a presentation visualization including data table 1208 is generated based on the second range of data 1204. In other embodiments, rule 119 can define the range of data selected from the data table based on matching values ​​of specific columns. For example, a data range including the "CA" value in the "Region" column can be selected as the first range, data 1202, and a data range including the "TX" value in the "Region" column can be selected as the second range, data 1204. Rule 119 can define column headers for both data tables 1206 and 1208. Presentation visualizations including data tables 1206 and 1208 can be included in the body of the layout template ("Title and Body") to generate the corresponding slides 1210 and 1212.

[0080] In some embodiments, certain text in data file 116 may be extracted and set as titles 1216 and 1212 of slides 1210 and 1212, including data tables 1206 and 1208. For example, as depicted, the descriptive text 1220 of data table 1214 in data file 116 may be extracted and set as titles 1216 and 1218 of slides 1210 and 1212. If no descriptive text exists in data file 116, one or more words, phrases, or sentences of text near data table 1214 in data file 116 may be extracted and set as titles 1216 and 1218 of slides 1210 and 1212.

[0081] Figure 13 A block diagram depicts an example computing system operating according to one or more aspects of this disclosure. In various illustrative examples, computer system 1300 may correspond to... Figure 1The system architecture 1300 can be any computing device within the system architecture 100. In one embodiment, the computer system 1300 can be any of the servers 112A-112Z or the training engine 115. In another embodiment, the computer system 1300 can be any of the user devices 120A-120Z.

[0082] In some implementations, computer system 1300 may be connected (e.g., via a network such as a local area network (LAN), intranet, extranet, or the Internet) to other computer systems. Computer system 1300 may operate as a server or client computer in a client-server environment, or as a peer-to-peer computer in a peer-to-peer or distributed network environment. Computer system 1300 may be provided by a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), cellular phone, web device, server, network router, switch, or bridge, or any device capable of (sequentially or otherwise) executing a set of instructions specifying actions to be taken by that device. Furthermore, the term "computer" should include any collection of computers that individually or jointly execute a set of instructions (or multiple sets of instructions) to perform any one or more methods described herein.

[0083] In a further aspect, the computer system 1300 may include a processing device 1302, a volatile memory 1304 (e.g., random access memory (RAM)), a non-volatile memory 1306 (e.g., read-only memory (ROM) or electrically erasable programmable ROM (EEPROM)), and a data storage device 1316, all of which may communicate with each other via a bus 1308.

[0084] The processing device 1302 may be provided by one or more processors, such as general-purpose processors (e.g., complex instruction set computing (CISC) microprocessors, reduced instruction set computing (RISC) microprocessors, very long instruction word (VLIW) microprocessors, microprocessors that implement other types of instruction sets, or microprocessors that implement combined types of instruction sets) or special-purpose processors (e.g., application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), or network processors).

[0085] The computer system 1300 may also include a network interface device 1322. The computer system 1300 may also include a video display unit 1310 (e.g., LCD), an alphanumeric input device 1312 (e.g., keyboard), a cursor control device 1314 (e.g., mouse), and a signal generation device 1320.

[0086] Data storage device 1316 may include a non-transitory computer-readable storage medium 1324 on which instructions 1326 encoding any one or more of the methods or functions described herein may be stored, including implementations of slide generation modules 118 (118A-118Z) and / or Figure 1 The instructions of the training engine 115 are used to implement any of the methods described herein.

[0087] Instruction 1326 may also reside wholly or partially in volatile memory 1304 and / or in processing device 1302 during execution by computer system 1300, thus volatile memory 1304 and processing device 1302 may also constitute machine-readable storage media.

[0088] Although computer-readable storage medium 1324 is shown as a single medium in the illustrative example, the term "computer-readable storage medium" will include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more sets of executable instructions. The term "computer-readable storage medium" will also include any tangible medium capable of storing or encoding a set of instructions that, when executed by a computer, cause the computer to perform any one or more methods described herein. The term "computer-readable storage medium" will include, but is not limited to, solid-state memory, optical media, and magnetic media.

[0089] Numerous details have been set forth in the foregoing description. However, it will be apparent to those skilled in the art who benefit from this disclosure that it can be practiced without these specific details. In some instances, well-known structures and devices are shown in block diagram form rather than in detail in order to avoid obscuring this disclosure.

[0090] Parts of a specific implementation have been presented based on algorithms and symbolic representations of operations on data bits within computer memory. These algorithmic descriptions and representations are means by which those skilled in the art of data processing most effectively communicate the essence of their work to others skilled in the art. Algorithms here are generally considered to be self-consistent sequences of steps leading to desired results. These steps are those requiring physical manipulation of physical quantities. Typically, but not necessarily, these quantities take the form of electrical or magnetic signals that can be stored, transmitted, combined, compared, and otherwise manipulated. For common reasons, these signals are sometimes referred to as bits, values, elements, symbols, characters, terms, numbers, etc., which has proven convenient.

[0091] However, it should be remembered that all these and similar terms are associated with appropriate physical quantities and are merely convenient labels applicable to those quantities. Unless otherwise stated, it is evident from the following discussion that, throughout this specification, the use of terms such as “receive,” “display,” “move,” “adjust,” “replace,” “confirm,” “play,” etc., refers to the actions and processes of a computer system or similar electronic computing device that manipulate data represented as physical (e.g., electronic) quantities in the registers and memory of the computer system and convert them into physical quantities similarly represented in the computer system's memory or registers or other such information storage, transmission, or display devices.

[0092] For the sake of simplicity, these methods are depicted and described herein as a series of actions. However, the actions according to this disclosure may occur in various orders and / or simultaneously, and may occur together with other actions not presented and described herein. Furthermore, not all the actions shown are necessary to implement the methods according to the disclosed subject matter. Moreover, those skilled in the art will understand and recognize that these methods may alternatively be represented as a series of interrelated states by one or more state diagrams or events. Furthermore, it should be understood that the methods disclosed in this specification can be stored on an article of writing to facilitate the transfer and assignment of such methods to a computing device. The term "article of writing" as used herein is intended to include a computer program accessible from any computer-readable device or storage medium.

[0093] Some embodiments of this disclosure also relate to apparatus for performing the operations herein. This apparatus may be constructed for the intended purpose, or it may comprise a general-purpose computer selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a computer-readable storage medium, such as, but not limited to, any type of disk, including floppy disks, optical disks, CD-ROMs and magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic cards or optical cards, or any type of medium suitable for storing electronic instructions.

[0094] Throughout this specification, references to "one embodiment" or "implementation" mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. Therefore, the phrases "in one embodiment" or "in an embodiment" appearing throughout this specification do not necessarily refer to the same embodiment. Additionally, the term "or" is intended to indicate an inclusive "or" rather than an exclusive "or." Furthermore, the words "example" or "exemplary" are used herein to indicate that something is used as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects or designs. Rather, the words "example" or "exemplary" are used to present concepts in a concrete manner.

[0095] It should be understood that the above description is intended to be illustrative and not restrictive. Many other embodiments will be apparent to those skilled in the art upon reading and understanding the above description. Therefore, the scope of this disclosure should be determined by reference to the appended claims and the full scope of their equivalents.

[0096] In addition to the descriptions above, users can be given control over whether and when the systems, programs, or features described herein can collect user information (such as information about the user's social networks, social behaviors or activities, professions, user preferences, or the user's current location), and whether content or communications are sent to the user from the server. Furthermore, some data may be processed in one or more ways before storage or use, thus removing personally identifiable information. For example, a user's identity may be processed so that personally identifiable information cannot be determined for the user, or the user's geographic location may be generalized at the location where location information is obtained (e.g., to the city, zip code, or state level), making it impossible to determine the user's specific location. Therefore, users can control what information about themselves is collected, how that information is used, and what information is provided to them.

Claims

1. A method for generating presentation slides with extracted content, comprising: Provide a user interface (UI) for presenting to users, the UI displaying slide-generating UI elements, the slide-generating UI elements allowing users to request slide generation using at least a portion of the content of a data file as source material; The UI receives the user's selection of UI elements for generating the slideshow. Multiple logical breakpoints in the content are identified based on at least one of the format or size of multiple content items in the content. The set of slides to be included in the slide presentation is determined based on the identified multiple logical breakpoints. The plurality of content items based on the content are each of the slides in the set of slides that are identified as a layout template; A trained machine learning model is applied to the plurality of content items of the content to obtain the output of the trained machine learning model, the output of which indicates one or more refined content items and a format for generating a presentation visualization item, the presentation visualization item including the one or more refined content items, wherein the one or more refined content items include a subset of the plurality of content items of the content, and wherein the trained machine learning model is trained to determine the text fragments to be included in the refined content items and the format to be used for the presentation visualization item including the refined content items; and The set of slides is generated based on each identified layout template and presentation visualization item, the presentation visualization item having an indicated format.

2. The method according to claim 1, wherein, The UI displays the data file having multiple parts and allows the user to select a first part of the multiple parts of the data file, wherein the multiple content items of the content are associated with the first part of the data file.

3. The method according to claim 2, further comprising: Receive user instructions regarding the selection of a second portion of the data file; as well as The second plurality of content items in the second part of the data file are refined into one or more additional refined content items to generate a second presentation visualization, wherein each of the one or more additional refined content items comprises a subset of the second plurality of content items in the second part of the data file, and wherein at least one slide in the set of slides is further generated based on the second presentation visualization.

4. The method according to claim 1, wherein, The trained machine learning model is trained using (i) identifying the training input of the set of training files and (ii) identifying the corresponding target output of the overview of the training files, wherein the corresponding target output further identifies the format of the overview of the corresponding training file.

5. The method according to claim 1, wherein, The plurality of content items of the content include a first set of sentences, and the one or more refined content items include a second set of sentences, the second set of sentences including fewer sentences than the first set of sentences, and wherein the generated presentation visualization items include a list based on the second set of sentences.

6. The method according to claim 1, wherein, The plurality of content items of the content include a data table, and the one or more refined content items include a range of data from the data table, wherein the generated presentation visualization includes a data chart based on the range of data.

7. The method according to claim 1, wherein, The plurality of content items of the content include images, and wherein the demonstration visualization item includes the images.

8. The method according to claim 1, further comprising: Receive interaction with the set of slides; as well as Use the interaction to apply heuristic rules to subsequent slide generation.

9. The method according to claim 1, further comprising: Set the text of the parent header in the data file as the title of the corresponding layout template of the slide; as well as Set the one or more refined content items, including the text associated with the parent header, as the body of the layout template.

10. A system for generating presentation slides with extracted content, comprising: Memory; as well as A processing device coupled to the memory, the processing device being configured to perform operations including the following: Provide a user interface (UI) for presenting to users, the UI displaying slide-generating UI elements, the slide-generating UI elements allowing users to request slide generation using at least a portion of the content of a data file as source material; The UI receives the user's selection of UI elements for generating the slideshow. Multiple logical breakpoints in the content are identified based on at least one of the format or size of multiple content items in the content. The set of slides to be included in the slide presentation is determined based on the identified multiple logical breakpoints. The plurality of content items based on the content are each of the slides in the set of slides that are identified as a layout template; A trained machine learning model is applied to the plurality of content items of the content to obtain the output of the trained machine learning model, the output of which indicates one or more refined content items and a format for generating a presentation visualization item, the presentation visualization item including the one or more refined content items, wherein the one or more refined content items include a subset of the plurality of content items of the content, and wherein the trained machine learning model is trained to determine the text fragments to be included in the refined content items and the format to be used for the presentation visualization item including the refined content items; and The set of slides is generated based on each identified layout template and presentation visualization item, the presentation visualization item having an indicated format.

11. The system according to claim 10, wherein, The trained machine learning model is trained using (i) identifying the training input of the set of training files and (ii) identifying the corresponding target output of the overview of the training files, wherein the corresponding target output further identifies the format of the overview of the corresponding training file.

12. The system according to claim 10, wherein, The plurality of content items of the content include a first set of sentences, and the one or more refined content items include a second set of sentences, the second set of sentences including fewer sentences than the first set of sentences, and wherein the generated presentation visualization items include a list based on the second set of sentences.

13. The system according to claim 10, wherein, The plurality of content items of the content include a data table, and the one or more refined content items include a range of data from the data table, wherein the generated presentation visualization includes a data chart based on the range of data.

14. The system according to claim 10, wherein, The plurality of content items of the content include images, and wherein the demonstration visualization item includes the images.

15. The system according to claim 10, wherein, The operation also includes: Receive interaction with the set of slides; and Use the interaction to apply heuristic rules to subsequent slide generation.

16. The system according to claim 10, wherein, The operation also includes: Set the text of the parent header in the data file as the title of the corresponding layout template of the slide; and Set the one or more refined content items, including the text associated with the parent header, as the body of the layout template.

17. A non-transitory computer-readable medium storing instructions, which, when executed by a processing device, cause the processing device to perform operations including: Provide a user interface (UI) for presenting to users, the UI displaying slide-generating UI elements, the slide-generating UI elements allowing users to request slide generation using at least a portion of the content of a data file as source material; The UI receives the user's selection of UI elements for generating the slideshow. Multiple logical breakpoints in the content are identified based on at least one of the format or size of multiple content items in the content. The set of slides to be included in the slide presentation is determined based on the identified multiple logical breakpoints. The plurality of content items based on the content are each of the slides in the set of slides that are identified as a layout template; A trained machine learning model is applied to the plurality of content items of the content to obtain the output of the trained machine learning model, the output of which indicates one or more refined content items and a format for generating a presentation visualization item, the presentation visualization item including the one or more refined content items, wherein the one or more refined content items include a subset of the plurality of content items of the content, and wherein the trained machine learning model is trained to determine the text fragments to be included in the refined content items and the format to be used for the presentation visualization item including the refined content items; and The set of slides is generated based on each identified layout template and presentation visualization item, the presentation visualization item having an indicated format.

18. The non-transitory computer-readable medium according to claim 17, wherein, The trained machine learning model is trained using (i) identifying the training input of the set of training files and (ii) identifying the corresponding target output of the overview of the training files, wherein the corresponding target output further identifies the format of the overview of the corresponding training file.

19. The non-transitory computer-readable medium according to claim 17, wherein, The plurality of content items of the content include a first set of sentences, and the one or more refined content items include a second set of sentences, the second set of sentences including fewer sentences than the first set of sentences, and wherein the generated presentation visualization items include a list based on the second set of sentences.

20. The non-transitory computer-readable medium according to claim 17, wherein, The plurality of content items of the content include a data table, and the one or more refined content items include a range of data from the data table, wherein the generated presentation visualization includes a data chart based on the range of data.

Citation Information

Patent Citations

  • Filmstrip automatic generation method based on electronic spreadsheet

    CN102169483A

  • Automated system for organizing presentation slides

    CN105531699A