User interfaces and tools for facilitating interaction with video content
The system addresses the challenge of finding specific content in recorded videos by incorporating interactive tools for real-time transcription, translation, and annotation, allowing for efficient searching and navigation of video content.
Patent Information
- Application Number
- JP2023562722
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-19
- Filing Date
- 2022-05-19
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-05-19
AI Technical Summary
Conventional recorded videos lack an efficient way for users to find specific content without watching or scanning the entire video.
The system provides interactive user interfaces and presentation tools that facilitate recording, sharing, viewing, searching, and casting of video content, including features like real-time transcription, translation, and annotation, which allow for precise synchronization and overlay of metadata with video content.
Enables users to efficiently search and navigate video content by providing searchable transcriptions and annotations, reducing the need to watch the entire video and enhancing the learning experience.
Smart Images

Figure 0007692498000001 
Figure 0007692498000002 
Figure 0007692498000003
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application is a continuation of U.S. Patent Application No. 17 / 303,075, filed on May 19, 2021, and claims the benefit thereof. The disclosure of that application is hereby incorporated by reference in its entirety.
Background Art
[0002] Background When making a presentation, the presenter often has to repeat instructions and information to explain a certain concept to a group of users. Then each user usually takes notes on the concept so that they can review the notes later. If a recording is made from the presentation, the number of times the presenter has to repeat the concept can be reduced. However, in a conventional recorded video, there is no easy way to provide the user with a way to find specific content within the video without watching and / or scanning the entire video. That is, when a user searches for a concept in the video, they have to watch or scroll through the entire recording to identify the location of that concept.
Summary of the Invention
[0003] Summary The systems and methods described herein can provide a number of user interfaces (UIs) and / or presentation tools that facilitate interaction with video content. For example, the tools can facilitate the recording, sharing, viewing, searching, and casting of video content. The video content can be educational, for presentation, and / or otherwise, based on information and input provided by any number of presenters and consumed by any number of users. The systems and methods described herein can provide, execute, and / or control the UIs and presentation tools based on commands received from an application (e.g., a browser, a web app, a native application, etc.) and / or commands received from an operating system (O / S) of a computing device. In some embodiments, the UIs and presentation tools described herein can be provided as a hybrid combination of information from both the application and the O / S. For example, some of the tools, UIs, and associated educational content (e.g., video content, files, annotations, etc.) may be provided by different application triggers or O / S trigger sources.
[0004] The systems and methods described herein can present a presentation tool that includes at least an interactive toolbar with a number of selectable tools (e.g., screen cast, recording of a screen cast, presenter camera (e.g., front (i.e., selfie) camera), real-time transcription, real-time translation, laser pointer tool, annotation tool, magnifying glass tool). The toolbar can be configured so that a presenter can easily present, record, and cast a presentation with a single input. Additionally, the toolbar can provide options to switch between presentation, recording, and / or casting. For example, certain tools and / or screen content can be configured to be switched on / off during recording. In some embodiments, specific tools can be provided to viewers of a recording (either in real-time or after the recording) to switch the toolbar, screen content, and / or video stream associated with the video. For example, certain elements of a recording (e.g., presenter's front camera stream, transcription stream, translation stream, annotation stream, etc.) can be switched on or off during recording and / or during a user's review of the recording.
[0005] The systems and methods described herein are configured such that a presentation tool can trigger the sharing of content from one or more computer displays. The presentation tool can enable a presenter and / or user to annotate (i.e., create annotations) the shared content in an effective manner. Annotations can be stored so that they can be retrieved later and aligned with timestamps and video content, as they are placed precisely on the shared content. For example, annotations can be made to the content during a video recording and / or casting of the content. Annotations can be layered on top of the content (e.g., underlying application content) and stored in metadata, so that the annotations can be appropriately positioned to move with the content when deleted or when a window event is detected (i.e., when the window is scrolled, resized, or moved across the UI). For example, when the presenter switches to another document (or scrolls within a document) during recording, e.g., when the presenter switches documents throughout the recording, metadata is used to save the annotation layer to trigger appropriate annotations to be overlaid on the appropriate content. This can enable multiple sources to be used to depict concepts and allow the presenter to place markup annotations on the content in an overlay layer (i.e., not in word processing editing), and when the presenter or user requests that the layer be deleted or reapplied, the overlay layer can be deleted and reapplied.
[0006] The systems and methods described herein enable a presenter or user to switch between a number of documents, applications, or other recorded content (accessed during recording) while annotating such content, where the annotations can be stored such that the annotations can be provided as an overlay that is retrieved and properly positioned as if executed during the video recording. Screen content, content captured by the presenter camera, transcription content, translation content, and annotation content can be configured to be switched on / off during and after recording (i.e., during viewing by the presenter and viewing by the user).
[0007] In some embodiments, the presentation tool described herein includes an annotation tool configured to enable a presenter or user to use one or more markup tools during recording to indicate chapters within the content, key ideas within the content. The markup tools can include any number of input mechanisms, including text input, laser pointer (and / or cursor, controller input, etc.), pen input, highlight input, graphic input, and the like.
[0008] In some embodiments, the systems and methods described herein can generate and display real-time transcriptions and / or translations of audio and video content. The transcriptions and / or translations can be depicted on a screen alongside other educational content. In some embodiments, the transcriptions and / or translations can be curated for later viewing after they are generated. For example, the transcription can be formatted to be easy to view and can be formatted to receive annotations from a presenter or user, where the annotations can indicate specific concepts of the content as important concepts to learn.
[0009] The systems and methods described herein can include tools for performing, formatting, and displaying translations and / or transcriptions of video content. When viewing a video (during or after recording), a user can scroll (e.g., video scroll) through the content (e.g., a web page, a document, etc.), and in response, the transcript portion can automatically scroll in synchronization with the video scroll. This synchronization of the video and the text content enables the corresponding text to be used for searching, thus facilitating an effective and resource-efficient search of the content included in the video.
[0010] In some embodiments, annotations and transcripts can be used to automatically generate a recap (e.g., summary) video representing a portion of the recorded video content. The systems and methods described herein can be searchable (and / or indexed) to surface in a search provided by an application (e.g., a browser) and / or the O / S of a computing device accessing the recorded video content, such that the annotations and the transcribed speech can be configured.
[0011] In some embodiments, the presentation tool described herein can include a magnifying glass tool that enables zooming in or out modes based on a single input. The magnifying glass tool can be used without manually changing the size of the window or web page. Additionally, the magnifying glass tool can be used in combination with an annotation tool. The annotation can be automatically resized with the video content to match the annotated content when the user exits either the zoom in or zoom out mode. This resizing enables the annotation to be stored via metadata, which can be later retrieved and applied as an overlay to the content without the annotated or zoomed content being the wrong size when reviewing the video content after recording ends.
[0012] A system of one or more computers can be configured to perform particular operations or actions by installing in the system software, firmware, hardware, or a combination thereof that cause the system to perform such actions during operation. One or more computer programs can be configured to perform particular operations or actions by including instructions that, when executed by a data processing apparatus, cause the apparatus to perform the actions.
[0013] In a first comprehensive aspect, a method implemented by a computer is described, which includes the step of starting a recording to capture video content, where the video content includes a presenter video stream, a screen-cast video stream, and an annotation video stream, and the step of generating a metadata record representing timing information used to synchronize at least one portion of the video content with an input received in at least one of the presenter video stream, the screen-cast video stream, or the annotation video stream during the capture of the video content based on the video content.
[0014] Embodiments can include any or all of the following features. In some embodiments, in response to the end of the recording, the method can include the step of generating a representation of the video content based on the metadata record, where the representation includes portions of the video content annotated by a user associated with the presenter video stream. In some embodiments, the timing information corresponds to a plurality of timestamps associated with each received input and at least one position in a document associated with the video content, and synchronizing the input includes, for each input, matching at least one of the plurality of timestamps to at least one position in the document.
[0015] In some embodiments, the video content further includes a transcription video stream, which is generated as modifiable transcription data configured to be displayed together with the screen-cast video stream during the recording of the video content, and includes real-time speech-to-text audio data from the presenter video stream. In some embodiments, the transcription video stream also includes real-time translated audio data from the presenter video stream, which is generated as text data configured to be displayed together with the screen-cast video stream and the speech-to-text audio data during the recording of the video content. In some embodiments, the transcription of the real-time speech-to-text audio data is performed by at least one speech-to-text application, the at least one speech-to-text application is selected from a plurality of speech-to-text applications determined to be accessible by the transcription video stream, and the modifiable transcription data and text data are stored in metadata records according to time stamps and configured to be searchable.
[0016] In some embodiments, the input includes annotation input associated with an annotation video stream, and the annotation input includes video marker data and teleprompter data generated by a user associated with the presenter video stream. In some embodiments, the presenter video stream, the screen-cast video stream, and the annotation video stream are configured to be switched on and off during recording, and the switching on and off triggers the display or removal from display of each presenter video stream, each screen-cast video stream, or each annotation video stream.
[0017] In a second general aspect, a system is described that includes a memory and at least one processor coupled to the memory, where the at least one processor is configured to generate a collaborative online user interface, the user interface including a renderer configured to render audio and video content associated with access to a plurality of applications from within the user interface, an annotation generation tool configured to receive annotation input in the user interface and generate a plurality of annotation data records for the received annotation input during rendering of the audio and video content, the annotation generation tool including at least one control for receiving the annotation input, a transcription generation tool configured to transcribe audio content during rendering of the audio and video content and display the transcribed audio content in the user interface, and is configured to receive commands from a content generation tool configured to generate a representation of the audio and video content in response to detecting the end of rendering. The representation can be based on the annotation input, the video content, and the transcribed audio content, and the representation includes portions of the rendered audio and video marked with the annotation input.
[0018] Embodiments can include any or all of the following features. In some embodiments, the content generation tool generates URL links to the representations of audio and video content and is further configured to index the representations to enable a search function to find at least a portion of the audio and video content in a web browser application. In some embodiments, a plurality of annotation data records include instructions for at least one application that receives annotation input in a plurality of applications and machine-readable instructions that overlay the annotation input on at least one image frame of a portion of the rendered video content depicting the at least one application according to respective timestamps.
[0019] In some embodiments, overlaying annotation input over at least one image frame includes retrieving at least one of a plurality of annotation data records, executing machine-readable instructions, and generating a document that enables a user to scroll at least one image frame with the annotation input overlaid over the at least one image frame according to at least one annotation data record. In some embodiments, the annotation generation tool is further configured to initiate a recording of rendered audio and video content, the rendered video content including data associated with a first application in a plurality of applications and data associated with a second application in the plurality of applications, receive a first set of annotations during a first segment of the recorded video content in the first application, store the first set of annotations according to respective timestamps associated with the first segment, receive a second set of annotations during a second segment of the recorded video content in the second application, and store the second set of annotations according to respective timestamps associated with the second segment.
[0020] In response to detecting that the cursor focus has switched from the first application to the second application, the annotation generation tool is further configured to retrieve the second set of annotations and data associated with the second application, align the timestamps associated with the second segment with the second set of annotations, and cause the retrieved second set of annotations to be displayed on the second application according to the respective timestamps associated with the second segment.
[0021] In some embodiments, the first set of annotations and the second set of annotations are generated by an annotation tool that enables marking, storing, and scrolling of the first set of annotations and the second set of annotations while maintaining an initial position on data associated with the first application or data associated with the second application for each annotation in the first set of annotations and the second set of annotations. In some embodiments, the annotation generation tool is further configured to retrieve the first set of annotations and the data associated with the first application in response to detecting that the cursor focus has switched from the second application to the first application, match the time stamps associated with the first segment to the first set of annotations, and cause a display of the retrieved first set of annotations on the first application according to each time stamp associated with the first segment.
[0022] In some embodiments, the annotation generation tool is further configured to receive additional annotations in the second application, where each additional annotation is associated with a respective time stamp, and generate a document from the second set of annotations and the additional annotations in response to detecting completion of the recording, where the document includes the second set of annotations and the additional annotations overlaid on data associated with the second application according to the respective time stamps associated with the second segment and the respective time stamps associated with the additional annotations, as well as a transcription of the recorded audio content associated with the second segment.
[0023] In a third general aspect, when executed by at least one processor, it causes a recording that captures video content to be started, where the video content includes a presenter video stream, a screen cast video stream, a transcription video stream, and an annotation video stream, and based on the video content, during the capture of the video content, at least one part of the video content is synchronized with an input received in at least one of the presenter video stream, the screen cast video stream, the transcription video stream, or the annotation video stream. A non-transitory computer-readable storage medium storing instructions configured to cause a computing system to execute instructions including generating a metadata record representing timing information used for synchronization.
[0024] Embodiments can include any or all of the following features. In some embodiments, the instructions further include, in response to the end of the recording, generating a summary video of the video content based on the metadata record, where the summary video includes parts of the video content annotated by a user associated with the presenter video stream.
[0025] In some embodiments, the timing information corresponds to a plurality of timestamps associated with each received input and at least one position in a document associated with the video content, and synchronizing the inputs includes, for each input, matching at least one of the plurality of timestamps to at least one position in the document.
[0026] In some embodiments, the transcription video stream is generated from the presenter video stream as real-time speech-to-text audio data, which is generated as text data configured to be displayed together with the screen-cast video stream during the recording of the video content, and real-time translated audio data from the presenter video stream, which is generated as text data configured to be displayed together with the screen-cast video stream and the speech-to-text audio data during the recording of the video content. In some embodiments, the real-time speech-to-text audio data is generated as editable transcription data configured to be displayed together with the screen-cast video stream during the recording of the video content, the transcription of the real-time speech-to-text audio data is performed by at least one speech-to-text application, the at least one speech-to-text application is selected from a plurality of speech-to-text applications determined to be accessible by the transcription video stream, and the editable transcription data and the text data are stored in a metadata record according to a time stamp and configured to be searchable.
[0027] In some embodiments, the input includes annotation input associated with an annotation video stream, and the annotation input includes video marker data and teleprompter data generated by a user associated with the presenter video stream. In some embodiments, the presenter video stream, the screen cast video stream, the transcription video stream, and the annotation video stream are configured to be switched on and off during recording, and the switching on and off triggers the display or removal from display of each presenter video stream, each screen cast video stream, each transcription video stream, or each annotation video stream.
[0028] In a fourth general aspect, when executed by at least one processor, to start a recording that captures audio content and video content, where the video content includes at least a presenter video stream, a screen cast video stream, a transcription video stream, and an annotation video stream; to cause the rendering of audio content and video content associated with the access of a plurality of applications from within a user interface; to receive an annotation input in the user interface during the rendering of the audio content and video content, where the annotation input is recorded in the annotation video stream; to perform speech-to-text on the audio content during the rendering of the audio content and video content, where the speech-to-text audio content is recorded in the transcription video stream; to translate the speech-to-text audio content during the rendering of the audio content and video content; and to cause the rendering in the user interface of the speech-to-text audio content and the translation of the speech-to-text audio content, along with the rendered audio content and the rendered video content. A non-transitory computer-readable storage medium storing instructions configured to cause a computing system to execute the instructions including these operations.
[0029] Embodiments can include any or all of the following features. In some embodiments, the computer-executable instructions are further configured to cause an online presentation system to generate representative content of at least a portion of the audio content and video content in response to detecting the end of the rendering of the video content and audio content. The representative content can be based on annotation input, video content, and transcribed and translated audio content, and the representative content includes portions of the rendered audio and video marked with annotation input. In some embodiments, the annotation input is rendered as an overlay on the video content, and the annotation input is configured to move with the video content in response to detecting a window event or cursor event that triggers a switch to other video content accessed during recording.
[0030] In a fifth comprehensive aspect, there is described a method implemented by a computer, the method including receiving at least one video stream, and receiving metadata representing timing information associated with an input detected in the at least one video stream, the timing information being configured to synchronize the detected input provided in the at least one video stream with a portion of the at least one video stream. In response to receiving a request to view at least one video stream, the method implemented by the computer can include generating a portion of the at least one video stream, the generating being based on the metadata and a detected user instruction requesting to view a representation of the at least one video stream, and causing rendering of the portion of the at least one video stream.
[0031] Embodiments can include any or all of the following features. In some embodiments, the timing information corresponds to a plurality of timestamps associated with each detected input in at least one video stream and at least one position in the content associated with at least one video stream, and synchronizing the detected inputs includes, for each input, aligning at least one timestamp with at least one position in a document associated with at least one video stream. In some embodiments, the at least one video stream includes a presenter video stream, a screen cast video stream, a transcription video stream, and an annotation video stream. In some embodiments, the representation of the at least one video stream includes a rendered portion of the at least one video stream annotated with the input, based on the detected input.
[0032] The systems, methods, computer-readable storage media, and aspects described above can be configured to perform any combination of the aspects described above, and each of them can be implemented in combination with any suitable combination of the features and aspects listed above.
[0033] Embodiments of the techniques described above can include hardware, a method or process, or computer software on a computer-accessible medium. Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features will be apparent from the description and drawings, and from the claims.
Brief Description of the Drawings
[0034]
Figure 1
Figure 2A
Figure 2B
Figure 3A
Figure 3B
Figure 3C
Figure 4
Figure 5A
Figure 5B
Figure 5C
Figure 6A
Figure 6B
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13A
Figure 13B
Figure 13C
Figure 13D
Figure 13E
Figure 13F
Figure 13G
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Best Mode for Carrying Out the Invention
[0035] The use of like or identical reference numerals in the various drawings is intended to indicate the presence of like or identical elements or features.
[0036] Detailed Description This specification describes a user interface (UI) and / or presentation tool that facilitates recording, sharing, viewing, interacting with, searching for, and casting video content. The UI and presentation tool can be provided in a presentation system that is online and capable of presenting content in real time. The presentation tool can be used to interact with presented (e.g., shared, cast, etc.) content. The systems and methods described herein can provide, execute, and / or control the UI and presentation tool based on commands received from an application (e.g., browser, web app, application, extension, native application, etc.) and / or commands received from an operating system (O / S) of a computing device. Accordingly, the systems and methods described herein can provide an online real-time presentation system as an application or as a set of user interfaces provided by the O / S.
[0037] In some embodiments, the systems and methods described herein can be used to generate educational content to be presented with a presentation tool. The content can all be transcribed, translated, and annotated in real time to distinguish important educational content. Annotations can be used to generate additional relevant content (e.g., educational content, learning guides, representative (e.g., recap, summary, snippet) videos and related content, video snippets, screenshots, image frames, etc.). For example, the application can automatically generate a recap video based on annotations provided to the content during the recording of a video (e.g., one or more presentations, lectures, seminars, etc.). Annotations can be provided by the presenter and / or the user. In operation, the presenter and / or the user can provide input for generating annotation markings in the form of text, importance indicators marked by the presenter (or marked by the user), and / or transcribed audio content markers, and the input is generated as a marker or overlay on the content being recorded in the video.
[0038] Conventional online educational videos cannot provide a convenient way for users to find specific content within a particular video without having to watch and / or scan the entire video. When a video is recorded, conventional techniques can generate a transcription that can be searched later, but cannot provide a real-time parallel view of the portion of the video related to the transcribed content. There is a need for a technical solution to provide live transcription and / or translation while recording a video. The systems and methods described herein provide such a technical solution that enables a parallel visual display of transcribed and / or translated content (e.g., a translation of the transcribed audio content) next to real-time annotated video content and / or screen share / screen cast content. This can provide the advantage of enhancing the learning and understanding of the video's content. The technical solution provided by the systems and methods described herein can enable video content (educational content, annotations, elements shown by a presenter, transcriptions, translations, etc.) to be quickly indexed and made searchable by a user. For example, the systems and methods described herein can provide a native application (or web application) configured to record the presented content and generate a presentation (e.g., a screen cast) function with features and tools for interacting with such content.
[0039] The techniques described herein provide a technical effect of enabling a single input command that simultaneously triggers the start of a screen cast (or screen share) presentation, the recording of the screen cast, and the transcription / translation of the content being screen cast. Multiple layers of the recorded content (e.g., document, website, nested video content layer, picture-in-picture layer, annotation layer, presenter camera (e.g., selfie) layer, participant (e.g., user) layer, transcription layer, and translation layer) can be captured separately to enable the presenter (i.e., the recorder) or the user (i.e., the participant or viewer) to switch the layers on and off. This can provide a more flexible approach to recording, with higher computational efficiency for recording all at once rather than having to record different layers separately or post-process the video, for example, to obtain a transcription. Additionally, the content of the recorded screen cast can be indexed so that a search task can retrieve and present the content while interacting with the recorded screen cast or while the recorded screen cast video is determined to have been recently accessed. This enables the integration of video content (not just a file name) into the OS-level search function in an efficient manner to avoid very long post-processing of the video. Such OS-level techniques can utilize signals received from a running application to adjust annotations (e.g., window events, etc.), and implementing the techniques described herein at the OS level can be more general-purpose than application-specific approaches to annotations.
[0040] The systems and methods of this specification can solve the technical problem (e.g., problem) of finding recent educational video content for a particular user. This can be useful when traditional classroom / lecture-based learning is replaced by home or "virtual" learning. For example, a user may not know where or how to retrieve previously captured video content when studying for a test or doing homework related to the educational content taught in the video content. In many cases, the user may have to study for a test using a large number of previously recorded videos. In conventional systems, the user may need to review, scan, and / or watch the entire of each video. However, the user can benefit from the key ideas and concepts from each video. Thus, the systems and methods described herein provide a technical solution in which representative videos annotated during the recording of one or more original videos are automatically generated to indicate key ideas and concepts. For example, the present systems and methods can enable the generation of one or more (e.g., a set of) curated searchable video content (e.g., summaries, snippets) considered important by a presenter or user (e.g., a presentation attendee). The generation of these representative videos is facilitated by the stream-based approach of capturing the content described herein.
[0041] The systems and methods described herein provide a technical solution to a technical problem by generating a repository of user interfaces (UIs) that can be used to present content (e.g., metadata, video content, etc.) and video snippets using the underlying O / S. The technical solutions described herein can provide technical effects such as improved content management, improved content access, and improved UI interaction. For example, the systems and methods described herein can generate representative videos that provide interactive explanations of parts of video content, presenter comments, annotations, and the like. Further, these snippets may be searchable using conventional file search or web browser applications.
[0042] Figure 1 is a block diagram illustrating an example of a real-time presentation system 100 according to an embodiment described herein. The system 100 can be provided by one or more applications 102 or an operating system O / S 104. In some embodiments, the system 100 can access and / or receive content from an online service, an online drive, an online library, and the like. The content can be depicted on one or more user interfaces (UIs) 106.
[0043] The real-time presentation system 100 can provide a control to the user so that the user can make a choice both when the system, operating system, application (e.g., program), and / or other functions described herein can enable the collection of user information (e.g., the user's social network, social actions or activities, occupation, user preferences, and / or information regarding the user's current location), and when content or communication is sent from the server to the user. Additionally, the system 100 can ensure that certain data is processed in one or more ways before being stored or used so that information that can identify an individual is removed. For example, the user's personal information can be processed so that it is not possible to determine information that can identify the user, or the geographical location of the user where location information is obtained (such as at the city, zip code, or state level) can be generalized so that the user's specific location cannot be determined. In this way, the user can control what information is collected about the user, how that information is used, and what information is provided to the user.
[0044] System 100 can generate any number of UIs (e.g., UI107) that can perform screen casting, screen sharing, and / or recording and upload to online resources either in real time or after recording. UI106 includes a toolbar 108, video and audio streams 110, representative content 112, annotations 114, and a library 116 and can be presented or otherwise accessed. For example, system 100 can be an online real-time presentation system (e.g., an application, UI, O / S-based portal) where a user can use toolbar 108, annotations 114, and library 116 to present content. A user can also use system 100 to generate video and audio content 110 that depicts annotations 114 provided by the user and / or presenter. The presentation content can be changed to provide specific representative content 112 that can include a recording, screen cast, share, and part of the presentation content. In some embodiments, representative content 112 is summary content (e.g., audio and / or video content with or without annotations) that summarizes all or part of a specific video content. In some embodiments, representative content 112 includes part of the video and / or audio content associated with a specific topic or category. In some embodiments, representative content 112 includes video and / or audio content that includes chapter information or title information of a specific video. In some embodiments, representative content 112 includes parts of a video that include markup (e.g., annotations), and such parts can include associated audio and / or metadata.
[0045] Generally, toolbar 108 can include an interactive toolbar that includes a number of selectable tools (e.g., screen cast, recording of screen cast, presenter camera (e.g., front camera (i.e., selfie) camera), real-time transcription, real-time translation, laser pointer tool, annotation tool, magnifying glass tool, etc.). The toolbar can be configured so that a presenter can easily present, record, and cast with a single input. Additionally, the toolbar may provide options to switch between presentation, recording, and / or casting. An example of a toolbar is shown as toolbar 11 in FIG. 1 7 as shown. Toolbar 11 7 includes a recording tool, a laser pointer tool, a pen tool (for generating annotation 114), an eraser tool, a magnifying glass tool, a selfie camera or other capture tool, and live transcription and translation tools, etc.
[0046] In some embodiments, toolbar 108 can include an annotation generation tool 108a configured to receive annotation input (e.g., annotation 120) in UI 107. (e.g., toolbar 11 7(selected from) Annotation generation tool 108a can generate an annotation data record (e.g., record 214) for the received annotation input 120 during the rendering of audio and video content (and as shown in UI 107). In some embodiments, annotation generation tool 108a can include at least one control (e.g., a software or hardware-based input control) that receives annotation input 120 and triggers the storage of a timestamp for the received annotation input. For example, system 100 can receive annotation 114 (e.g., annotation 120), and in response, store metadata (e.g., annotation data record 214) that includes one or more timestamps indicating when input 120 was received and in which application input 120 was received. Later, the metadata can be used to generate video snippets and / or representative content 112 based on when the input was received, what the input indicated, and / or the importance level of the input and / or the content related to the input. In some embodiments, since the user can select any number of tools to generate annotations for the content, for example, any number of tools on toolbar 11 7 Any number of the tools above may be part of annotation generation tool 108a.
[0047] In some embodiments, presentation system 100 can also generate and modify video streams and audio streams 110. For example, system 100 can be used to present content using various libraries 116 and accessed applications, images, or other resources. The content can be on toolbar 11 7It can be recorded using. The recorded content can be accessed by the presenter or another user. Using the recorded content, the system 100 can automatically generate representative content 112.
[0048] In some embodiments, the computing device hosting the system 100 can include a front camera tool (e.g., a selfie camera). Using the selfie camera, a presenter video stream can be generated as shown in presenter video stream example 122. The consumer of the content depicted on the UI 107 on the system 100, or the presenter (shown in stream 122), can switch the view of stream 122 on or off. For example, if stream 122 overlaps with content 124, for example, the presenter or the consumer of the content depicted on the UI 107 can make stream 122 not be displayed in order to secure a greater view of content 124. Similarly, the participant video stream 126 may be depicted on the UI 107. The participant video stream 126 can also be switched on or off by any of the participants or by the presenter.
[0049] During operation, the presenter (e.g., the user shown in stream 122) can access the system 100 so that, for example, the UI 107 and the toolbar 11 7 are presented. The presenter can present the content, annotate the content, record the content and / or the annotation, and upload the content and / or the annotation for future review using the toolbar 11 7Using, any or all of the content within UI107 can be cast, screen-cast, or otherwise shared. In this example, the presenter accesses system 100 via a browser application and has selected to share (e.g., cast) the entire browser application including presentation 101, tab 128, stream 122, stream 126, and previously entered annotations 120. Toolbar 11 7 is also presented in the shared content and can be toggled on / off.
[0050] Figures 2A and 2B are block diagrams showing an example computing system 200 configured to generate and operate a real-time online presentation system 100 according to the embodiments described herein. System 100 can operate on any of the computing systems described herein in a desktop operating system, a mobile operating system, an application extension, or other software. Using system 200, a computing device (e.g., computing system 201, computing system 202, and server computing system 204), and / or other devices (not shown in FIG. 2A) for operating system 100 (and corresponding UI) can be configured. For example, system 200 can generate a number of UIs that allow a presenter to share, annotate, and record audio and video using system 100.
[0051] As shown in FIG. 2A, computing system 202 includes an operating system (O / S) 216. Generally, O / S 216 can function to execute and / or control applications, UI interactions, accessed services, and / or device communications not shown. For example, O / S 216 is an application 21 7and the UI generator 220 can be executed and / or otherwise managed. In some embodiments, the O / S 216 can also execute and / or otherwise manage the real-time presentation system 100. In some embodiments, one or more applications 21 7 may execute and / or otherwise manage the real-time presentation system 100. In some embodiments, the browser 222 may execute and / or otherwise manage the real-time presentation system 100.
[0052] Application 21 7 can be any type of computer program that can be executed / distributed by the computing system 202 (or by the server computing system 204 or via an external service). Application 21 7 can provide a user interface (e.g., an application window, menu, video stream, toolbar, etc.) so that the user can interact with the functions of each application 21 7 . The application window of a particular application 21 7 can display application data along with any type of control such as a menu, icon, toolbar, widget, etc. Application 21 7 can include or access app information 224 and session data 226, both of which can generate content and / or data and can be used to provide such content and / or data to the user and / or the O / S 216 via the device interface. App information 224 is for a particular application 21 7It can correspond to information being executed by or otherwise accessed. For example, the app information 224 can include text, images, video content, metadata (e.g., metadata 228), inputs, outputs, or control signals associated with the interaction with the application 21 7 and can include control signals associated with the interaction with the application 21. In some embodiments, the app information 224 can include data downloaded from a cloud server, server 204, service, or other storage resources. In some embodiments, the app information 224 can include data associated with a particular application 21, including, but not limited to, metadata, tags, timestamp data, URL data, etc. 7 In some embodiments, the application 21 7 can include a browser 222. Using the browser 222, the system 100 can configure content for presentation, casting, and / or other sharing.
[0053] The session data 226 is for the application 21 7It can be related to the user session 230 with. For example, the user can access the user account 232 via a user profile 234 on or related to the computing system 202, or alternatively via the server computing system 204. Accessing the user account 232 can include providing a username / password or other type of authentication credentials and / or permission data 236. A login screen can be displayed so that the user can supply the user credentials, and upon authentication, the user can access the functions of the computing system 202. The session can start in response to determining that the user account 232 has been accessed or when one or more user interfaces (UIs) of the computing system 202 are displayed. In some embodiments, the session and the user account can be authenticated and accessed using the computing system 202 without communicating with the server computing system 204.
[0054] In some embodiments, the user profile 234 can include multiple profiles for a single user. For example, the user can have a work user profile and a personal user profile. Both profiles can utilize the real-time presentation system 100 to use and access content items stored from both user profiles. Thus, if the user opens a browser session with a business profile and opens an online file or application with a personal user profile, the system 100 can access the content on both profiles.
[0055] During a session (and if the user permits), session data 226 is generated. Session data 226 includes information about session items used / enabled by the user during a particular computing session 230. Session items can include clipboard content, browser tabs / windows, documents, online documents, applications (e.g., web applications, native applications), virtual desktops, display states (or modes) (e.g., split screen, picture in picture, full screen mode, selfie mode, etc.), and / or other graphical control elements (e.g., files, windows, control screens, etc.).
[0056] When the user launches, enables, and / or operates these session items on the user interface, session data 226 is generated. Session data 226 can include an identification of which session items (e.g., documents, browser tabs, etc.) were launched, configured, or enabled. Session data 226 can also include metadata that defines the position of a window, the size of a window, whether a session item is placed in the foreground or the background, whether a session item is focused or not, the time a session item was used (or last used), and / or the recency or last appearance order of a session item, and / or any or all of these details of the session. In some examples, session data 226 may include recorded content for the session, such as audio stream recording 110a and video stream recording 110b. Such recordings can be stored on a server (such as server 204 or a cloud server), locally (e.g., on device 201 or 202), or in a particular library 116 configured to store the recorded content and metadata of system 100.
[0057] In some examples, the session data 226 is sent to the server computing system 204 via the network 240, where it can be stored in the memory 242 in association with the user account 232 according to the user permission data 236 of the user in the server computing system 204. For example, when a user launches and / or operates a session item at a user interface (e.g., of the system 100) on the computing system 202, the session data 226 regarding the session item can be sent to the server computing system 204. In some embodiments, the session data 226 is instead (or also) stored within the memory device 244 on the computing system 202.
[0058] The UI generator 220 can generate representations of content items and toolbars that are rendered in a UI associated with and / or provided by the system 100. The UI generator 220 can perform searches, content item analysis, browser process initiation, and other processing activities to ensure that content items are rendered accurately and efficiently within a particular region or order in the UI associated with the system 100. For example, the generator 220 can determine how a particular content item is depicted in the UI associated with the system 100. In some embodiments, the generator 220 may add formatting to the content items depicted by the system 100. In some embodiments, the generator 220 may remove formatting from the content items depicted by the system 100.
[0059] As shown in FIG. 2A, O / S 216 includes or can access a service (not shown), a communication module 248, a camera 250, a memory 244, and a CPU / GPU 252. Computing system 202 can also include or access metadata 228 and preferences 256. Additionally, computing system 202 can also include or access an input device 258 and / or an output device 260.
[0060] Services (not shown) that system 200 can access can include online storage, content item access, account session or profile access, permission data access, etc. In some embodiments, the service may function to replace the server computing system 204 to which user information and account 232 are accessed via the service. Similarly, the real-time presentation system 100 may access via one or more services.
[0061] Camera 250 can include one or more image sensors (not shown) that can detect changes in background data associated with camera capture (and video capture) performed by computing system 202 (or another device communicating with computing system 202). Camera 250 can include a rear capture mode and a front capture mode.
[0062] Computing system 202 can generate and / or distribute specific policies, permissions, and preferences 256. The policies, permissions, and preferences 256 can be configured by the computing system 202, the device manufacturer of system 100, and / or the user accessing system 202. The policies and preferences 256 can include routines (i.e., a set of actions) that are triggered based on audio commands, visual commands, schedule-based commands, or other configurable commands. For example, a user can set a specific UI to be displayed and start recording an interaction with the UI in response to a specific action. In response to detecting such an action, system 202 can display the UI and trigger the recording. Other policies and preferences 256 may be configured to modify and / or control the content associated with system 202 that is composed of policies, permissions, and / or preferences 256.
[0063] Input device 258 can provide data received via, for example, a touch input device that can receive tactile user input, a keyboard, a mouse, a hand controller, a wearable controller, a mobile device (or other portable electronic device), a microphone that can receive audible user input, etc., to system 202. Output device 260 can include, for example, a device that generates content for a display for visual output, one or more speakers for audio output, etc.
[0064] In some embodiments, computing system 202 can store specific applications and / or O / S data in a repository. For example, annotations 114, data records 214, metadata 228, audio stream recordings 110a, and video stream recordings 110b can be stored for later retrieval and / or extraction. Similarly, screen captures and annotation video streams can also be stored in and retrieved from such repositories.
[0065] Server computing system 204 can include any number of computing devices in a number of different device forms, such as standard servers, groups of such servers, or rack server systems. In some examples, server computing system 204 can be a single system that shares components such as processor 262 and memory 242. User account 232 can be associated with the configuration of system 204 and / or session 230 and / or the configuration of profile 234 according to user permission data 236, and can be provided to system 202, for example, in response to requests from the user of user account 232.
[0066] Network 240 can include other types of data networks, such as the Internet and / or local area network (LAN), wide area network (WAN), cellular network, satellite network, or other types of data networks. Network 240 can also include any number of computing devices (e.g., computers, servers, routers, network switches, etc.) configured to receive and / or transmit data within network 240. Network 240 can further include any number of wired connections and / or wireless connections.
[0067] The server computing system 204 can include one or more processors 262 formed on a substrate, an operating system (not shown), and one or more memory devices 242. The memory device 242 can represent any type of (or types of) memory (e.g., RAM, flash, cache, disk, tape, etc.). In some examples (not shown), the memory device 242 can include an external storage device, for example, a memory that is physically separate from the server computing system 204 but accessible by the server computing system 204. The server computing system 204 can include one or more modules or engines representing specially programmed software.
[0068] Generally, the computing systems 100, 201, 202, and 204 can communicate with each other via the communication module 248 and / or transfer data wirelessly via the network 240, for example, using the systems and techniques described herein. In some embodiments, each of the systems 100, 201, 202, and 204 can be configured to communicate within the system 200 with other devices associated with the system 200.
[0069] FIG. 2B represents an architecture example 263 that records video and audio and stores the resulting recorded content (e.g., audio stream recording 110a, video stream recording 110b, recorded annotation 114, and other recorded video streams) together with associated metadata 228. In this example, the real-time presentation system 100 is accessed via a native application for the O / S and uses a recording tool associated with the native application. Recordings (e.g., video and audio streams) may be uploaded in real time to an online drive.
[0070] As shown in FIG. 2B, O / S 216 can include or access the real-time presentation system 100 and any number of applications 21 7 For example, the application 21 7 can also include a browser 222. The browser 222 represents a web browser configured to access information on the Internet. The browser 222 can launch one or more browser processes 264 to generate browser content or other browser-based operations. The browser 222 can also launch browser tabs 266 within the context of one or more browser windows 268.
[0071] The application 21 7 can include a web application 270. The web application 270 represents an application program stored, for example, on a remote server (such as a web server) and distributed over the network 240 via the browser tab 266. In some embodiments, the web application 270 is a progressive web application that can be saved on the device and used offline. The application 21 7 can also include non-web applications that can be at least partially stored (e.g., locally stored) on the computing system 202. In some examples, the non-web applications may be executable by the O / S 216 (or executable on top of the O / S 216).
[0072] The application 21 7may further include a native application 272. The native application 272 represents a software program developed to be used on a specific platform or device. In some examples, the native application 272 is a software program developed for multiple platforms or devices. In some examples, the native application 272 is a software program developed to be used on a mobile platform and configured to also execute on a desktop or laptop computer.
[0073] In some embodiments, the real-time presentation system 100 can be executed as an application. In some embodiments, the system 100 can be executed within a video conferencing application. In some embodiments, the real-time presentation system 100 can be executed as a native application. Generally, the system 100 can be configured to support the selection, modification, and recording of audio data or text, HTML, images, objects, tables, or other content items within the application 21 7 In some embodiments, the real-time presentation system 100 can be configured to support the selection, modification, and recording of audio data or text, HTML, images, objects, tables, or other content items within the application 21.
[0074] The presentation system 100 shown in FIG. 2B includes a recording 273, a real-time transcription 274, a real-time translation 275, drawings 276, and key idea metadata 278. Each of the elements 273-278 can be recorded during a session of the system 100. The recorded elements 273-278 can represent a video and / or audio stream that can be annotated by a first user (e.g., a presenter) during the session and provided (shared, cast, streamed, etc.) in real time to any number of other users (data consumers, participants, etc.).
[0075] In some embodiments, the recorded streams associated with elements 273-278 can be generated using one or more tools associated with system 100. System 100 includes and / or can access a memory and at least one processor coupled to the memory, and the at least one processor is configured to generate a collaborative online user interface (e.g., system 100). The user interface is configured to receive commands from a renderer and tool / toolbar 108 (e.g., annotation generation tool 108a, transcription generation tool 108b, video content generation tool 108c). Each tool / toolbar 108 can be accessible via a UI or toolbar presented by system 100.
[0076] The renderer (e.g., UI generator 220) can be configured to render audio and video content associated with access to one or more of a plurality of applications within the user interface of system 100. For example, the renderer can utilize the UI generator 220 to render an application, annotation, cursor, input, video stream, or other UI content within system 100 or associated with computing system 202.
[0077] (e.g., toolbar 11 7The above annotation generation tool 108a can be configured to receive annotation input (e.g., annotation input 120) in the user interface. And the annotation generation tool 108a can use that input to generate any number of annotation data records for the received annotation input during the rendering of audio and video content. The annotation generation tool 108a can include at least one control that receives the annotation input and results in the storage of a timestamp for each received annotation input. The timestamp can be used to match the video content with annotations, transcriptions, translations, and / or other data associated with the system 100.
[0078] In some embodiments, annotation data record 211 (e.g., generated from annotation 114 and / or metadata 228) can include an indication of at least one application that is receiving an annotation input and is being accessed. The annotation data record 211 can also include machine-readable instructions to overlay the annotation input on at least one image frame of a portion of the rendered video content depicting the indicated application (according to respective timestamps). For example, the annotation data record 211 can utilize any number of video streams, metadata, and annotation inputs to determine, for example, which specific application is receiving the annotation and at what point in time, in order to determine the proper positioning of an overlay (e.g., a video stream overlay) for a particular frame of one or more other video streams depicting the application. Using these image frames and annotation overlays, representative content 112 can be generated to enable the user to quickly review the annotated concept, thereby avoiding the user having to review the entire video stream.
[0079] Overlaying the annotation input on at least one image frame can include retrieving at least one of a plurality of annotation data records and executing machine-readable instructions for performing the overlay. The system 100 can then generate a document (e.g., an online document, video snippet, transcription snippet, image, etc.) by which the user can scroll through at least one image frame including the annotation input overlaid on at least one image frame (based on annotation data records indicating timestamps, annotations, etc.).
[0080] The transcription generation tool 108b can be configured to transcribe audio content captured during the rendering of audio and video content and display the transcribed audio content on a user interface associated with the system 100. In some embodiments, the transcription generation tool 108b can also provide markers, highlights, or other indicators overlaid on the transcribed text to indicate to a user viewing the presentation a specific location in the transcription corresponding to audio speech rendered by the system 100 and spoken by a presenter. In some embodiments, additional indicators can be provided with or on top of the transcribed text to indicate important concepts or language. A user accessing the recording later can utilize such indicators to quickly find important concepts or language. Additionally, the system 100 can use such indicators as triggers to obtain audio content, video content, transcription content, translation content, and / or annotation content that occurs within a time threshold associated with the marking of a particular indicator. Such indicators can be used to generate summary content and / or other representations of the video stream (e.g., audio and video content).
[0081] For example, the summary generation tool 108c can be configured to extract such indicators (and / or annotations) to generate representative content 112 in response to detecting the end of the rendering of audio and / or video. The representative content can be based on annotation input, video content, and transcribed audio content. In some embodiments, the summary content can include portions of the rendered audio and video marked with annotation input (or other indicators). In some embodiments, the video content generation tool 108c is further configured to generate a URL link to the representative content 112. For example, the system 100 can trigger portions of the video and / or audio content of one or more video streams, particularly compiled, curated, or otherwise combined, to be uploaded to a website or online storage memory so that the portions can be conveniently and later accessed. In some embodiments, the tool 108c can also index the representative content 112 to enable a search function to find at least a portion of the representative content 112, for example, using a web browser application 222.
[0082] During operation, a first user (e.g., of the presenter computing system 279) can trigger a session of the real-time presentation system (e.g., via an application trigger or an O / S trigger). The system can be operated so that the presenter of system 279 presents and records content. For example, system 279 can trigger a recording 273 to generate video and / or audio content in the form of a recorded presenter video stream (e.g., content captured by a selfie camera), a screen cast video stream (e.g., the drawing 276 and screen cast 277 content), an annotation video stream (annotation data record 214 and / or key idea markers and corresponding metadata 278), a transcription video stream (e.g., real-time transcription 274), and / or a translation video stream (e.g., real-time translation 275). The presenter can turn any of these streams on / off during recording. In some embodiments, metadata 228 can be captured and stored during recording. Metadata 228 can be associated with any number of video streams. Each video stream can also include audio data and / or annotation data. However, in some embodiments, the annotation data may be recorded separately as a video layer.
[0083] When recording is triggered and presentation and / or annotation of content is initiated, system 100 can trigger the casting application 280 to cast the presentation and / or annotation on a separate device (e.g., the television 281 in the boardroom or other device). System 100 can also trigger the transcription of the video / audio content 282, which can be generated in real time and provided to the online storage 283. The content can be formatted in real time by the formatting application 284 for presentation within system 100, and the formatting application 284 can also provide such transcribed (and / or translated) data to the application 285 (or other applications accessible to a user using, for example, the computing system 286). In some embodiments, translation and transcription may not need to be requested by the user to be provided in the view of the UI of system 100. In that case, the presenter computing system 279 can directly provide the recording content in real time to the formatting application 284 and then to the user computing system 286 (in some examples via the application 285).
[0084] In some embodiments, system 100 can initiate a recording 273 that captures video content (and / or audio content). The video content (and / or audio content) can be represented as a presenter video stream, a screen cast video stream, a transcription video stream, a translation video stream, an audio stream, and / or an annotation video stream. Any suitable combination of these streams can form the video content, and if the presenter selects to turn one or more streams off or on during recording 273, the streams within the video content can change. By being able to select different streams in such a simple manner, a flexible approach is provided for recording content and generating additional representative content from the recorded content. System 100 can generate at least one metadata record during the capture of the video content (and / or audio content) based on the video content (and / or audio content). Each metadata record can represent timing information used to synchronize at least one portion of the video content with an input received in at least one of the recording video streams (e.g., annotation 114 / record 214, key idea metadata 278). In other words, the timing information can be used to synchronize an input received in at least one of the presenter video stream, the screen cast video stream, or the annotation video stream (or any other stream) with the video content. The timing information can be used later to generate a learning guide (representative content 112), an overlay of annotations on snippets of the video content, searchable video content, etc.
[0085] Figures 3A - 3C are screenshots showing the switching between an example of a user interface (UI) of a real - time presentation system and content with annotations according to the embodiments described herein. In this example, the presenter (shown in presenter video stream 122) can trigger a presentation (e.g., a screen cast, screen share, video conference, etc.) to start the presentation and recording of content for consumption by the users shown in participant stream 126. In some embodiments, system 100 is configured to trigger the start of recording of specific audio and video content rendered by system 100. For example, the presenter can indicate the start of sharing of content from system 100 with a single control, thereby triggering the automatic recording of such content.
[0086] As shown in FIG. 3A, the presenter in stream 122 is presenting a first application 302 and a second application 304. The first application 302 is annotated with annotations 306 and 308. The presenter in stream 122 can actively annotate using the cursor 310a, for example, by using a pen tool 312 from an annotation generation tool (e.g., toolbar 314). During operation, the rendered video content can include data (maps as well as annotations 306 and 308) associated with the first application 302 from any number of open or available applications accessible to system 100. The rendered video content can also include data (e.g., geographical concepts) associated with the second application 304. Motion The rendered video content can also include data (e.g., geographical concepts) associated with the second application 304.
[0087] The presenter (or consumer of the presented content) can annotate any number of applications, documents, content items, or display portions presented by the system 100, so the system 100 is configured to track which of the above items receive the annotation. By tracking the annotation to the annotated item, the annotation can be captured as a layer of video content (e.g., stream), and when the user accesses the recorded content later, the layer can be overlaid later or not displayed. Such overlay switching can ensure that the user can appropriately display the application content and the annotation for the appropriate application content. In addition, the user can use a scroll control (e.g., control 316) associated with an application (e.g., application 304). The presenter can scroll the content in a specific application having a cursor focus for scrolling the content and scroll (e.g., move) the annotation with the content. In this way, a set of overlaid annotations can be captured and scrolled with the application content to ensure that the annotated application content is saved.
[0088] As shown in FIG. 3B, the presenter (shown in presenter stream 122) is presenting application content in application 304. In this example, the presenter has annotated the content in application 304 using toolbar 314, as shown by annotations 318, 319, and 320. Annotations 318 - 320 are depicted as text writing with a selected pen tool, but any number of annotations and annotation types can be entered using markup tools and / or selections within the application content. For example, the content can be highlighted, drawn on, corrected, marked, etc. In some embodiments, certain content can include an indicator for marking the content. For example, some content may relate to a paragraph of text. In such an example, the entire paragraph can be marked by selecting an indicator presented on or near the paragraph within the application content. Each of annotations 318 - 320 can be associated with one or more timestamps representing the time within the recorded video at which the respective annotation was entered by the user. The timestamps can indicate how system 100 can track and search for specific content that includes the annotations.
[0089] For example, by tracking the annotations, system 100 can, in real time, in a first application, receive a first set of annotations (e.g., annotations 306 and 308) in a first segment of the recorded video content and store the first set of annotations (e.g., annotations 114 and / or annotation data record 214) according to the respective timestamps associated with the first segment. System 100 can also, in real time, in a second application (e.g., application 304), receive a second set of annotations (e.g., annotations 318, 320, and 322) in a second segment of the recorded video content and store the second set of annotations according to the respective timestamps associated with the second segment. At some point, system 100 can detect that the cursor focus has switched between applications. For example, system 100 may determine that the presenter has switched from using application 302 where cursor 310a is in focus to using application 304 where cursor 310b is in focus. Since the annotations may be provided as a layer on top of the application content, the annotations can be applied and removed in response to the change in cursor focus to avoid having annotation - attached content that is no longer applicable to the application or application content that most recently received the cursor focus.
[0090] In response to detecting that the cursor focus has switched from the first application 302 to the second application 304, the system 100 can retrieve the second set of annotations 318, 320, and 322 and retrieve data associated with the second application (e.g., application content, metadata, or other settings for the content). Next, the system 100 can match the timestamps associated with the second segment with the second set of annotations 318, 320, and 322. To appropriately display the annotations received at the previous timestamp, the system 100 matches the content being displayed at the time of the timestamp (such as a screen cast) and overlays the annotations (e.g., annotations 318, 320, and 322). The system 100 can then display the retrieved second set of annotations (e.g., annotations 318, 320, and 322) on the second application 304 according to each timestamp associated with the second segment. Additionally, the system 100 may delete the annotations applied to different applications associated with the system 100. For example, the system 100 may delete the annotations associated with the application 302 when the presenter switches the cursor focus to the application 304. As shown in FIG. 3A, when the user switches back to the application 302, the system 100 deletes the annotations 318, 320, and 322 and instead retrieves and renders the annotations 306 and 308 to ensure, for example, that the application 302 depicts the exact annotations from the previous markup. In an example where the applications 302, 304 are arranged in parallel within the UI (i.e., not overlapping), the annotations 306, 308 can be displayed in the application 302, and at the same time, the annotations 318, 320, 322 can be displayed in the application 304.In this way, the user can view all the annotations for the displayed content simultaneously.
[0091] In some embodiments, a presenter using the system 100 can trigger the generation of a first set of annotations (e.g., annotations 306 and 308) and a second set of annotations (e.g., annotations 318, 320, 322) via an annotation tool (e.g., from one or more tools of toolbar 314 or another toolbar). The annotation tool can enable the marking, storage, and scrolling of the first set of annotations (e.g., annotations 306 and 308) and the second set of annotations (e.g., annotations 318, 320, 322) while maintaining the initial positions in the data associated with the first application or the data associated with the second application for each annotation of the first set of annotations (e.g., annotations 306 and 308) and the second set of annotations (e.g., annotations 318, 320, 322). That is, the annotation tool can store metadata indicating the location (i.e., position) in the data content presented by a particular application where each annotation can be found for each annotation. In this way, the system 100 can generate an overlay of annotations that can be restored on top of the data content, for example, when summary content (or other representative content) is generated. In another example, the system 100 can generate such an overlay of annotations at the appropriate position in the data content when the presenter scrolls the data content and / or switches between applications.
[0092] In some embodiments, the second application 304 can receive additional annotations (e.g., annotation 324). In this example, the presenter added notes regarding library code, resource links, and changes to office hours. The additional annotations (e.g., annotation 324) can also be associated with respective timestamps corresponding to when the annotation 324 was added to the content of the application 304 during recording. In response to detecting the completion of the recording, the system 100 can generate a document 328, as shown in FIG. 3C. The document 328 can be generated from a set of second annotations (e.g., annotations 318, 320, and 322) and additional annotations (e.g., annotation 324). The document can include a set of second annotations 318 - 322 and additional annotation 324 overlaid on data associated with the second application 304 according to the respective timestamps associated with the second segment and the respective timestamps associated with the additional annotation. In some embodiments, one or more still image frames or video snippets 330 may be generated to be executed within the document 328, or provided as a link or search result associated with the document 328. The inputs (such as annotations 318 - 322 and additional annotation 324) can be synchronized with the video content (i.e., overlaid at the correct position of the data from the application 304) by matching the timestamps to respective positions within the document 328 associated with the video content.
[0093] In some embodiments, system 100 can also generate a transcription 332 of the recorded audio content associated with the second segment. Generally, document 328 can be configured to be changed at any time. For example, the presenter can later make changes to the recorded presentation, such as changed audio, additional markup or annotations, and / or other changes. Such changes can be configured to trigger the regeneration of document 328 to include the changes. Document 328 can also be referred to as a summary content document or a representative content document.
[0094] FIG. 4 is a screenshot showing an example presenter toolbar 400 provided by a real-time presentation system according to the embodiments described herein. Presenter toolbar 400 includes at least a laser pointer tool 402, a pen tool 404, a magnifying glass tool 406, an eraser tool 408, a screen cast recording tool 410, a chapter creation tool 412, a self-shot (e.g., presenter) camera tool 414, a closed caption tool 416, a transcription tool 418, and a marker tool 420. Each tool 402-420 of toolbar 400 may be part of annotation generation tool 108a. For example, each tool can be used to create an annotation for the content being presented.
[0095] Using the laser pointer tool 402, the cursor can be configured as a laser pointer during a presentation in the system 100. The laser pointer tool 402 can provide a visual focus to the consumers of the presentation provided by the system 100. The pen tool 404 can provide an annotation function for any content or part of the presented screen (e.g., window, application, full screen, etc.). The pen tool 404 can include any number of selectable pens, color contents, sizes of contents and / or texts, shapes, etc. The magnifying glass tool 406 can provide a zoom function for all small texts and graphics enlarged by the presenter during the presentation. The eraser tool 408 can provide a deletion and erasure function similar to a manual eraser to correct errors or delete annotations, for example, to secure a place for generating more annotations.
[0096] The screen cast recording tool 410 can provide a recording function to start a recording and upload the recorded content locally, to a cloud server, or to other selected locations. In some embodiments, the screen cast recording tool 410 not only triggers the recording but also triggers a screen cast, screen share, or other presentation modes. For example, when the presenter selects the tool 410, the presentation and the recording can start simultaneously. This allows the user to select a single control input to quickly start the presentation of the content while recording the content and / or related audio content, thus providing the advantage that the presentation and recording are easy for the user (e.g., the presenter).
[0097] Generally, the screen or window that is shared when tool 410 is selected can be the last detected sharing setting or the last screen used before selecting tool 410. That is, the presenter's recording scope can match the previously selected display scope (e.g., tab, window, full screen, etc.). In some embodiments, a confirmation UI can be presented when tool 410 is selected so that the presenter can select which display scope to share and / or record. In some embodiments, the presenter can stop the presentation by reselecting tool 410. However, this action may not stop the recording. This can be advantageous as it allows the presenter to add additional notes, audio, or additional content that the viewer may wish to have when accessing the recording at another time.
[0098] To end the recording, the presenter can select another tool or command (not shown). When the recording is ended (e.g., stopped) in system 100, toolbar 400 may be removed from view. Further, when an instruction to stop the recording is detected, system 100 can automatically trigger the upload, transmission, or other finalization of the recording. The recording is generally uploaded as it occurs rather than at the end of the recording, so the delay for upload completion can be minimal. In some embodiments, system 100 may be offline, and in such situations, a local copy of the recording can be generated instead.
[0099] The chapter creation tool 412 can be used by a presenter to annotate a recording video with respect to time. For example, the presenter can select the tool 412 at any point during the presentation to generate chapters for the recording video. In some embodiments, the chapter creation tool 412 (or post-recording tool) can be used to create chapters for the recording after the recording is completed (e.g., after recording). Thus, the presenter may wish to further annotate the presentation with chapters to facilitate the user's future search and review of content from the presentation. A chapter represents a section of the video. The chapter can provide a preview image frame to assist the user in identifying the content of the chapter. The chapter can also include metadata, title data, or identification data added by the user or the system. The video split by chapters can be presented in a timeline view to enable the user to select previously configured chapter indicators presented on the timeline. Conventional systems that provide chapter generation provide such functionality after recording. That is, conventional systems do not provide the option to generate chapters in real time (e.g., on the fly) while recording the video.
[0100] The self - capture (e.g., presenter) camera tool 414 can trigger the function of the front camera on a computing device (e.g., device 202) that executes the real - time presentation system 100. The tool 414 can be switched on and off by the presenter and / or user (e.g., consumer) of the presented content. The video stream captured by the tool 414 can be used by the closed captioning tool 416 and / or the transcription tool 418 to generate captions, transcriptions, and translations of the audio data presented from the video / audio stream (e.g., stream 122) captured by the tool 414 (e.g., via camera 250).
[0101] The transcription tool 418 represents the transcription generation tool 108b described herein. The presenter of the system 100 can switch the real - time transcription of audio on and off. In some embodiments, the transcription tool 418 can trigger a live transcription with full translation by using the closed captioning tool 416 in combination with the transcription generation tool 108b. The transcription tool 418 can work in cooperation with the UI generator 220 to generate a specially formatted transcription for rendering, for example, with the content presented via a screen share presentation from the system 100.
[0102] The marker tool 420 can be selected by a presenter to mark, for example, specific content, ideas, slides, annotations, or other presentation parts of the screen as key ideas. The key ideas can represent elements that the presenter considers useful, important, learning guide materials, and / or selectable as representative content 112. When the presenter selects the marker tool 420, other markings (such as highlighting, annotations, etc.) can be made to the presented content so that it can be stored as a key idea in the system 100. In some embodiments, the marker tool 420 can provide user feedback in the form of a backlight or other markings on the tool 420 to let the presenter understand that the tool 420 is active. Other feedback options are also possible.
[0103] The toolbar 400 can also include a close menu control (not shown) that can function to close or minimize the toolbar. The toolbar 400 can be moved and / or rotated for use in any presentation provided by the system 100. In some embodiments, the toolbar 400 can be made invisible when the cursor is dragged over the toolbar, for example, when a mouse over event occurs on the toolbar. This can provide the advantage that the presenter and the viewers of the presentation (such as users) can view the content without the need to manually move the toolbar 400.
[0104] Figures 5A - 5C show screenshots of an example of sharing a screen in a UI of a real - time presentation system according to the embodiments described herein. Figure 5A shows a browser 500 in which a user is accessing the home page of presentation 101 (e.g., P101). The user is also accessing the content within browser tab 502 and browser tab 504. The user can decide to present the content to one or more other users. For example, the user can be a presenter who plans to provide a presentation to a number of users.
[0105] The presenter can access a menu UI 506 provided by a computing system 202 (e.g., via O / S 216 or an application 21 that hosts the real - time presentation system 100). The UI 506 may be presented from a quick - settings UI. From UI 506, the presenter can select a presentation control 508 with a cursor 510 so that an additional screen for configuring a screen cast and / or screen share for presenting the content from presentation 101 is provided. 7
[0106] Figure 5B shows a presentation UI 512 where the presenter can select to cast 514 the content or share the content via a video conference 516. For example, the presenter can select to present presentation 101 via screen cast to a TV in an executive conference room (e.g., TV 281). Alternatively, the presenter can select to present presentation 101 via a video - conferencing application (e.g., using a native application or a browser application). In this example, the presenter has selected to cast presentation 101 as indicated by cursor 518.
[0107] FIG. 5C shows a casting UI 520 that allows a presenter to select which display focus to cast. Since the user has selected to share content, system 100 can populate toolbar 522 to indicate that presentation tools are available. UI 520 includes options for sharing the screen. The options include at least built-in display option 524 and external display option 526. In this example, the presenter has selected built-in display 524 as indicated by cursor 528. The presenter can also be provided with options regarding which scope of the screen to share. Example options depicted include full screen option 530, browser tab option 532, and application window 534. Other options are possible and are based on the content that is cursor focused behind UI 520. The presenter can be provided with an option 536 to share (or not share) audio content. The presenter can also be provided with an option 538 to render (or not render) presenter tools. The presenter can select an option and use save control 540 to save the selected option.
[0108] FIGS. 6A and 6B show screenshots of an example toolbar provided by real-time presentation system 100 according to embodiments described herein. FIG. 6A shows a shared presentation of browser tab 600 having a rendered toolbar 602. The presenter can access tools on toolbar 602 similar to toolbar 400. In this example, the presenter has selected pen tool 604. In response, system 100 provides a sub-panel 606 for pen tool 604 so that the presenter can select options for the pen. Sub-panel 606 also includes a trash can option 609 to delete the selected annotation.
[0109] As shown in FIG. 6A, the presenter provides annotation inputs such as drawing 610, text 612, and drawings (e.g., a circle at line 614). The presenter also draws additional markings 616 that appear to be errors or extra pen strokes. In this case, the user can select the marking 616 and then select option 609 to delete the marking 616.
[0110] Annotations from toolbar 602 can be generated for content within the scope of the shared window or screen. If the presenter starts drawing or annotating outside of that scope, system 100 can trigger an indication that the annotation is out of view. Additionally, the annotations can be made scrollable and configured to keep the content annotated during a recording / casting session. To align the content with the annotations to enable the recorded content and annotations to be accessed after recording / casting, an annotation video stream with corresponding metadata can be captured. In some embodiments, system 100 may be configured to capture the annotations within the annotation stream, but may not display the annotations during recording / casting if a scroll event is detected. In some embodiments, system 100 can enable each user to manually purge the annotations, for example, after recording.
[0111] In some embodiments, window switching can be triggered so that when switching from one window or application to another, the annotation is deleted (e.g., made invisible). Subsequently, when switching back to the window or application associated with the annotation, the annotation can be replaced (e.g., redisplayed). Additionally, the annotation can be resized according to the resized window. In some embodiments, the annotation may remain visible (i.e., rendered and displayed) as long as the underlying application content is visible to the user. In other words, the annotation can be visible even if the associated application is overlapped in another window or application or is otherwise not in the foreground.
[0112] FIG. 6B shows a toolbar example 602 having another sub-panel example 620. In this example, the toolbar 602 includes, among other examples, a trash can option 622 to delete a specific annotation, a redo / undo button to redo or undo annotation input, a static pen 626, an erasable pen 628, a fluorescent pen 630, and any number of selectable colors 632, 634, and 636. Additional sub-panels may be provided for display to allow a presenter to select other options related to, for example, color, font, line style, or pen tool 604.
[0113] FIG. 7 shows a screenshot of an example use of the toolbar 108 provided by the real-time presentation system 100 according to the embodiments described herein. The UI 700 depicts a partial map of the United States. The presenter can use the toolbar 702 to interact with the UI 700 and the depicted content of the UI 700. In this example, the presenter selected the chapter creation tool 704 during the recording of the presentation to generate a chapter, as indicated by the indicator message 708 that notifies the presenter that two chapters have been generated.
[0114] The chapter creation tool 702 can be used by the presenter to annotate the recording video with respect to time. For example, the presenter can select the tool 702 at any point during the presentation to generate a chapter of the recording video. A chapter represents a section of the video. The chapter can provide a preview image frame that helps the user identify the content of the chapter. The chapter can also include (or trigger the storage of) metadata, title data, or identification data added by the user or the system. The video split by chapters can be presented in a timeline view, enabling the user to select the previously configured chapter indicators presented on the timeline.
[0115] As shown in FIG. 7, a pass-through view 706 provided in any part of the presentation UI space can be generated using a self-shot camera stream (e.g., a presenter video stream). The presenter can be the presenter or the presenter of video and audio content. The presenter video stream can be automatically placed throughout the recording, for example, at a location on the screen that ensures that the stream does not interfere with the view of the content being annotated. In some embodiments, the presenter can drag the presenter video stream of view 706 within the presented UI content. In some embodiments, the presenter can shrink or expand view 706. In some embodiments, the presenter can crop view 706. In some embodiments, the presenter can hide view 706.
[0116] FIG. 8 shows a flow diagram of an example of using a real-time presentation system according to the embodiments described herein. In this example, a presenter can use system 100 to present ideas or content. During operation, a user can access system 100 via a quick settings UI (such as UI506 or UI512). The user can select (804) the destination of the presentation. For example, the user can present via casting or via a video conference. Next, the user can select (806) the scope of the screen to be shared. For example, the user can select to share one or more screens, one or more browser tabs, one or more applications, one or more windows, etc.
[0117] In some embodiments, a user may wish to record a screen cast of a presentation and can do so by selecting (808) to also record the presentation. Then, the screen cast recording can be started. In some embodiments, the quick settings UI can provide the option to cast, share, and record with a single input command. Thereafter, the user can perform the presentation and can generate annotations, chapters, and other data (810). The user can choose to stop the presentation by selecting a presentation stop control (812). If the user has selected to record a presentation (e.g., a screen cast), the presentation can be ended by stopping the recording, thereby triggering system 100 to end the recording and complete the upload of the recording to the repository (814).
[0118] FIG. 9 is a screenshot 900 showing an example of a transcript 902 generated by a real-time presentation system according to an embodiment described herein. The view of the screenshot 900 can be provided after recording a presentation / screen cast. The system 100 may generate the transcript 902 in real time as the recording occurs. Additionally, the presenter may create annotations that mark key ideas 904 and 906 during the recording. The presenter may perform post-recording annotation and markup to make the video content useful to other users. For example, the presenter may decide to generate additional annotations and / or key idea markings such as key ideas 908 and 910, and may do so after recording. The new key ideas and / or annotations can be part of a video stream that can be added to the recording data. Similarly, the presenter may add additional audio data by recording additional content. The transcription 902 may be updated with the new audio data. Additionally, the transcription 902 may be changed in other ways to add or remove content after recording.
[0119] In some embodiments, the system 100 can automatically highlight specific content that is being accessed after recording. The highlighted content can indicate some kind of mistake or error to the presenter. The highlighting draws attention to the mistake or error so that the presenter can correct the error, for example, before spreading additional information (such as representative content 112, video stream, etc.) with the recording. In some embodiments, the system 100 can indicate areas for providing additional information. For example, the presenter can add titles, labels, etc. to the key ideas.
[0120] In some embodiments, system 100 can utilize machine learning techniques to learn and correct specific errors. In some embodiments, system 100 can utilize machine learning techniques to learn what content should be presented to the presenter in order to provide a list of items to be updated and / or corrected. In some embodiments, system 100 can utilize machine learning techniques to automatically generate a title and additional content from a recording so that the presenter can select what updates should be applied or added to the recording.
[0121] As shown by UI912, the presenter can also add closed-captioned content and / or translated content. In some embodiments, the user can use control 914 to select one or more languages and provide transcript content, closed-captioned content, and / or translated content for as many languages as the presenter has decided to provide.
[0122] FIG. 10 is a screenshot showing an example of presenting recorded content to a user of a real-time presentation system according to an embodiment described herein. In this example, the presenter may have completed a recording, a part of which is shown in screenshot 1000. In response, system 100 can analyze and index the content of the recording (e.g., any or all video streams, annotations, transcripts, translations, audio, presentation content, or resources accessed during the presentation, etc.). The analysis can further include determining which content of the recording should be used to generate a part of the video content (e.g., a representative or recap video or snippet, a learning guide, an audio track, etc.). Such content can be generated based on the metadata recording and can include parts of the video content annotated by the presenter (or by a user associated with the presenter video stream). In some embodiments, the summary video can also include other parts of the video content that are not annotated but are instead selected to be included in the representative content.
[0123] As shown in FIG. 10, system 100 generated a video snippet 1002 that discusses translation and transcription related to ribosomes within a cell. The presenter can provide an indicator, title, and / or message to be presented along with video snippet 1002, as indicated by presented item 1004. The item may be presented based on annotations generated by the presenter. A user who receives presented item 1004 can select a link, video, or other information to obtain the information presented by item 1004 and / or to respond or comment regarding the item.
[0124] The user can also use control 1006 to search for content, metadata, or other streams associated with the recording within the recording. In this example, the user is entering a search query for the term "cell structure". In response, system 100 can provide, along with item 1004 presented as search results, a highlighted portion of a transcription (or translation) that includes the search term, as indicated by highlight 1008. Additionally, system 100 can highlight additional transcription or translation content 1010 that may be relevant to the search query.
[0125] FIG. 11 is a screenshot showing another example of presenting recorded content to a user of a real-time presentation system according to an embodiment described herein. In this example, a web browser application 1102 executing the system 100 depicts educational content in a window 1104, for example. The system 100 can generate representative content 112, as shown by the menu 1106 and the UI 1108. The representative content of the menu 1106 can include an example menu 1106 accessed by a user viewing the content within the window 1104. The menu 1106 includes available video snippets 1110 related to the topic presented in the window 1104. In some embodiments, the video snippets 1110 can include snippets or image frames of content presented for a particular topic or date. In some embodiments, any number of video snippets and / or links can be embedded in the menu 1106 to provide the user with quick answers and content. Thus, instead of presenting results from the Internet, the system 100 can present search results from previously accessed content accessed locally, from an online library, from an online drive, and / or from another repository. In some embodiments, the system 100 can prioritize displaying recently accessed or viewed key idea snippets (e.g., video clips). The menu 1106 can be provided at a time useful to a user accessing the menu. Additionally, related searches may be presented as options in the menu 1106. For example, a user accessing the menu 1106 is provided with a search for the term "ribosome" 1112 based on the topic being discussed in the content of the window 1104.
[0126] System 100 can present the recorded content to the user in other ways. For example, the menu 1114 provided by the O / S can present additional content associated with the recording corresponding to the content provided in the window 1104 or to the window 1104. In this example, the O / S presented the search results in the UI1108. In some embodiments, the system 100 can present content in the UI1108 based on the search query 1120 input by the user. For example, the input search query 1120 can be matched with the key ideas from the video recording associated with the window 1104 and presented as search results generated by the O / S.
[0127] As shown, the UI1108 includes, as the top search results, a timeline 1116 of the video and key ideas. The user can select any of the events listed in the timeline 1116 to be guided to the video portion containing such content in the window 1104 or a new window. Additionally, the UI1108 also includes one or more videos 1118 related to the content accessed in the window 1104.
[0128] In some embodiments, the content presented in a UI, such as menu 1106 and / or UI 1108, can also be retrieved from sources other than the specific recorded video accessed in window 1104. For example, system 100 can retrieve content from another presenter or another presentation similar to the presentation being accessed in window 1104 (or similar to the content in the presentation) to populate menu 1106 and / or UI 1108. Thus, system 100 can utilize content from other presenters, companies, users, and / or one or more authoritative sources or resources related to the topic determined to be related to the content accessed in window 1104.
[0129] FIG. 12 is a screenshot showing an example of presenting key ideas and content marked during the recording of a session generated by a real-time presentation system according to the embodiments described herein. In this example, the user may be using an extension, application, or O / S that provides and initiates a screen cast. For example, browser window 1200 may be shared using system 100. The shared content includes at least a timeline 1202 having key ideas 1204, 1206, and 1208 corresponding to respective timestamps 1210, 1212, and 1214. Timeline 1202 may be generated, for example, by the presenter 1216 of the content during the presentation. Alternatively, the presenter may generate the key ideas and timeline 1202 after the completion of the video recording. It can be seen that the transcript is synchronized with timeline 1202 such that scrolling of one of content 1216 or the transcript causes the corresponding scrolling of the other.
[0130] Figures 13A - 13G show screenshots depicting marked content composed by a user accessing the real - time presentation system 100 according to the embodiments described herein. In this example, the user may be using an extension, application, or O / S that provides and initiates a screen cast. While the browser window 1304 is being cast by the online real - time presentation system 100, the toolbar 1302 is depicted. The toolbar 1302 can be initiated when the cast of the browser window 1304 is started, thereby enabling the presenter to select tools to initiate a tele - presentation (e.g., annotation of video or still video content). In some embodiments, for example, if the presenter uses a stylus, smart pen, or other such tool to provide input to the presentation content, the toolbar described herein may be omitted.
[0131] Referring to FIG. 13A, the toolbar 1302 includes a pointer tool, an eraser pen tool, a pen tool, a closed caption tool, a mute tool, and a key idea marker tool 1306. The marker tool 1306 can represent a control that the presenter can select to mark, for example, specific content, ideas, slides, annotations, or other presented portions of the screen as key ideas. A key idea can represent an element that the presenter considers useful, important, a learning guide material, and / or selectable with respect to representative content 112. Generally, key ideas can be organized by date, time stamp, and / or topic.
[0132] In this example, the presenter is using a pen tool to input text 1308 and / or highlights 1310 and 1312. The presenter then selects the marker tool 1306 and then marks the annotations of text 1308 as well as highlights 1310 and 1312, which may indicate that such content is a key idea. In response, the system 100 can provide an indicator message 1314 to provide feedback to the presenter regarding the ideas marked as key ideas. In some embodiments, the marker tool 1306 can be used to generate a chapter (such as a video marker that generates marker data, a chapter marker that generates marker data, etc.) that can be provided as annotation input alongside the telemetry data (i.e., highlights 1310 and 1310 and / or text 1308). The presenter can use the marker tool 1306 and / or other toolbar tools in real time and during recording to mark such annotation input with telemetry and key ideas. For example, during a presentation, the presenter can interactively mark chapters, annotations, key ideas, etc. Annotations obtained from the bidirectionality can be used to generate learning guides, representative content 112, video snippets, and searchable content so that the system 100 enables a user (such as a presentation participant) to easily access a recap video of the key ideas and / or annotations.
[0133] Referring to FIG. 13B, browser window 1304 is shown with an additional transcript section 1316. Transcript section 1316 can be generated in real time while a presenter is speaking using system 100 and presenting content within window 1304. Transcript section 1316 can represent the currently recorded transcript video stream. Transcript section 1316 may highlight the currently spoken sentence, as indicated by highlight 1318. If a user is accessing the recorded video after the recording is complete, the currently spoken sentence can be highlighted and updated as the speech (e.g., audio) is provided throughout the video. This can provide the advantage that the user can follow along with what is being said in transcript section 1316. As the audio progresses, the highlight updates to indicate the specific audio being spoken.
[0134] In some embodiments, the presenter or user can access the recording after completion and navigate the transcript to update the content within window 1320 according to the selected transcript in section 1316. For example, the user can select a paragraph in the transcript, navigate to the beginning of the paragraph, and trigger the matching content within window 1320. Additionally, the user can access search control 1322 to search the transcript for content. Browser window 1304 also depicts a sharing option 1324 that allows the presenter or user to share a specific full recording, a portion of the transcript, a portion of window 1320, or other parts of the video recording.
[0135] Referring to FIG. 13C, a browser window 1304 is shown, which includes additional options. For example, a marker tool 1326 is provided for the paragraphs of the transcript, enabling the user to mark (or unmark) specific portions of the transcript (and the resulting video portion associated with the transcript) as key ideas. For example, the user marks a paragraph as a key idea 1328 by selecting the marker tool 1326. The user can mark or unmark paragraphs within the transcript throughout the video. The marked portions can be accessed by the system 100 to generate representative content 112. Marking a transcript portion can function to automatically select the related video stream at the same timestamp (or multiple timestamps). Thus, if a specific transcript paragraph is marked as a key idea, other content can also be marked as a key idea at the same timestamp or in its vicinity. That is, marking one video stream can function to mark other video streams as key ideas, including but not limited to annotations (e.g., via an annotation video stream), translations (e.g., via a translation video stream), screen content (e.g., via a screen cast video stream), and camera views (e.g., via a presenter video stream).
[0136] Referring to FIG. 13D, again a browser window 1304 is shown, and the key idea marking shown in FIG. 13D is depicted on a timeline 1330 where a key idea 1328 is marked at a time stamp 1332 within the video. An indicator 1334 depicts a portion of the transcript 1316. The indicator can be a video snippet or an image frame that assists the user in identifying the content at the key idea's time stamp 1332. In some embodiments, the user can use the timeline 1330 to mark, unmark, or otherwise change the marked key ideas.
[0137] Referring to FIG. 13E, again a browser window 1304 is shown, and additional key ideas are marked. For example, a Partial order key idea 1336 and an Untitled key idea 1338 are marked by a user using the system 100. Corresponding time stamps 1340 and 1342 are also generated for the timeline 1330. In one example, the user selected a paragraph 1344 to trigger the concept 336. Additionally, when the user selects a particular translated paragraph (or other content the user uses to generate a key idea), an editing tool 1346 can be provided. Any portion of the transcript can be edited using the editing tool 1346. In some embodiments, the editing tool 1346 can be used to combine and / or split portions of the transcript, thus triggering possible changes to the key ideas.
[0138] Referring to FIG. 13F, the user selected an editing tool 1346 to edit a transcription portion 1344 that can trigger an edit to the key idea 1336 of the timeline 1330. In response to selecting an editing tool in portion 1344, system 100 can present a UI 1348. UI 1348 can provide an entry for changing the title of the key idea using control 1350, and an entry for changing any portion of the actual transcript shown in control 1352. Additionally, UI 1348 can provide a control for joining or splitting a portion of the transcription that can trigger joining or splitting of the key idea. Such changes to the key idea can change the underlying video frames, text, and context of the key idea.
[0139] Referring to FIG. 13G, in response to the user entering a search 1360, a number of search results 1354, 1356, and 1358 are presented. Such search results can be generated by system 100. For example, after a presenter (or other user) generates key ideas and annotations for a video provided by system 100, system 100 can be configured to make the video (as well as the underlying video stream and associated metadata) searchable. When the user searches for content associated with the video (using a search engine), the search engine can return search results (text, video, images, etc.) that include the video and / or a portion of the associated content.
[0140] As shown in FIG. 13G, the search includes a search term set and subsets. The system 100 can provide search results 1354-1358 because it enables a search function to find at least a portion of representative content using a web browser application, and thus can perform or trigger indexing of a portion of representative video content (e.g., key ideas, transcriptions, annotations, inputs, etc.). Specific URL links can be generated to direct the user to a portion of a video or text that includes representative content. In some embodiments, video search results can be provided that are selected to direct the user to a location (e.g., a timestamp) in the video that correlates the searched terms to matching key ideas. Each search result can be configured to include a video thumbnail and timestamp, a title, highlighted transcriptions (e.g., highlights 1362, 1364, and 1366), a username, and a timestamp of the uploaded video.
[0141] FIG. 14 is a screenshot showing translated text presented in real time during recording of a session generated by a real-time presentation system 100 according to embodiments described herein. For example, in addition to a closed caption version 1402 of the recording and / or the audio being presented, the system 100 can also generate and render a real-time translation 275 shown as text 1404. The user can use control 1406 to select the language in which a particular translation is displayed. The translation in the selected language may, in some examples, form part of a transcription video stream or be provided as a separate translation stream.
[0142] Closed captions can be toggled on or off with tool 1408 in toolbar 1410. By providing closed caption content 1402, it can be made easier for the user to follow the presentation. With real-time translation content 1404, a user learning the presenter's language can follow the presentation. In some embodiments, the user can access a previously recorded video that includes a translation in a first language and can select a second language to display the translation in the second language. This can be useful for a user seeking assistance from a parent or other user who does not speak the language of the presentation.
[0143] FIG. 15 shows a flowchart of an example process 1500 for generating and recording a screencast according to an embodiment described herein. The presenter can configure computing system 202 to generate a screencast that starts, for example, from one or more libraries 116 associated with real-time presentation system 100. The library can include content associated with the presenter that can be stored on a local storage drive, an online storage drive, server computing system 204, or another location accessible to computing system 201 and / or computing system 202. The presenter can enter library 116 and select to start recording a screencast (1502). Next, the presenter can select the scope of the content to record (e.g., window, tab, full screen, etc.) (1504). System 100 may activate a screencast / screen share tool to trigger a UI for selecting the scope. The user is recording a screencast, but for example, if the screencast recording is for later viewing by the user, the user can select not to share the screen.
[0144] Next, system 100 can start recording according to the selected scope and can present one or more toolbars (e.g., toolbar 108). The presenter can use a screen casting tool (e.g., toolbar 108) to annotate the content (1506). The presenter can choose to end the recording at some point. When the recording ends, system 100 can automatically upload the video (as well as the corresponding video stream and metadata) to library 116 as a newly available file. In some embodiments, system 100 is configured such that the video is viewed by and shared with other people.
[0145] FIG. 16 shows a flowchart of a process example 1600 for generating metadata records associated with multiple video streams, according to embodiments described herein. Generally, process 1600 utilizes the systems and algorithms described herein to generate the metadata records used by real-time presentation system 100. Process 1600 can utilize one or more computing systems comprising at least one processing device and a memory storing instructions that, when executed, cause the processing device to perform the multiple operations and computer-implemented steps recited in the claims. Generally, system 100, system 200, system 263, and / or system 1900 can be used in the description and execution of process 1600.
[0146] In block 1602, process 1600 includes starting a recording to capture video content. The video content can include any or all of a presenter video stream, a screen-cast video stream, a transcription video stream, and / or an annotation video stream. For example, system 100 can be accessed by a user (e.g., a presenter) to start a recording that captures video content. Such video content can include a presenter video stream (e.g., content captured by a selfie camera), a screen-cast video stream (e.g., the content of drawing 276 and screen-cast 277), an annotation video stream (annotation data record 214 and / or key idea marker and corresponding metadata 278), a transcription video stream (e.g., real-time transcription 274), and / or a translation video stream (e.g., real-time translation 275).
[0147] In block 1604, process 1600 includes generating a metadata record representing timing information during capture of the video content based on the video content. The timing information can be used to synchronize an input received in at least one of a presenter video stream, a screen cast video stream, a transcription video stream, or an annotation video stream with a portion of the video content. In some embodiments, the input includes annotation input associated with the annotation video stream. In some embodiments, the annotation can include drawings 276, text, audio input, reference links, and the like. In some embodiments, the annotation input includes video marker data and / or teleprompter data generated by a user associated with the presenter video stream. For example, a presenter can input an annotation using a teleprompter that inputs drawings, text, etc. as an overlay to the video content. Similarly, a presenter can mark chapters using a marker tool during recording. The chapters can be stored as video marker data that can be used to generate chapters of the video content.
[0148] In some embodiments, each metadata record represents timestamp data used to synchronize an input received in at least one of the recorded video streams (e.g., annotation 114 / record 214, key idea metadata 278). In some embodiments, metadata 228 can be captured and stored during recording. Metadata 228 can be associated with any number of video streams and annotations received during or after the recording of the video stream. Each video stream can also include audio data. In some embodiments, the video stream can store annotation data as metadata. However, in some embodiments, the annotation data may be recorded separately as a video layer, and thus, metadata 228 may be obtained from the video layer.
[0149] In some embodiments, process 1600 includes generating content representative of a portion of video and / or audio content based on metadata records. For example, representative content can include portions of video content annotated by a user (e.g., a presenter) associated with the presenter video stream in response to the end of the recording. The video content can include representative content 112 and can be generated based on timing information, metadata 228, and / or other video content or annotations of the video content. Generation may be performed automatically in response to the end of the recording or may be initiated by the user or in another manner in response to user input when the recording has ended. In some embodiments, representative video content can include an overlaid image frame depicting the rendered video content and / or annotations on the screen content. In some examples, representative content can also include one or more portions of the video content from immediately before and / or after each portion of the video content annotated by the user.
[0150] In some embodiments, the timing information corresponds to a plurality of timestamps associated with each received input. For example, the timing information can correspond to annotations received during a recording and / or screen cast (e.g., provided by a presenter). The received annotations can be provided at a particular one or more timestamps. The timing information can also correspond to at least one position in the content or document associated with the presenter video stream, screen cast video stream, or annotation video stream in which the input was received (i.e., in the content or document associated with the video content). For example, the timing of annotation creation also corresponds to the (spatial) position within the screen / video / content where the annotation was placed during the period including the timestamp. In some embodiments, synchronizing the inputs includes matching at least one of the plurality of timestamps for each input to at least one position in the content or document. For example, system 100 can perform a matching process that matches an annotation or marker input to a position in the video content and the point in time associated with receiving the annotation or marker input during the recording of the video content.
[0151] In some embodiments, the video content further includes a transcription video stream in addition to a plurality of other video streams. The transcription video stream can include real-time captioned audio data from the presenter video stream. The real-time captioned audio can be generated as modifiable transcription data (e.g., text data) configured to be displayed along with the screen-cast video stream during recording of the video content. That is, the transcription can be generated and rendered in real-time or near real-time as the presenter records and presents the content. In some embodiments, real-time translated audio data from the presenter video stream is generated as text data configured to be displayed along with the screen-cast video stream and the captioned audio data during recording of the video content. For example, during recording, the transcription can be rendered along with other video stream content from the screen-cast. In some embodiments, system 100 can also perform and render a translation of the transcription using the text data of the transcription video stream. Thus, the text (transcription) data can be rendered with or without translation.
[0152] In some embodiments, the transcription of real-time speech-to-text audio data is performed by at least one speech-to-text application. The at least one speech-to-text application can be selected from any number of speech-to-text applications determined to be accessible by the transcription video stream. For example, system 100 can determine which speech-to-text application can provide an accurate and convenient transcription for the audio content. Such determination can be made based on the audio content, the language of the audio content, demographics provided by the user presenting or accessing the video stream, and the like. The modifiable transcription data and text data can be stored according to timestamps within the metadata record and configured to be searchable. This can facilitate the search for content within the video stream in an effective and resource-efficient manner.
[0153] In some embodiments, the presenter video stream, the screen-cast video stream, and the annotation video stream are configured to be switched on and off during recording. The on-off switching can trigger the display (or removal from display) of each presenter video stream, each screen-cast video stream, or each annotation video stream.
[0154] FIG. 17 is a flowchart of an example process for generating and recording a video presentation in a real-time presentation system according to an embodiment described herein. Generally, process 1700 utilizes the systems and algorithms described herein to generate metadata records used by real-time presentation system 100. Process 1700 can utilize one or more computing systems comprising at least one processing device and a memory storing instructions that, when executed, cause the processing device to perform the recited operations and computer-implemented steps. Generally, in the description and execution of process 1700, system 100, system 200, system 263, and / or system 1900 can be used.
[0155] Real-time online presentation system 100 can be a system that includes at least one camera, at least one microphone, at least one speaker, at least one display screen, and one or more user interfaces configured to be displayed on the at least one display screen. System 100 can execute the instructions of process 1700 using at least one processor and one or more computer-readable hardware storage devices storing computer-executable instructions executable by the at least one processor.
[0156] In block 1702, process 1700 includes initiating a recording that captures audio content and video content. For example, a presenter can access system 100 to trigger a presentation and / or recording to begin capturing the audio content and video content being presented, which can ultimately result in generating recordings 110, 110b, and / or annotations 114. As described throughout this disclosure, the video content can include at least a presenter video stream, a screen cast video stream, a transcription video stream, and an annotation video stream. In some embodiments, as discussed with reference to FIG. 16, metadata records can be generated based on the video content.
[0157] In block 1704, process 1700 includes causing the rendering of audio content and video content associated with access to a plurality of applications from within a user interface. For example, during the presentation and recording of audio and video content, system 100 may trigger content sharing (e.g., screen share, video conference share, screen cast, etc.). The video data can be rendered via a screen that provides various UIs, and the audio content can be rendered via speakers. In some embodiments, the audio content is also rendered as text that is captioned and / or translated near or within a threshold distance of the remaining content being presented by system 100.
[0158] In block 1706, process 1700 includes receiving annotation input in the user interface during the rendering of audio and video content. The annotation input may be recorded in an annotation video stream. For example, when a user annotates video content (e.g., annotations 306, 308 in FIG. 3A), system 100 may record the annotation in a separate stream that can be represented as an overlay that can be placed on content from other video streams captured by system 100. In some embodiments, the annotation input is rendered as an overlay on the video content. The annotation input can also be configured to move with the video content in response to detecting a window event or cursor event that triggers a switch to other video content (e.g., an application, window, browser tab, etc.) accessed during the recording. For example, a window event or other signal indicating a scroll of the window can be received, and the annotation input can be configured to scroll with the content of the underlying application such that the annotation remains in a fixed position relative to the underlying, annotated, application content.
[0159] In block 1708, process 1700 includes performing speech-to-text on the audio content during the rendering of the audio content and video content. For example, the audio content is speech-to-text in real time. The speech-to-text audio content can be recorded to a transcription video stream and rendered and marked in real time by system 100. For example, a presenter (or a user watching the presentation) can mark, annotate, modify the transcription data, or otherwise interact with the transcription data presented in the UI provided by system 100.
[0160] In block 1710, process 1700 optionally includes translating the audio content during the rendering of the audio content and video content. For example, the translation can be performed in real time. The translation can include translating the text presented in a screencast (or other sharing mechanism) in addition to translating the audio information occurring during the presentation.
[0161] In block 1712, process 1700 includes causing, in a user interface, the rendering of the speech-to-text audio content (and optionally the translated audio content) in real time along with the rendered audio content and video content. For example, educational / presentation content, speech-to-text content, and optional translated content can be depicted in a single UI so that a presenter and a user watching the presentation can conveniently access the video stream presented in one view. In some embodiments, additional video streams, such as a presenter video stream, an annotation video stream, a participant video stream, etc., are added to such views.
[0162] In some embodiments, process 1700 may also include causing online presentation system 100 to generate summary content in response to detecting the end of the rendering of video content and audio content. The summary content may be, for example, representative content 112, and content 112 may be based on annotation input, video content, transcribed audio content, and translated audio content (i.e., content 112 may include portions of video content selected or determined based on annotation input, transcribed audio content, etc.). The summary content may be generated based on the generated metadata record. In some embodiments, the summary content includes portions of the rendered audio and video marked with annotation input.
[0163] FIG. 18 is a flowchart of an example process 1800 of presenting a video presentation in a real-time presentation system according to an embodiment described herein. Generally, process 1800 utilizes the systems and algorithms described herein to generate metadata records used by real-time presentation system 100. Process 1800 can utilize one or more computing systems comprising at least one processing device and a memory storing instructions that, when executed, cause the processing device to perform the recited plurality of operations and computer-implemented steps. Generally, in the description and execution of process 1800, systems 100, 200, 263, and / or 1900 can be used.
[0164] In step 1802, process 1800 includes receiving at least one video stream. For example, a user can access system 100 to view presentation content (e.g., video and audio content). The user may select a recording to view or use system 100 to view the recording live. In response to indicating which recording to view, system 100 can trigger system 202 to receive, for example, one or more of a plurality of video streams. The video stream can include, but is not limited to, at least a presenter video stream, a screen cast video stream, a transcription video stream, and an annotation video stream, as described throughout the present disclosure.
[0165] In step 1804, process 1800 includes receiving metadata representing timing information associated with an input detected in at least one video stream. For example, system 100 can trigger system 202 to receive metadata 228 representing timing information. The timing information can be configured to synchronize a detected input provided in at least one video stream with the content (e.g., video, audio, data, metadata, etc.) of at least one video stream. For example, the timing information can include information and / or instructions configured to synchronize a detected input (e.g., an annotation, a marker, etc.) with at least one of a plurality of video streams.
[0166] In step 1806, process 1800 includes generating a portion of at least one video stream based on metadata. That portion can be generated in response to receiving a request to view any or all of the at least one video stream. For example, a user can request to view content associated with a video stream. In response, system 100 can generate a summary video, recap video, or other representative video (and / or audio) as a compilation or other combination of portions of the video stream based on the metadata.
[0167] In some embodiments, system 100 can generate and present UI 302, and annotations 306 and 308 retrieved from the metadata are depicted as an overlay on the content shown in UI 302. UI 302 can depict annotations 306 and 308 as an overlay on the content within UI 302 at the timestamps indicated in the metadata in response to a detected user instruction to display compiled content (e.g., summarized content, recap content, and / or other representative content) associated with a plurality of video streams. The generated portion can include annotation content, video content, or other video and / or audio content representing content requested by the user and / or provided by system 100. In some embodiments, the generated portion includes content based on the detected input and includes a rendered portion of the video stream annotated with the input.
[0168] In some embodiments, the entire screenshot shown in FIG. 3A can be provided as an image frame in response to detecting a request to display compiled or otherwise curated content, since the frame contains annotated content. The annotated content can be an indicator indicating that the information within the image frame contains key data, as shown by a presenter associated with the content of at least one video stream.
[0169] In step 1808, process 1800 includes causing the rendering of the portion of the at least one video stream in at least one user interface. For example, UI generator 220 formats and displays the portion shown as compiled (e.g., recapitulated, summarized) content using a renderer. In response to a request to display a compilation or other combination of content, other portions of the video stream can also be displayed or alternatively displayed. For example, video and / or audio content such as presenter video streams, translated video streams, transcription video streams, another annotated video stream, and / or video and / or audio content associated with other video streams generated by system 100 can also be depicted.
[0170] In some embodiments, the timing information corresponds to a plurality of timestamps associated with each input detected in one or more of the video streams and at least one position in the content or document associated with at least one of the one or more video streams (i.e., in the content or document associated with at least one video stream). In some embodiments, synchronizing the detected inputs includes matching at least one timestamp to at least one position in the document for each input.
[0171] In some embodiments, the recorded video can be opened with a native application on a device (e.g., desktop, tablet, mobile device, wearable device, etc.). The native application can provide additional tools that allow the user to read the transcript of the video recording, navigate the video recording by selecting the transcript, skip / skim between key ideas, search within and between videos, and / or view key ideas across the extent of the video (e.g., show all moments of "this is what will be on the test" from a presentation to prepare employees for a test). In some embodiments, the recorded video and system 100 may be provided as an application extension instead of a native application.
[0172] During operation of system 100, the presenter can be provided with the option to mark key ideas, draw them in real time over the recording, and store such annotations and the recording online as any number of separate video streams to facilitate the generation of content 112 for recording. At the end of the recording, the presenter can review the recording and upload it to an online drive to share directly with one or more applications and / or users. With system 100, the presenter can create a narrated screencast that the user can view later, record and share presentations and related content asynchronously, conduct in-person presentations, and prepare remote presentations via video conferencing software and related applications.
[0173] The systems and methods described herein can provide a screen share scope selection tool (e.g., presentation system 100). The tools of system 100 can provide the user with the option to select a presentation mode (e.g., extended display or mirror display mode, etc.) while connected to an external display (e.g., a TV or projector hardware) that also includes access to a presenter toolbar. The presenter toolbar can include a casting destination tool, a screen share panel, a screen share recording tool, a screen share stop tool, a telestretching tool, a laser pointer tool, a closed caption tool, a camera tool, a markup tool, and any number of annotation tools (e.g., pen, highlighter, shapes, etc.). The telestretching tool can enable the user to telestretch to any location on the screen. Alternatively, the presenter toolbar can be omitted and a stylus can be used directly for annotation. The closed caption tool option can provide live captions and translations on the device over highlighted text, for example, based on input from a microphone associated with system 100. The language of the translation can be selected by the user or provided in text form. In some examples, the translated text can be synthesized and output to the user as audio data.
[0174] When the user selects the recording option from the presenter toolbar or the screen share panel, the current screen share scope becomes active, and the tool asks the user whether to record and upload to the cloud server. The toolbar can provide the first user with the option to move to a screen share scope selection tool for trimming and publishing the recording when the recording is triggered via the screen capture tool. The markup option (i.e., the star option in toolbar 400) can enable the user to markup important / key ideas presented on the screen and display indicator text for confirming the marking.
[0175] The toolbar can automatically perform speech recognition on the captured recording, allow the user to highlight the text for accuracy checking, and ask the user to provide a title for the key idea before uploading to the repository to share the recording with other users of system 100.
[0176] Since the key ideas in system 100 are organized by date and topic, another user can search the transcript via the search bar provided when the user accesses the recording, navigate through the transcript and / or key ideas, and be able to view recap (summary, representative part) videos of all key ideas on a pre-determined time basis (e.g., daily, weekly, monthly, quarterly, annually, etc.). System 100 can highlight the current sentence (being read) in the transcript and enable the user to edit the title, transcript, and mark the key ideas of the paragraphs. The system can display the recording clip as a search result or a quick answer in the browser when the user's query matches the recorded key idea.
[0177] In some embodiments, system 100 can provide a reading assistance UI for parallel displays. For example, system 100 can provide reference assistance with parallel display e-books to hold the context for reading and referring to content while reading. A user can select any text within system 100 and upload the text. System 100 can proactively propose useful learning moments using the uploaded text. For example, system 100 can provide key concepts and present articles and videos related to those concepts, such as glossary-style related content. In some embodiments, system 100 can adjust the Lexile® level of specific text. For example, system 100 can replace particularly advanced words in the text with simpler terms to tailor the content to users with a limited vocabulary. In some embodiments, system 100 can replace specific content with less advanced content to help the reader understand the sentences of the content. Then, system 100 can switch back to the original content to enable further understanding of the usage of the vocabulary in the text.
[0178] In some embodiments, system 100 can also provide context learning moments. For example, system 100 can incorporate paragraph translation for users who have a first learning language different from the language of the text. System 100 can also provide quick links for vocabulary search and / or answer search.
[0179] In some embodiments, system 100 can provide access to accessibility features such as text-to-speech with speed, pitch, and accent adjustments. In some embodiments, system 100 can provide fonts that assist dyslexic readers in reading sentences, and can also highlight sentences and / or words read aloud by system 100. To assist the user in learning the presented concepts, other highlighting, annotation, and data synthesis can be performed by system 100.
[0180] FIG. 19 shows an example of a computer device 1900 and a mobile computer device 1950 that can be used with the techniques described herein. Computing device 1900 is intended to represent various forms of digital computers, such as laptops, desktops, tablets, workstations, personal digital assistants, smart devices, appliances, electronic sensor-based devices, televisions, servers, blade servers, mainframes, and other appropriate computing devices. Computing device 1950 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. The components shown here, their connections and relationships, and their functions are intended to be exemplary only, and are not intended to limit the embodiments of the invention described and / or claimed herein.
[0181] Computing device 1900 includes a processor 1902, a memory 1904, a storage device 1906, a high-speed interface 1908 that connects to the memory 1904 and a high-speed expansion port 1910, and a low-speed interface 1912 that connects to a low-speed bus 1914 and the storage device 1906. The processor 1902 can be a semiconductor-based processor. The memory 1904 can be a semiconductor-based memory. Each of the components 1902, 1904, 1906, 1908, 1910, and 1912 can be interconnected using various buses and can be mounted on a common motherboard or in other manners as required. The processor 1902 can process instructions executed within the computing device 1900, including instructions stored in the memory 1904 or the storage device 1906 for displaying graphical information for a GUI on an external input / output device such as a display 1916 coupled to the high-speed interface 1908. In other embodiments, a plurality of processors and / or a plurality of buses can be used, along with a plurality of memories and a plurality of types of memories, as required. Also, a plurality of computing devices 1900 can be connected and each device can provide a portion of the required operations (e.g., as a server bank, a blade server group, or a multiprocessor system).
[0182] The memory 1904 stores information within the computing device 1900. In one embodiment, the memory 1904 is one or more volatile memory units. In another embodiment, the memory 1904 is one or more non-volatile memory units. The memory 1904 can also be another form of computer-readable medium, such as a magnetic disk or an optical disk. Generally, the computer-readable medium can be a non-transitory computer-readable medium.
[0183] The memory device 1906 can provide large-capacity storage for the computing device 1900. In one embodiment, the memory device 1906 can be or include a computer-readable medium, such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory, or other similar solid-state memory devices, or an array of devices including a storage area network or other configured devices. The computer program product can be tangibly embodied in an information carrier. The computer program product can also include instructions that, when executed, perform one or more of the methods as described above and / or methods implemented by a computer. The information carrier is a computer or machine-readable medium such as the memory 1904, the memory device 1906, or the memory on the processor 1902.
[0184] The high-speed controller 1908 manages operations that use a large amount of the bandwidth of the computing device 1900, while the low-speed controller 1912 manages operations that do not use as much bandwidth. Such functional assignments are merely exemplary. In one embodiment, the high-speed controller 1908 is coupled to the memory 1904, to the display 1916 (e.g., via a graphics processor or accelerator), and to a high-speed expansion port 1910 that can receive various expansion cards (not shown). In this embodiment, the low-speed controller 1912 is coupled to the memory device 1906 and a low-speed expansion port 1914. The low-speed expansion port, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices such as a keyboard, a pointing device, a scanner, or a network device such as a switch or router, e.g., via a network adapter.
[0185] As shown in the figure, computing device 1900 can be implemented in a number of different forms. For example, it may be implemented as a standard server 1920, or multiple times in such a server group. It may also be implemented as part of a rack server system 1924. In addition, it may be implemented in a computer such as a laptop computer 1922. Alternatively, the components of computing device 1900 may be combined with other components within a mobile device (not shown), such as device 1950. Each of these devices can include one or more of computing devices 1900 and 1950, and the entire system can be composed of a plurality of computing devices 1900 and 1950 that communicate with each other.
[0186] Computing device 1950 includes other components, but particularly includes a processor 1952, a memory 1964, input / output devices such as a display 1954, a communication interface 1966, and a transceiver 1968. In addition, device 1950 can be equipped with a storage device such as a microdrive or other device to provide additional storage capacity. Each of components 1950, 1952, 1964, 1954, 1966, and 1968 is interconnected using various buses, and some of the components can be mounted on a common motherboard or in other ways as needed.
[0187] Processor 1952 can execute instructions within computing device 1950, including instructions stored in memory 1964. The processor can be implemented as a chipset of separate multiple analog and digital processors. The processor can provide adjustment of other components of device 1950, such as, for example, control of the user interface, applications executed by device 1950, and wireless communication by device 1950.
[0188] Processor 1952 can communicate with a user via a control interface 1958 and a display interface 1956 coupled to a display 1954. The display 1954 can be, for example, a TFT LCD (Thin Film Transistor Liquid Crystal Display) or an OLED (Organic Light Emitting Diode) display, or other suitable display technology. The display interface 1956 can comprise appropriate circuitry for driving the display 1954 to present graphical and other information to the user. The control interface 1958 can receive commands from the user and convert them for submission to the processor 1952. Additionally, an external interface 1962 can be provided that communicates with the processor 1952 to enable short-range communication of the device 1950 with other devices. The external interface 1962 can provide, for example, wired communication in some embodiments and wireless communication in other embodiments, and can also use multiple interfaces.
[0189] Memory 1964 stores information within computing device 1950. Memory 1964 can be implemented as one or more of one or more computer-readable media, one or more volatile memory units, or one or more non-volatile memory units. Extended memory 1974 is also provided and can be connected to device 1950 via expansion interface 1972, which can include, for example, a SIMM (Single In-line Memory Module) card interface. Such extended memory 1974 can provide additional storage space for device 1950 or can also store applications or other information for device 1950. Specifically, extended memory 1974 can include instructions for executing or complementing the processes described above and can also include secure information. Thus, for example, extended memory 1974 may be provided as a security module for device 1950 and may be programmed with instructions that enable secure use of device 1950. In addition, a secure application may be provided via the SIMM card along with additional information, such as by placing identification information on the SIMM card in a non-hackable manner.
[0190] The memory can include, for example, flash memory and / or NVRAM memory, as described below. In one embodiment, the computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more of the methods as described above. The information carrier is a computer-readable or machine-readable medium, such as, for example, memory 1964, extended memory 1974, or memory on processor 1952, which can be received via transceiver 1968 or external interface 1962.
[0191] Device 1950 can communicate wirelessly via a communication interface 1966 that can include a digital signal processing circuit if necessary. The communication interface 1966 can provide communication in various modes or protocols, such as GSM voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA (registered trademark), CDMA2000, or GPRS. Such communication can be performed, for example, via a radio frequency transceiver 1968. In addition, short-range communication may be performed, such as by using Bluetooth, Wi-Fi, or other such transceivers (not shown). In addition, a GPS (Global Positioning System) reception module 1970 can provide additional navigation and location-related wireless data to device 1950, and such wireless data can be used as needed by applications running on device 1950.
[0192] Device 1950 can also communicate audibly using an audio codec 1960, which can receive voice information from a user and convert it into usable digital information. The audio codec 1960 can similarly generate audible sounds for the user, such as via a speaker within the handset of device 1950. Such sounds can include sounds from voice calls, recorded sounds (such as voice messages, music files, etc.), and sounds generated by applications operating on device 1950.
[0193] As shown in the figure, computing device 1950 can be implemented in many different forms. For example, it may be implemented as a mobile phone 1980. It may also be implemented as part of a smartphone 1982, a personal digital assistant, or other similar mobile devices.
[0194] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include embodiments implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a memory system, at least one input device, and at least one output device.
[0195] These computer programs (also known as modules, programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or in assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0196] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, or an LED (light emitting diode)) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other types of devices can be similarly used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and the input from the user can be received in any form including acoustic, voice, or tactile input.
[0197] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a client computer having a graphical user interface or a web browser by which a user can interact with an embodiment of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), and the Internet.
[0198] A computing system can include clients and servers. The clients and servers are generally remote from each other and typically interact via a communication network. The relationship between the client and the server is created by computer programs that run on their respective computers and have a client-server relationship with each other.
[0199] In some embodiments, the computing device shown in FIG. 19 can include sensors that interface with virtual reality or a headset (VR headset / AR headset / HMD device 1990). For example, one or more sensors included in computing device 1950 or other computing devices shown in FIG. 19 can provide input to the AR / VR headset 1990 or, generally, provide input to the AR / VR space. The sensors can include, but are not limited to, touchscreens, accelerometers, gyroscopes, pressure sensors, biometric sensors, temperature sensors, humidity sensors, and ambient light sensors. Computing device 1950 can use the sensors to determine the absolute position and / or detected rotation of the computing device in the AR / VR space that can be used later as input to the AR / VR space. For example, computing device 1950 can be incorporated into the AR / VR space as a virtual object such as a controller, laser pointer, keyboard, weapon, etc. By positioning the computing device / virtual object by the user when incorporated into the AR / VR space, the user can position the computing device to view the virtual object in several ways in the AR / VR space.
[0200] In some embodiments, one or more input devices included in or connected to computing device 1950 can be used as input into the AR / VR space. Examples of input devices include, but are not limited to, touchscreens, keyboards, one or more buttons, trackpads, touchpads, pointing devices, mice, trackballs, joysticks, cameras, microphones, earphones or buds with input capabilities, game controllers, or other connectable input devices. When the computing device is incorporated into the AR / VR space, a user interacting with the input devices included in computing device 1950 can generate specific actions in the AR / VR space.
[0201] In some embodiments, one or more output devices included in computing device 1950 can provide output and / or feedback to the user of the AR / VR headset 1990 in the AR / VR space. The output and feedback can be visual, tactical, or auditory. Output and / or feedback can include, but are not limited to, rendering of the AR / VR space or virtual environment, vibration, turning on / off or flashing and / or strobing of one or more lights or strobes, sounding of an alarm, chiming, playing of a song, and playing of an audio file. Examples of output devices include, but are not limited to, vibration motors, vibration coils, piezoelectric devices, electrostatic devices, light-emitting diodes (LEDs), strobes, and speakers.
[0202] In some embodiments, to create an AR / VR system, computing device 1950 can be placed within AR / VR headset 1990. The AR / VR headset 1990 can include one or more positioning elements that enable a computing device 1950, such as a smartphone 1982, to be placed in an appropriate position within the AR / VR headset 1990. In such embodiments, the display of the smartphone 1982 can render a stereoscopic image representing the AR / VR space or virtual environment.
[0203] In some embodiments, computing device 1950 can appear as another object within a computer-generated 3D environment. Interaction by the user with the computing device 1950 (e.g., rotating, shaking, touching the touch screen, swiping a finger across the touch screen) can be interpreted as interaction with an object in the AR / VR space. As an example, the computing device can be a laser pointer. In such an example, computing device 1950 appears as a virtual laser pointer within the computer-generated 3D environment. When the user operates the computing device 1950, the user within the AR / VR space sees the movement of the laser pointer. The user receives feedback on the computing device 1950 or the AR / VR headset 1990 from the interaction with the computing device 1950 in the AR / VR environment.
[0204] In some embodiments, computing device 1950 can include a touch screen. For example, a user can interact with the touch screen in a specific way that mimics what happens on the touch screen in the AR / VR space. For example, the user can use a pinching gesture to zoom in on content displayed on the touch screen. This pinching gesture on the touch screen can cause the information provided in the AR / VR space to zoom. In another example, the computing device may be rendered as a virtual book in a computer-generated 3D environment. In the AR / VR space, the pages of this book can be displayed in the AR / VR space, and a user's finger swipe across the touch screen can be interpreted as turning / flipping the pages of the virtual book. As each page is turned / flipped, in addition to seeing the content of the page change, audio feedback such as the sound of turning the pages of the book can be provided to the user.
[0205] In some embodiments, in addition to the computing device, one or more input devices (e.g., a mouse, a keyboard) can be rendered in the computer-generated 3D environment. The rendered input devices (e.g., a rendered mouse, a rendered keyboard) can be used to be rendered within the AR / VR space to control objects within the AR / VR space.
[0206] Although many embodiments have been described, it will be understood that various changes can be made without departing from the spirit and scope of the invention.
[0207] In addition, the logical flow shown in the figures is not required to be in the specific order, or sequential order, shown to achieve the desired result. In addition, other steps may be provided, steps may be removed from the described flow, other components may be added to or removed from the described system. Accordingly, other embodiments are within the scope of the following claims.
[0208] In addition to the above description, the user is provided with controls that allow the user to make a selection both when the systems, programs, devices, networks, or functions described herein are capable of collecting user information (e.g., information regarding the user's social network, social actions or activities, occupation, user preferences, or the user's current location) and when content or communications are transmitted to the user from a server. In addition, certain data can be processed in one or more ways before it is stored or used so that the user information is deleted. For example, the user's personal information can be processed so that the user information cannot be determined about the user, or the user's geographical location where location information is obtained (such as at the city, zip code, or state level) can be generalized so that the user's specific location cannot be determined. In this way, the user can control what information is collected about the user, how that information is used, and what information is provided to the user.
[0209] A computer system (e.g., a computing device) can be configured to wirelessly communicate with a network server via a communication link established with the network server using any known wireless communication technologies and protocols, including radio frequency (RF), microwave frequency (MWF), and / or infrared frequency (IRF) wireless communication technologies and protocols adapted for communication over a network.
[0210] In accordance with aspects of the present disclosure, embodiments of the various techniques described herein can be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations thereof. Embodiments can be implemented as a computer program product (e.g., an information carrier, a machine-readable storage device, a computer-readable medium, a computer program tangibly embodied on a tangible computer-readable medium) for processing by a data processing apparatus (e.g., a programmable processor, a computer, or multiple computers) or for controlling the operation of a data processing apparatus. In some embodiments, the tangible computer-readable storage medium can be configured to store instructions that, when executed, cause the processor to execute a process. A computer program, such as the computer programs described above, can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, such as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program can be deployed to be processed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a communication network.
[0211] The specific structural and functional details disclosed herein are merely representative for purposes of describing example embodiments. However, example embodiments may be embodied in many alternative forms and should not be construed as limited to only the embodiments shown herein.
[0212] The terms used in this specification are for the purpose of describing particular embodiments only and are not intended to limit the embodiments. As used in this specification, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. The terms "comprises", "comprising", "includes", and / or "including" as used in this specification specify the presence of the stated features, steps, acts, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, acts, elements, components, and / or groups thereof.
[0213] When an element is referred to as being "coupled", "connected", or "responsive" to another element, or "on" another element, it can be directly coupled, connected, or responsive to, or on, the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly coupled", "directly connected", or "directly responsive" to another element, or "directly on" another element, intervening elements are absent. The term "and / or" as used in this specification includes any and all combinations of one or more of the associated listed items.
[0214] In this specification, spatially relative terms such as "directly below," "below," "lower," "above," "upper," etc. may be used for ease of explanation to describe one element or feature in relation to another as illustrated in the figures. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is turned upside down, elements described as "below" or "directly below" other elements or features will be oriented "above" the other elements or features. Thus, the term "below" can encompass both upward and downward orientations. The device may be otherwise oriented (rotated 70 degrees or at other orientations), and the spatially relative descriptors used herein can be interpreted accordingly.
[0215] In this specification, embodiments of these concepts are described with reference to cross-sectional views that are schematic illustrations of idealized embodiments (and intermediate structures) of the embodiments. Thus, for example, variations from the shapes of the figures are to be expected as a result of manufacturing techniques and / or tolerances. Accordingly, embodiments of the described concepts should not be construed as being limited to the particular shapes of the regions illustrated herein, but should include, for example, departures in shapes due to manufacturing. Thus, the regions shown in the figures are essentially schematic in nature, and their shapes are not intended to show the actual shape of a region of the device nor are they intended to limit the scope of the embodiments.
[0216] In this specification, terms such as "first," "second," etc. may be used to describe various elements, but it will be understood that these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. Thus, a "first" element can be referred to as a "second" element without departing from the teachings of the present embodiment.
[0217] Unless otherwise defined, terms (including technical and scientific terms) used herein shall have the same meaning as commonly understood by one of ordinary skill in the art to which these concepts belong. Terms as defined in commonly used dictionaries shall be interpreted to have a meaning consistent with the context of the relevant art and / or the present specification, and it is further understood that they shall not be interpreted in an idealized or overly formal sense unless explicitly defined herein.
[0218] Some features of the described embodiments have been illustrated as described herein, but one of ordinary skill in the art will, here, envision many variations, substitutions, modifications, and equivalents. Accordingly, it should be understood that the appended claims are intended to encompass such variations and modifications as fall within the scope of the embodiments. They are presented by way of example and not limitation, and it should be understood that various changes in form and detail can be made. Any part of the apparatus and / or method described herein can be combined in any combination, except mutually exclusive combinations. The embodiments described herein can include various combinations and / or sub-combinations of the functions, components, and / or features of the different embodiments described.
Claims
1. A step of starting a recording for capturing video content, the video content including a presenter video stream, a screen cast video stream, and an annotation video stream, the presenter video stream including an image of a presenter captured by a camera, the screen cast video stream including a presentation of two or more applications by the presenter, including the step of generating, based on the video content, during the capture of the video content, a metadata record representing timing information used to synchronize at least one portion of the video content with an input received in the annotation video stream, the timing information representing when the input was received and in which of the two or more applications the input was received, a method implemented by a computer.
2. A step of generating, in response to the end of the recording, based on the metadata record, a representation of the video content, the representation including portions of the video content annotated by a user associated with the presenter video stream The method implemented by a computer according to claim 1, further including.
3. The method implemented by a computer according to claim 1 or 2, wherein the timing information corresponds to at least one position in a document associated with the video content.
4. The video content further includes a transcription video stream, and the transcription video stream During the recording of the video content, real-time speech-to-text audio data from the presenter video stream, generated as editable transcription data configured to be displayed together with the screen cast video stream, During the recording of the video content, real-time translated audio data from the presenter video stream, which is generated as text data configured to be displayed together with the screen-cast video stream and the real-time speech-to-text audio data, and The method implemented by a computer according to claim 1 or 2, comprising.
5. The transcription of the real-time speech-to-text audio data is performed by at least one speech-to-text application, and the at least one speech-to-text application is selected from a plurality of speech-to-text applications determined to be accessible by the transcription video stream, The method implemented by a computer according to claim 4, wherein the modifiable transcription data and the text data are stored in the metadata record according to a time stamp and configured to be searchable.
6. The input includes annotation input associated with the annotation video stream, and the annotation input includes video marker data and teleprompter data generated by a user associated with the presenter video stream. The method implemented by a computer according to claim 1 or 2.
7. The presenter video stream, the screen-cast video stream, and the annotation video stream are configured to be switched on and off during the recording, and the switching on and off triggers the display or deletion from the display of the presenter video stream, the screen-cast video stream, or the annotation video stream. The method implemented by a computer according to claim 1 or 2.
8. When executed by at least one processor, Starting a recording to capture video content, the video content including a presenter video stream, a screen-cast video stream, a transcription video stream, and an annotation video stream, starting Generate a metadata record representing timing information used to synchronize at least one portion of the video content with the input received in the annotation video stream during capture of the video content based on the video content; including instructions configured to cause a computing system to execute instructions including: The presenter video stream includes images of the presenter captured by a camera; The screen cast video stream includes presentations of two or more applications by the presenter; The timing information is a computer program representing when the input was received and in which of the two or more applications the input was received;
9. The instructions are: Generating, in response to the end of the recording, a representation of the video content based on the metadata record, the representation including portions of the video content annotated by a user associated with the presenter video stream; The computer program according to claim 8, further comprising:
10. The computer program according to claim 8 or 9, wherein the timing information corresponds to at least one position in a document associated with the video content;
11. The transcription video stream is: Real-time speech-to-text audio data from the presenter video stream, generated as text data configured to be displayed together with the screen cast video stream during the recording of the video content; Real-time translated audio data from the presenter video stream, generated as text data configured to be displayed together with the screen cast video stream and the real-time speech-to-text audio data during the recording of the video content; The computer program according to claim 8 or 9, including:
12. The real-time speech-to-text audio data is generated as editable transcription data configured to be displayed together with the screen-cast video stream during the recording of the video content. The transcription of the real-time speech-to-text audio data is performed by at least one speech-to-text application, and the at least one speech-to-text application is selected from a plurality of speech-to-text applications determined to be accessible by the transcription video stream. The computer program according to claim 11, wherein the editable transcription data and the text data are stored in the metadata record according to a time stamp and configured to be searchable.
13. The input includes annotation input associated with the annotation video stream, and the annotation input includes video marker data and teleprompter data generated by a user associated with the presenter video stream. The computer program according to claim 8 or 9.
14. The presenter video stream, the screen-cast video stream, the transcription video stream, and the annotation video stream are configured to be switched on and off during the recording, and the switching on and off triggers the display or deletion from the display of the presenter video stream, the screen-cast video stream, the transcription video stream, or the annotation video stream. The computer program according to claim 8 or 9.
Citation Information
Patent Citations
Video and handout PPT and voice content accurate matching method and system
CN107920280A
Multimedia data objects for real-time slide presentations, and systems and methods for recording and viewing multimedia data objects
JP2005533398A
Method for creating annotated transcript of presentation, information processing system, and computer program
JP2008282397A
Technology for generating visual compositions for multimedia conference events
JP2011514043A
System, Method and Computer Program Product for a Universal Call Capture Device
US20150003595A1