Auto-completion for content expressed in video data
By automatically generating text from spoken content in videos, the problem of low user interaction efficiency in existing video commenting systems is solved, and the accuracy of comments and the utilization of computing resources are improved.
Patent Information
- Application Number
- CN202080029806.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-19
- Filing Date
- 2020-03-30
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2040-03-30
AI Technical Summary
Existing video commenting systems have low user interaction efficiency. Users need to manually transcribe video content and replay it multiple times to generate comments, resulting in wasted computing resources and inaccurate comments.
The system automatically generates text for the spoken content in the video and inserts it into the text input field through a computing device. The system identifies the spoken content contained in the video based on keywords and generates text, reducing user manual interaction.
It improves user engagement and comment accuracy, optimizes the use efficiency of computing resources, reduces manual operations and unintentional input, and reduces the consumption of computing resources.
Smart Images

Figure CN113748425B_ABST
Abstract
Description
Background Art
[0001] Commenting on videos is becoming popular and ubiquitous across many social, educational, and entertainment platforms. Many video-based commenters reference the video content to contextualize and specify their messages. Commenters can reference visual entities or specific sound clips in a variety of ways. For example, users can reference someone's voice or quote at a specific time, or provide a timestamp to allow viewers to play a video starting at a specific point in time. This feature plays a key role in influencing user engagement and, ultimately, user traffic and retention.
[0002] Although existing video-based platforms provide features that allow users to provide comments, most user interfaces that exist today are simplistic in nature and do not provide tools for optimizing the user experience. Many sites follow a traditional model that includes a video display area, a text entry field, and a comment section. Users are often required to manually type text in the text entry field, which is both cumbersome and inefficient in terms of user productivity and computing resources. This inefficiency is exacerbated when complex tasks are to be performed. For example, if a user wants to provide a quote from the spoken content of a video, the user is required to play the video step by step and manually transcribe the spoken content. This traditional approach can result in inaccuracies regarding the comments. In addition, this traditional approach can result in inefficient use of computing resources because the commentator may be required to replay parts of the video multiple times to transcribe the content. This problem can have a negative impact on many performance metrics of the site.
[0003] It is in response to these and other technical challenges that the present disclosure is proposed. Summary of the Invention
[0004] The technology disclosed herein provides an improvement over existing systems by enabling a computing device to perform an auto-completion process that generates the text of the spoken content of a video and inserts it into a text input field. By providing the quoted content in the text input field, the system can alleviate the user from having to perform the tedious process of listening to the spoken content of the video and manually typing the spoken content into the computing device. In some configurations, the system can receive one or more keywords from the user input and identify the spoken content in the video that contains the keywords. The system can provide the text of the spoken content based on the relevance level and fill in one or more input fields with the text of the spoken content.
[0005] The technology described herein provides many benefits. For example, by providing an automatic completion process for generating spoken content from a video and inserting the spoken content into a text input field, the technology disclosed herein can increase user participation from both a personal and a community perspective. Specifically, by providing a mechanism that automates the process of generating input text containing spoken content, the system can enable users to publish more accurate statements in the comment section of a video platform while minimizing the amount of manual interaction required to generate comments. From a community perspective, user participation can also be optimized. Some usage data shows that comments containing spoken content from a video (also referred to as "quotes" herein) are more likely to receive responses than comments that do not include spoken content from a video. The system described herein not only helps users provide more accurate comments containing spoken content, but also by providing suggested line completion content to the user's input, the system can encourage users to provide such information when they may not otherwise be provided with quotes. This feature can encourage certain types of user activities, which ultimately enhances user participation in video-based systems.
[0006] For purposes of this description, the term "spoken" content may include any type of language, melody, or sound that can be produced by an entity or person. Spoken content may be interpreted based on any form of input received from an input device (e.g., a microphone) or any type of sound that can be interpreted based on audio data to generate any type of notation, including symbols, text, images, code, or any other data that can represent sound.
[0007] The technology described herein can lead to more efficient use of computing systems. In particular, by automating the generation of input strings for quotations with video content, user interaction with computing devices can be improved. The technology disclosed herein can eliminate multiple manual steps that require additional computing resources. For example, for someone transcribing audio content from a video stream, the user may have to play the video multiple times to ensure they can accurately capture the content. This results in the computing device retrieving video data and using multiple computing resources (including memory resources and processing resources) to play and replay the video and corresponding audio while transcribing the content. The elimination of these manual steps leads to more efficient use of computing resources (e.g., memory usage, network usage, and processing resources) because it eliminates the need for a person to retrieve, render both audio and video data, and view the rendered data. Additionally, the reduction in manual data entry and the improvement in user interaction between people and computers can bring several other benefits. For example, by reducing the need for manual entry, unintentional entry and human error can be reduced. Fewer manual interactions and the reduction in unintentional entry can avoid the consumption of computing resources that might be used to correct or retype data created by unintentional entry.
[0008] By reading the following detailed description and viewing the associated drawings, other features and technical advantages in addition to those explicitly described above will also be apparent. This summary is provided to introduce, in a simplified form, a selection of concepts further described below in the detailed description. This summary is not intended to identify the key or important features of the claimed subject matter, nor is it intended to be used to assist in determining the scope of the claimed subject matter. For example, the term "technology" may refer to (multiple) systems, (multiple) methods, computer-readable instructions, (multiple) modules, algorithms, hardware logic, and / or (multiple) operations as allowed by the context described above and the entire document. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The detailed description is described with reference to the accompanying drawings. In the drawings, the left-most digit(s) of a reference number identifies the drawing in which the reference number first appears. The same reference numbers in different drawings indicate similar or identical items. References to individual items of a plurality of items may use reference numbers with letters of the alphabetical sequence to refer to each individual item. General references to items may use specific reference numbers without the alphabetical sequence.
[0010] Figure 1A An example scenario is shown in which the system may be used in an auto-completion process for providing spoken content from a video.
[0011] Figure 1B The steps of an auto-completion process for providing spoken content from a video are shown.
[0012] Figure 1C Aspects of a playback process for rendering an audio output in response to receiving a selection of a link within a text portion are shown.
[0013] Figure 1D Aspects of an auto-completion process utilizing a graphical menu to obtain a caption are shown.
[0014] Figure 1E Additional aspects of the auto-completion process for obtaining captions using graphical menus or other types of input are shown.
[0015] Figure 2 is a block diagram illustrating components of a process for generating text data from video data.
[0016] Figure 3 An example of a user interface that displays related text portions based on text data having a timeline associated with multiple text portions is shown.
[0017] Figure 4An example graphical user interface is shown having a menu of ordered options for allowing a user to select spoken content.
[0018] Figure 5 An example graphical user interface is shown having a filtered menu of options for selecting spoken content.
[0019] Figure 6A An example of a user interface with input text indicating a specified time for a video is shown.
[0020] Figure 6B An example of a text portion being selected based on input text indicating a specified time of a video is shown.
[0021] Figure 7A An example of a user interface with input text indicating a specified time and entity of a video is shown.
[0022] Figure 7B An example of a text portion being selected based on input text indicating a specified time and entity of a video is shown.
[0023] Figure 8A A user interface is shown that displays a portion of text selected based on characteristics of an audio file.
[0024] Figure 8B A user interface is shown that displays a portion of text formatted based on characteristics of an audio file.
[0025] Figure 9A Shown are forms of tokens that may be generated based on characteristics of user input captured by a computer's microphone.
[0026] Figure 9B An auto-completion process for generating output tokens based on analysis of audio content having a threshold level of correlation with tokens generated from user input is shown.
[0027] Figure 9C One example of how the generated tokens may be used to fill in one or more portions of a document is shown.
[0028] Figure 9D An example of how audio content associated with generated tokens may be rendered is shown.
[0029] Figure 10 is a flow chart illustrating aspects of a routine for computationally efficiently generating spoken content for a video.
[0030] Figure 11 is a computing system diagram showing aspects of an illustrative operating environment for the techniques disclosed herein.
[0031] Figure 12 is a computing architecture diagram illustrating aspects of the configuration and operation of a computing device that may implement aspects of the techniques disclosed herein. DETAILED DESCRIPTION
[0032] Figure 1A and Figure 1B An example scenario is shown in which the system can be used in an auto-completion process for providing quotes from a video. Generally, the system can analyze user input to identify quotes expressed in a video. To help users create comments that include quotes from the video, the quotes can be automatically populated into an input field. Rather than requiring users to manually transcribe the audio content of a video, the system can receive one or more keywords from the user input and identify quotes from the video that contain the keywords. The quotes can then be populated into the input field.
[0033] like Figure 1A , system 100 can cause display of user interface 130 including video display area 140, text entry field 150, and comment section 160. System 100 can receive video data 110 having video content 111 and audio content 112. The system can also receive text data 113 associated with video data 110. In one illustrative example, text data 113 can be in the form of closed caption text and have a plurality of different phrases associated with the timeline of video content 111 and audio content 112. System 100 can process video content 111 to generate rendered video content 116 for display within video display area 140. Additionally, system 100 can process audio content 112 to generate a rendering of audio content 112 via an endpoint device (e.g., a speaker).
[0034] The user interface 130 can be configured to receive input text 151 at a text entry field 150. The input text 151 includes at least one keyword 152. In some embodiments, the keyword 152 can be distinguished from other words in the input text 151 by using special characters (e.g., single quotes or double quotes). In this example, the keyword 152 GREATEST is identified because it immediately follows the first quote of the phrase within double quotes.
[0035] The system 100 can then identify portions 115 of the text data 113 based on the keywords 152. Figure 1B, the system 100 can then insert a portion 115 of the text data 113 having at least one keyword 152 into the text entry field 150. In this example, the portion of the text data, "PLAY IN HISTORY," is identified in the text data 113 based on the keyword 152, "GREATEST." If the user wishes to continue manually typing the remainder of the quotation, the user can press a predetermined key (e.g., the ESC key), and the system will remove the portion 115 of the text data 113.
[0036] In some configurations, the user interface 130 may include an interface element 131 for receiving input. The user interface 130 may also be configured to display a portion 115 of the text data 113 and a user input 151 in the comment section 160 in response to receiving input at the interface element 131. For illustrative purposes, a demarcated portion of text (e.g., a sentence with punctuation) may be referred to herein as a "portion 115" of the text data 113, a "segment 115" of the text data 113, or a "text portion 115." In some configurations, the portion 115 of the text data 113 may be inserted into a generated comment 143 or other graphical element. The comment 143 may be configured with a link that calls for playback of audio data associated with the portion 115 of the text data 113. The link may cause the audio content 112 to be played back at specific time intervals. In some embodiments, the time intervals may be derived from the timestamp data 121 associated with the text data 113. The timestamp data may include a specific point in time, or the timestamp data may indicate an interval in which the system 100 may generate the audio output 126 from the speaker 125 of the system 100 .
[0037] Figure 1C 1. Aspects of a playback process for rendering an audio output in response to receiving a selection of a link within a text portion are shown. In this illustrative example, when a user selects a generated comment 143, the system can render an audio output 126 from a speaker 125 of the system 100. In some configurations, the playback of the audio content 112 can be controlled using timestamp data 121. The playback can be based on the timestamp data 121.
[0038] In some configurations, the auto-completion process may be based on one or more user inputs. Figure 1D 1 shows aspects of an automatic completion process for obtaining captions using a graphical menu. In this example, after the user includes at least one keyword 152, the user can take one or more actions, such as selecting a graphical element 122. Figure 1E, in response to selection of the graphical element 122, the system 101 can obtain the portion 115 of the text data 113 to insert into the input text field 150. Such an embodiment is optional, as it will be appreciated that the system 100 can automatically fill in the input text field in response to receiving at least one keyword or any other text that can be identified using the portion of the text data 113. In other embodiments, as opposed to displaying the graphical element 122, the system can also receive a predetermined input, such as a special key or a special key sequence (e.g., shift-control-Q), to invoke the system 100 to automatically fill in the input text field using the portion 115 of the text data 113.
[0039] In some configurations, the system 100 may generate the text data 113 by analyzing the video data 110 . Figure 2 An example of a process for generating text data 113 is shown. In this example, processor 101 may analyze audio content 112 associated with video data 110 to generate text data 113. For example, if audio content 112 contains dialogue, processor 101 may convert the dialogue into a plurality of phrases 114. Any suitable technique for transcribing audio signals may be utilized.
[0040] In this example, audio content 112 contains conversations between players of a video game. One or more criteria can be used to parse phrases 114 of text data 113 into sentences. Sentences can be generated based on phrases 114 transcribed from audio content 112, wherein sentences can include punctuation marks and other identifiers for delimiting phrases. Therefore, the criteria can include universal grammatical rules for a specific language for identifying where punctuation marks or other identifiers can be placed to identify a specific quote (e.g., to identify the beginning and end of a quote). By defining sentences, the beginning and end of a specific quote can be utilized to identify the portion of text data that should be selected for insertion into text entry field 150. In some configurations, system 100 can select sentences with keywords 152 for insertion into text entry field 150, which keywords 152 are provided as part of the input.
[0041] Comments section 160 is also referred to herein as "text field 160," "text section 160," or "notation section 160." Comments section 160 may include any portion of a user interface, including text or any other type of notation associated with video content or audio content. For example, comments section 160 may be part of a word processing document, a OneNote file, a spreadsheet, a blog, or any other form of media or data that enables a computer to render text in conjunction with the rendering of a video.
[0042] In other embodiments, the sentence or any other delimited portion of text data 113 can be identified by the characteristic of voice.For example, system 100 can analyze at least one of the pitch, inflection point (inflection point) or the volume of the audio content to detect the audio content.If system 100 detects the change about the threshold level of voice, then the system can identify the starting point or ending point of a sentence or the delimited portion.Similarly, if there is a change (for example, inflection of volume or any type) about the threshold level of any other type of characteristic, then the system can identify the starting point or ending point of a sentence or the delimited portion.This technology can help identification to be inserted into the quotation in the text input field 150.
[0043] Except text data 113 is parsed into the text portion of sentence or any other type of delimitation, system 100 can also identify the entity associated with each sentence.For example, system 100 can analyze at least one of the pitch, inflection point or the volume of the audio content to detect the audio content.Based on the change of the threshold level of at least one of the pitch, inflection point or the volume, system 100 can identify the entity associated with the sentence or the delimited text portion, for example, specific people.Then, system 100 can insert identifier 117 into text data 113, for specific sentence or any other text portion of delimitation.
[0044] The system 100 can also identify specific identifier names by interpreting the audio content. For example, if a name is repeated multiple times within a specific context, the system 100 can associate the name with a specific portion of text. The system can also identify specific speech by detecting predetermined tones, pitches, inflection characteristics, etc. The system 100 can also associate specific speech with a name and associate the name with a portion of text associated with speech having specific characteristics.
[0045] In some configurations, the user's intent may be used to identify keywords for an input entry to be analyzed for text data. The user's intent may be inferred from one or more characters of the text input. For example, a single quote character or a double quote character may be used to identify the user's intent. Figure 3 In the example shown in , the input entry includes the following text: Ilike the quote, "Greatest, where the entry includes a double quote character only before the word Greatest. In this example, the double quotes indicate that the following words are part of the spoken content that the user wishes to include in their review. Based on this type of input, the system can search for keywords that immediately follow a double quote character, a single quote character, and so on.
[0046] This example is provided for illustrative purposes and should not be construed as limiting. It will be appreciated that other characters or other visual indicators can hint at user intent to identify keywords. For example, formatted text (e.g., bold text, italic text, or other types of text formatting) can be used to identify user intent. In one illustrative example, if a user text entry includes one or two bold words, these words can be used to generate a search query to identify the spoken content of a video.
[0047] In some configurations, the system 100 may utilize time stamps associated with text portions to identify the most relevant text portions for quoting. To illustrate aspects of this feature, Figure 3 An example set of text data 113 is shown. Such a data set can be generated by processor 101 by recording the timestamp of each text portion transcribed from the audio content. In this particular example, text data 113 includes three sentences that include the keyword "greatest," and the system records a time stamp in each sentence with the keyword, for example, at 3:20, 7:50, and 9:03, respectively.
[0048] In some embodiments, the portion of text selected for insertion into the text entry field 150 may be selected based on a selected time marker 301 relative to a time marker of a particular portion of the text data 113. Figure 3 In the example shown in FIG. 3 , the system selects the first sentence (“Greatest achievement in history!”) because its timing is closer to the selected time marker 301 than the timings of the other sentences (“Greatest player ever!” and “I am the Greatest!”).
[0049] The selected time marker 301 can be based on a number of factors. In one illustrative example, the selected time marker 301 can be based on a time indicated by a user input. For example, if the user input includes the text "I like the player's quote at time marker 3:20, Greatest," the system 100 can designate 3:20 as the selected time marker and then select the portion of text that is closest to the selected time marker and also includes the specific keyword provided by the user input (e.g., "greatest"). In this way, even if multiple sentences within the text data 113 include keywords from the user input, the system 100 can be more accurate in selecting the portion of text.
[0050] In another example, the selected time marker 301 may be based on the current time associated with the rendered content displayed in the video display area. Figure 3 , the position of the playback cursor 305 represents the current time of the video rendering. Thus, if a user is watching a video and pauses the video at a specific time (e.g., at the 4:00 mark), the system can select the portion of text containing the keyword closest to that specific time. The selected time marker 301 can also be based on a combination of a number of different factors, including the current time of the video player and the user. In this way, if the user input is not completely accurate with respect to the specified time, the system can analyze the time specified by the input in conjunction with the player's current time and determine the selected time marker 301.
[0051] In some configurations, the system can select, sort, and display multiple portions of text for the user. For example, Figure 4 As shown in , the system can identify more than one portion of text that contains the keyword provided in the input. In this scenario, the system can generate a user interface 130 including a menu 401 that shows each sentence with the keyword. In some configurations, the sentences with the keyword can be sorted based on relevance level. In this example, since the first sentence (3:20) is closer to the selected time marker 301 than the second sentence (7:50), the first sentence can be placed first within the menu 401. The menu 401 can also be configured to receive user input. In response to user input indicating selection of a particular sentence or portion of text, the system can populate the text entry field 150 with the selected portion.
[0052] This example is provided for illustrative purposes and should not be construed as limiting. It will be appreciated that other variations of the technology disclosed herein may be within the scope of the present disclosure. For example, although the selected time marker 301 is indicated as a specific point in time, it will be appreciated that the selected time marker 301 may include a time interval. Thus, the portion of text closest to the selected range or the portion of text closest to a point within the time interval may be selected to fill in the text entry field 150.
[0053] In some embodiments, the menu 401 options may be sorted according to time stamps associated with the text portions. Figure 4 In the example shown in , each sentence is sorted based on the proximity of the associated time relative to the selected time marker 301 .
[0054] In some embodiments, the system may filter the different menu 401 options based on one or more factors. Figure 5An example of such a feature is shown. In this example, the system can analyze text data 113 to determine time tags for portions of the text data that contain at least one keyword. The system can then determine whether the associated time tag for each text portion is within a threshold duration of the current time tag 301. The system can then insert each text portion with an associated time tag within the threshold duration into menu 401.
[0055] In one illustrative example, the system can insert the selected portion of text data in menu 401 in response to determining that the time stamp for the selected portion of text data is within a predetermined threshold of the current time stamp. Thus, the system can filter certain text portions from the ordered list of menu 401 options even if those text portions have a threshold level of relevance and / or common keywords with the input text.
[0056] While the examples described herein illustrate embodiments in which text portions are selected based on keywords, it will be appreciated that other techniques for identifying relevant text portions may be utilized. For example, in some embodiments, the system may select text portions based on relevance levels. Relevance levels may be based on many different factors, which may include context interpreted by user input. Figure 6A 、 Figure 6B 、 Figure 7A and Figure 7B An example of such an embodiment is shown.
[0057] In some embodiments, the system can select portions of the text data 113 based on user-defined time markers. For example, consider a scenario in which a user provides the following input text: "I like the quote at time marker 3:30 when he said, "Greatest___." In this example, the system can select sentences that have the word "Greatest" and have a time marker that is closest to the time marker indicated in the input. Figure 6A and Figure 6B Another example of this feature is shown.
[0058] exist Figure 6A In the example shown in , the input text includes "@7:50". Based on analysis of this input, the system may select the text portion "Greatest play ever!" because this portion has an associated time that is equal to, or within a threshold duration relative to, the time indicated in the input text. Figure 6BAs shown in , the system selects the portion of text data at the time mark indicated in the input text. The selected portion of text is then inserted into the text entry field 150.
[0059] In other embodiments, the system may select one or more portions of text based on a combination of indicators provided in the input text. Figure 7A and Figure 7B An example of how to provide multiple indicators within the input text is shown in FIG. Figure 7A , the input text includes “I LIKE WHAT PLAYER 1 SAID AT 7:50”. Based on analysis of the input, the system can select portions of text based on the time indicated in the input text and entities identified in the input text.
[0060] In this example, if Figure 7A As shown in , the selected portion may include the quote from Player 1 at 7:49 (GREATEST PLAY EVER!) because the time of the text portion is within the threshold duration from the time indicated in the input text and because the portion is associated with the entity indicated in the input text. Other text portions may be excluded based on the fact that they are outside the threshold duration from the time indicated in the input text or they are associated with entities not indicated in the input text. Figure 7B As shown in , the system selects the portion of the text data at the time indicated in the input text and inserts it into the text entry field 150 .
[0061] These examples provide for illustrative purposes and should not be interpreted as restrictive. It will be appreciated that the text portion can be selected based on other factors. In another example, the word selected for automatic completion of typing can be determined by the characters of the voice associated with the part of the text data. In a specific explanation, the word selected for automatic completion of typing can be based on the inflection, tone or volume of the voice associated with this text portion.
[0062] Figure 8AShown is the example of this embodiment.Here, system can analyze audio content to detect at least one in the tone, inflection point or the volume of the part of audio content.Then, system can change the starting point or the ending point in the text data based on the threshold value about at least one in tone, inflection point or the volume.The starting point 801 and the ending point 802 determined can define the boundary of the part of text data.In this example, because tone and volume exceed threshold value after a certain time point, so system can select the text that is associated with the characteristic (for example, tone and / or volume) that presented before tone and / or volume change, and utilize the characteristic that presents after the change to filter text.Such embodiment can be useful when text data 113 may not comprise punctuation marks or other text delimiters therein.Therefore, if text data 113 comprises a long string of text, or if punctuation marks are incorrect, then system can select context-dependent text based on the characteristic of audio content.
[0063] In some configurations, the system can select a style, arrangement, appearance, or punctuation for the selected text. Such characteristics of the text can be based on an analysis of the audio content. For example, if the system determines that a portion of the text has a raised voice, the system can generate a visual indicator to indicate the raised voice. Figure 8B An example of this feature is shown. As shown, as one or more characteristics of speech change, the system can automatically format a selected portion of text inserted into the text entry field 150. In this example, given that the rate of change in pitch and / or the rate of change relative to volume exceeds a threshold, the system formats the word "history" to emphasize the associated text.
[0064] This example is provided for illustrative purposes and should not be construed as limiting. It will be appreciated that the display of the text in the text entry field 150 can be formatted using other characteristics, pitch or volume of the voice or sound associated with the text portion. It will also be appreciated that the layout of any displayed text can be selected using a change in the threshold level of a characteristic and / or the rate of change (shown as "slope" in the accompanying drawings). The selected layout can include any technology that arranges the text to make it more prominent, legible, readable and / or attractive when displayed. The arrangement of text involves selecting a font, font size, line length, line spacing and letter spacing, as well as adjusting the spacing between pairs of letters. The term layout also applies to the style, arrangement and appearance of the letters, numbers and symbols created by this process.
[0065] While the examples disclosed herein illustrate embodiments involving text entry, the technology disclosed herein can identify any type of content related to a video based on any user input indicating specific content. In another illustrative example, user input indicating a melody can be utilized to identify specific audio content related to a video. Figures 9A-9D Such an example is shown in .
[0066] Now refer to Figure 9A , shows an example scenario in which a user 901 provides input 903 such as a melody. In this example, the user 901 provides (e.g., chants, sings, chats, speaks, hums, talks) a melody to an input device such as a microphone 902. In some configurations, the melody can include a vocal input comprising a series of tones. The input 903 can include any audible sound captured by the microphone or text received from the input device. The input 903 can be received in association with a character input or a predetermined key input or a selection of a menu item. In response to the input 903, the system 100 analyzes the melody and determines a sequence of notes 181 that defines the melody. This process can utilize any suitable technology to transcribe the user's speech into any type of notation.
[0067] like Figure 9B As shown in FIG, system 100 can identify audio segments of audio content 112 that have a threshold level of relevance to a note sequence 181. Any suitable technique for processing audio content 112 to identify specific audio content based on a note sequence or melody provided by a user can be utilized. In some configurations, system 100 can initiate one or more processes that compare the note sequence and / or the user's melody to different portions of audio content 112. Confidence scores can be generated for the different portions of audio content 112. Any portion of audio content 112 with a confidence score above a threshold can be identified as relevant audio content. The system can then generate an output 182 defining the portion of audio content 112 that has the threshold level of relevance to the user's melody and / or note sequence 181. Output 182 can be in the form of a graphical representation of the portion of audio content 112 and can be in any format that conveys or models a melody, a series of notes, a series of pitches, a series of pitch changes, etc.
[0068] Next, if Figure 9C, the user input may cause the computing device 100 to generate an entry 183 within the notation portion 160 of the user interface 130. It will be appreciated that the user input may be based on a selection of a user interface element 131 or any other type of input. For example, the input may include a voice command, a gesture, or any other type of user input that provides an indication of the user's intent to add the entry 183 within the notation portion 160 or any other portion of the user interface. It will be appreciated that the entry 183 may also include a link associated with a portion of the audio content 112 that has a threshold level of relevance to the user's melody and / or note sequence 181. Thus, as Figure 9D As shown in , in response to user input (eg, selection of item 183 ), system 100 may render audio output 126 portion of audio content 112 that has a threshold level of relevance to the user's melody and / or note sequence 181 .
[0069] Figure 10 1 is a diagram illustrating aspects of a routine 1000 for computationally efficiently generating and managing portions of text. It will be understood by those skilled in the art that the operations of the methods disclosed herein are not necessarily presented in any particular order, and that it is possible and contemplated to perform some or all of the operations in an alternative order. For ease of description and illustration, the operations have been presented in the order shown. Operations may be added, omitted, performed together, and / or performed simultaneously without departing from the scope of the appended claims.
[0070] It should also be understood that the illustrated method can be terminated at any time and need not be performed in its entirety. As defined herein, some or all of the operations of the method and / or substantially equivalent operations may be performed by executing computer-readable instructions contained on a computer storage medium. The term "computer-readable instructions" and variations thereof as used in the specification and claims are used broadly herein to include routines, applications, application modules, program modules, programs, components, data structures, algorithms, and the like. The computer-readable instructions may be implemented on a variety of system configurations, including single-processor or multi-processor systems, minicomputers, mainframe computers, personal computers, handheld computing devices, microprocessor-based programmable consumer electronics, combinations thereof, and the like.
[0071] It should be recognized, therefore, that the logical operations described herein are implemented as: (1) a sequence of computer-implemented actions or program modules running on a computing system such as those described herein; and / or (2) interconnected machine logic circuits or circuit modules within a computing system. The manner of implementation is a matter of choice depending on the performance and other requirements of the computing system. Thus, the logical operations may be implemented in software, firmware, dedicated digital logic units, and any combination thereof.
[0072] Additionally, the example presentation UI described above may be implemented in association with Figure 10 For example, the various devices and / or modules described herein may generate, send, receive, and / or display data associated with the content of a video (e.g., real-time content, broadcasted events, recorded content, etc.) and / or present a UI that includes a rendering of one or more participants of a remote computing device, avatar, channel, chat session, video stream, image, virtual object, and / or application associated with the video.
[0073] The routine 1000 begins at operation 1002, where the system may cause display of a user interface having a video display area and a text entry field. Figure 1A An example of a user interface is shown in FIG. In some configurations, the user interface may also include a comment section. The user interface may be displayed on a client device such as a tablet computer, a mobile phone, a desktop computer, or the like.
[0074] Next, at operation 1004, the system may receive input text at a text entry field. In some configurations, the text input may include keywords or phrases. The input text may be received by an input device such as a touch screen, a keyboard, or any other suitable input device. The input text may also be received by a gesture captured by a camera of the device, or by capturing an audio signal by a speaker of the device.
[0075] Next, at operation 1006, the system can analyze the text data to determine whether the input text and the part of the text data have a threshold level of relevance. In some configurations, the threshold level of relevance can be based on the common keywords between the input text and the part of the text data. The threshold level of relevance can also be based on a selected time mark or a predetermined time line. For example, if the time within the threshold of the part of the text data and the selected time mark is associated, then the text portion can be selected to be used for the text entry field. Alternatively, if the text portion is within the predetermined time line, then the text portion can be selected to be used for the text entry field. The selected time mark can be based on the current position of the video player, the time mark indicated in the input text, or the time mark otherwise indicated by the user. Text data (e.g., closed caption text) can be received by the system, or text data can be generated by the system by analyzing the audio content associated with the video data.
[0076] Next, at operation 1008, the system can analyze the video data to determine parameters of the selected text portion. For example, the pitch or volume of an audio track (e.g., audio content) associated with the text portion can be used to select specific words of the text portion to be inserted into the text entry field. Such features can be used when identifiers such as the speaker's name are not provided with the text data. In another example, the pitch or volume of an audio track associated with the text portion can be used to format the text to highlight certain words or phrases.
[0077] Next, at operation 1010, the system may fill the selected portion of text into the text entry field. In some configurations, a portion of the selected portion of text may be added to the existing text within the text entry field to serve as an auto-complete (e.g., line completion) feature. For example, if the user types an initial keyword, and the initial keyword is the first word of the selected text, the system may maintain the initial keyword typed by the user and only add the portion of the selected text that does not include the initial keyword.
[0078] Next, at operation 1012, the system can populate the comment section of the user interface with the selected text portion. In some configurations, the system can populate the comment section in response to user input receiving the selected text portion. The user input can be a voice command, a gesture captured by a camera, or any other suitable interaction with the computer. At operation 1012, the text portion displayed within the comment section can be formatted based on the analysis performed at operation 1008.
[0079] It should be appreciated that the subject matter described above can be implemented as a computer-controlled device, a computer process, a computing system, or as an article of manufacture such as a computer-readable storage medium. The operations of the example methods are shown in separate blocks and are summarized with reference to those blocks. The methods are shown as a logical flow of blocks, each of which can represent one or more operations that can be implemented in hardware, software, or a combination thereof. In the context of software, an operation represents a computer-executable instruction stored on one or more computer-readable media that, when executed by one or more processors, enables the one or more processors to perform the recited operation.
[0080] Generally, computer-executable instructions include routines, programs, objects, modules, components, data structures, etc. that perform specific functions or implement specific abstract data types. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations may be performed in any order, combined in any order, subdivided into multiple sub-operations, and / or performed in parallel to implement the described processes. The described processes may be performed by resources associated with one or more devices (e.g., one or more internal or external CPUs or GPUs) and / or one or more hardware logic units (e.g., a field programmable gate array ("FPGA"), a digital signal processor ("DSP"), or other type of accelerator).
[0081] All of the methods and processes described above can be embodied in software code modules executed by one or more general-purpose computers or processors and fully automated via the software code modules. The code modules can be stored in any type of computer-readable storage medium or other computer storage device, such as those described below. Some or all of the methods can alternatively be embodied in specialized computer hardware, such as those described below.
[0082] Any conventional description, element or block in the flowcharts described herein and / or depicted in the accompanying drawings should be understood to potentially represent a module, segment or portion of code, which includes one or more executable instructions for implementing the specific logical functions or elements in the routine. Alternative implementations are included within the scope of the examples described herein, wherein elements or functions may be deleted from those shown or discussed or performed in a different order, including substantially simultaneously or in reverse order, depending on the functionality involved as understood by those skilled in the art.
[0083] Figure 11 1004 .
[0084] As shown, a communication session 1104 (where N is a number having a value of two or greater) can be implemented between a plurality of client computing devices 1106(1) through 1106(N) associated with or as part of the system 1102. The client computing devices 1106(1) through 1106(N) enable users (also referred to as individuals) to participate in the communication session 1104. Although this embodiment shows a communication session 1104, it will be appreciated that a communication session 1104 is not required for every embodiment disclosed herein. It will be appreciated that a video stream can be uploaded by each client 1106 and comments can be provided by each client 1106. It will be appreciated that any client 1106 can also receive video data and audio data from the server module 1130.
[0085] In this example, a communication session 1104 is hosted by a system 1102 on one or more networks 1108. That is, the system 1102 can provide services that enable users of client computing devices 1106(1) to 1106(N) to participate in the communication session 1104 (e.g., via real-time viewing and / or recorded viewing). Thus, "participants" of the communication session 1104 can include users and / or client computing devices (e.g., multiple users can participate in a communication session in a room via the use of a single client computing device), each of which can communicate with other participants. Alternatively, the communication session 1104 can be hosted by one of the client computing devices 1106(1) to 1106(N) using peer-to-peer technology. The system 1102 can also host chat conversations and other team collaboration functionality (e.g., as part of an application suite).
[0086] In some implementations, such chat conversations and other team collaboration functions are considered to be external communication sessions distinct from the communication session 1104. The computerized agent used to collect participant data in the communication session 1104 may be able to link to such an external communication session. Thus, the computerized agent may receive information such as date, time, session-specific information, etc. that enables connectivity to such an external communication session. In one example, a chat conversation may be conducted based on the communication session 1104. Additionally, the system 1102 may host the communication session 1104, which includes at least a plurality of participants that are co-located at a meeting location (e.g., a conference room or auditorium) or located at different locations. In the examples described herein, some embodiments may not utilize the communication session 1104. In some embodiments, a video may be uploaded from at least one of the client computing devices (e.g., 1106(1), 1106(2)) to the server module 1130. When the video content is uploaded to the server module 1130, any client computing device may access the uploaded video content and display the video content within a user interface such as those described above.
[0087] In the examples described herein, client computing devices 1106(1) to 1106(N) participating in a communication session 1104 are configured to receive and render communication data for display on a user interface of a display screen. The communication data may include a collection of various instances or streams of real-time content and / or recorded content. The collection of various instances or streams of real-time content and / or recorded content may be provided by one or more cameras (e.g., video cameras). For example, a single stream of real-time content or recorded content may include media data associated with a video feed provided by a video camera (e.g., audio data and visual data capturing the appearance and voice of users participating in the communication session). In some implementations, the video feed may include such audio data and visual data, one or more still images, and / or one or more avatars. The one or more still images may also include one or more avatars.
[0088] Another example of a single stream of real-time content or recorded content may include media data including avatars of users participating in a communication session and audio data capturing the users' voices. Yet another example of a single stream of real-time content or recorded content may include media data including a file displayed on a display screen and audio data capturing the users' voices. Thus, the various streams of real-time content or recorded content within the communication data enable facilitating remote conferencing between a group of people and sharing content within a group of people. In some implementations, the various streams of real-time content or recorded content within the communication data may originate from multiple co-located cameras located in a space (e.g., a room) for recording or streaming a presentation that includes one or more personal presentations and one or more personal consumption of the presented content.
[0089] Participants or attendees can view the content of the communication session 1104 in real time as the activity occurs, or alternatively, at a later time after the activity occurs via a recording. In the examples described herein, client computing devices 1106(1) to 1106(N) participating in the communication session 1104 are configured to receive and render communication data for display on a user interface of a display screen. The communication data may include a collection of various instances or streams of real-time content and / or recorded content. For example, a single stream of content may include media data associated with a video feed (e.g., audio data and visual data capturing the appearance and voice of users participating in the communication session). Another example of a single stream of content may include media data including avatars of users participating in the conference session and audio data capturing the voice of the users. Yet another example of a single stream of content may include media data including content items displayed on a display screen and / or audio data capturing the voice of the users. Thus, the various streams of content within the communication data enable facilitating a conference or broadcast presentation between a group of people dispersed across remote locations. Each stream may also include text, audio, and video data, such as data transmitted within a channel, chat board, or private messaging service.
[0090] Participants or attendees of a communication session are people who are within range of a camera or other image and / or audio capture device so that the person's movements and / or sounds produced when the person is viewing and / or listening to content shared via the communication session can be captured (e.g., recorded). For example, a participant can be sitting in a crowd, viewing shared content being performed in real time at a broadcast location where a stage presentation is taking place. Alternatively, a participant can be sitting in an office conference room, viewing shared content of a communication session with other colleagues via a display screen. Even further, a participant can be sitting or standing in front of a personal device (e.g., a tablet computer, smartphone, computer, etc.) and viewing the shared content of a communication session alone in their office or at home.
[0091] System 1102 includes device(s) 1110. Device(s) 1110 and / or other components of system 1102 may include distributed computing resources that communicate with each other and / or with client computing devices 1106(1) through 1106(N) via one or more networks 1108. In some examples, system 1102 may be a standalone system that is responsible for managing aspects of one or more communication sessions (e.g., communication session 1104). As an example, system 1102 may be managed by an entity such as YOUTUBE, FACEBOOK, SLACK, WEBEX, GOTOMEETING, GOOGLE HANGOUTS, etc.
[0092] Network(s) 1108 may include, for example, a public network (e.g., the Internet), a private network (e.g., an institutional and / or personal intranet), or some combination of private and public networks. Network(s) 1108 may also include any type of wired and / or wireless network, including, but not limited to, a local area network ("LAN"), a wide area network ("WAN"), a satellite network, a cable network, a Wi-Fi network, a WiMax network, a mobile communication network (e.g., 3G, 4G, etc.), or any combination thereof. Network(s) 1108 may utilize communication protocols, including packet-based and / or datagram-based protocols, such as the Internet Protocol ("IP"), the Transmission Control Protocol ("TCP"), the User Datagram Protocol ("UDP"), or other types of protocols. Furthermore, network(s) 1108 may also include a plurality of devices that facilitate network communications and / or form the hardware foundation of the network, such as switches, routers, gateways, access points, firewalls, base stations, repeaters, backbone equipment, and the like.
[0093] In some examples, network(s) 1108 may also include devices that enable connection to wireless networks, such as wireless access points (“WAPs”). Examples support connectivity via WAPs that send and receive data over various electromagnetic frequencies (e.g., radio frequencies), including WAPs that support the Institute of Electrical and Electronics Engineers (“IEEE”) 802.21 standards (e.g., 802.11g, 802.11n, 802.11ac, etc.) and other standards.
[0094] In various examples, device(s) 1110 may include one or more computing devices operating in a cluster or other grouping configuration to share resources, balance loads, improve performance, provide failover support or redundancy, or for other purposes. For example, device(s) 1110 may belong to various categories of devices, such as traditional server-type devices, desktop-type devices, and / or mobile-type devices. Thus, although device(s) 1110 are shown as a single type of device or server-type device, device(s) 110 may include a wide variety of device types and is not limited to a particular type of device. Device(s) 1110 may represent, but are not limited to, a server computer, a desktop computer, a web server computer, a personal computer, a mobile computer, a laptop computer, a tablet computer, or any other type of computing device.
[0095] The client computing device (e.g., one of client computing devices 1106(1) through 1106(N)) can belong to various classes of devices, which can be the same as or different from device(s) 1110, such as, for example, a traditional server-type device, a desktop-type device, a mobile-type device, a dedicated-type device, an embedded-type device, and / or a wearable-type device. Thus, the client computing device can include, but is not limited to, a desktop computer, a game console and / or gaming device, a tablet computer, a personal data assistant ("PDA"), a mobile phone / tablet hybrid device, a laptop computer, a telecommunications device, a computer navigation-type client computing device (e.g., a satellite-based navigation system including a Global Positioning System ("GPS") device), a wearable device, a virtual reality ("VR") device, an augmented reality ("AR") device, an implantable computing device, an automotive computer, a web-enabled television, a thin client, a terminal, an Internet of Things ("IoT") device, a workstation, a media player, a personal video recorder ("PVR"), a set-top box, a camera, an integrated component (e.g., a peripheral device) for inclusion in a computing device, a home appliance, or any other kind of computing device. Furthermore, the client computing device may include a combination of the earlier listed examples of client computing devices, such as a desktop-type device or a mobile-type device combined with a wearable device, and the like.
[0096] Client computing devices 1106(1) to 1106(N) of various classes and device types may represent any type of computing device having one or more data processing units 1192 operably connected to a computer-readable medium 1194 (e.g., via a bus 1116), which in some instances may include one or more of a system bus, a data bus, an address bus, a PCI bus, a Mini-PCI bus, and any of various local, peripheral, and / or independent buses.
[0097] Executable instructions stored on computer-readable medium 1194 may include, for example, operating system 1119 , client module 1120 , profile module 1122 , and other modules, programs, or applications that may be loaded and executed by data processing unit(s) 1192 .
[0098] The client computing devices 1106(1)-1106(N) may also include one or more interfaces 1124 to enable communications between the client computing devices 1106(1)-1106(N) and other networked devices (e.g., device(s) 1110) over the network(s) 1108. Such network interfaces 1124 may include one or more network interface controllers (NICs) or other types of transceiver devices to send and receive communications and / or data over the network. Additionally, client computing devices 1106(1) to 1106(N) may include input / output ("I / O") interfaces (devices) 1126 that enable communication with input / output devices (e.g., user input devices including peripheral input devices (e.g., game controllers, keyboards, mice, pens, sound input devices such as microphones, cameras for obtaining and providing video feeds and / or still images, touch input devices, gesture input devices, etc.), and / or output devices including peripheral output devices (e.g., displays, printers, audio speakers, touch output devices, etc.). Figure 11 The client computing device 1106 (N) is shown connected in some manner to a display device (e.g., display screen 1129 (1)) that can display a UI according to the techniques described herein.
[0099] exist Figure 11 In the example environment 1100 of FIG. 1 , client computing devices 1106 ( 1 ) through 1106 (N) can use their respective client modules 1120 to connect to each other and / or to other external devices in order to participate in a communication session 1104 or contribute activities to a collaborative environment. For example, a first user can utilize client computing device 1106 ( 1 ) to communicate with a second user of another client computing device 1106 ( 2 ). When executing client modules 1120 , the users can share data, which can result in client computing device 1106 ( 1 ) connecting to system 1102 and / or other client computing devices 1106 ( 2 ) through 1106 (N) via network(s) 1108 .
[0100] The client computing devices 1106(1) to 1106(N) (each of which is also referred to herein as a "data processing system") may use their respective profile modules 1122 to generate participant profiles ( Figure 111102 ), and provides the participant profile to other client computing devices and / or device(s) 1110 of system 1102. The participant profile may include one or more of the following: an identity of a user or a group of users (e.g., name, unique identifier (“ID”), etc.), user data (e.g., personal data), machine data such as location (e.g., IP address, room in a building, etc.), and technical capabilities, etc. The participant profile may be utilized to register a participant for a communication session.
[0101] like Figure 11 As shown in FIG, the device(s) 1110 of the system 1102 include a server module 1130 and an output module 1132. In this example, the server module 1130 is configured to receive media streams 1134(1) to 1134(N) from various client computing devices (e.g., client computing devices 1106(1) to 1106(N)). As described above, the media streams can include a video feed (e.g., audio and visual data associated with a user), audio data to be output along with the presentation of the user's avatar (e.g., an audio-only experience where no video data of the user is sent), text data (e.g., a text message), file data, and / or screen sharing data (e.g., a document, slide layout, image, video, etc. displayed on a display screen), etc. Thus, the server module 1130 is configured to receive a collection of various media streams 1134(1) to 1134(N) (the collection being referred to herein as "media data 1134") during the real-time viewing communication session 1104. In some scenarios, not all client computing devices participating in the communication session 1104 provide media streams. For example, a client computing device may be a consuming device or a “listening” device only, such that it only receives content associated with the communication session 1104 but does not provide any content to the communication session 1104.
[0102] In various examples, the server module 1130 can select aspects of the media stream 1134 to be shared with individual client computing devices of the participating client computing devices 1106(1)-1106(N). Accordingly, the server module 1130 can be configured to generate session data 1136 based on the stream 1134 and / or pass the session data 1136 to the output module 1132. The output module 1132 can then transmit communication data 1139 to the client computing devices (e.g., the client computing devices 1106(1)-1106(3) participating in the real-time viewing of the communication session). The communication data 1139 can include video, audio, and / or other content data provided by the output module 1132 based on content 1150 associated with the output module 1132 and based on the received session data 1136.
[0103] As shown, output module 1132 sends communication data 1139(1) to client computing device 1106(1), and sends communication data 1139(2) to client computing device 1106(2), and sends communication data 1139(3) to client computing device 1106(3), etc. The communication data 1139 sent to the client computing devices may be the same or may be different (e.g., the location of the content stream within the user interface may be different between devices).
[0104] In various implementations, the device(s) 1110 and / or the client module 1120 may include a GUI presentation module 1140. The GUI presentation module 1140 may be configured to analyze communication data 1139 for delivery to one or more of the client computing devices 1106. Specifically, the GUI presentation module 1140 at the device(s) 1110 and / or the client computing device 1106 may analyze the communication data 1139 to determine an appropriate manner for displaying the video, images, and / or content on the display screen 1129 of the associated client computing device 1106. In some implementations, the GUI presentation module 1140 may provide the video, images, and / or content to a presentation GUI 1146 that is rendered on the display screen 1129 of the associated client computing device 1106. The GUI presentation module 1140 may cause the presentation GUI 1146 to be rendered on the display screen 1129. The presentation GUI 1146 may include the video, images, and / or content analyzed by the GUI presentation module 1140.
[0105] In some implementations, the presentation GUI 1146 can include multiple sections or grids that can render or include video, images, and / or content for display on the display screen 1129. For example, a first section of the presentation GUI 1146 can include a video feed of a presenter or individual, and a second section of the presentation GUI 1146 can include a video feed of an individual consuming conference information provided by the presenter or individual. The GUI presentation module 1140 can populate the first and second sections of the presentation GUI 1146 in a manner that appropriately simulates an environment experience that the presenter and individual can share.
[0106] In some implementations, the GUI presentation module 1140 can zoom in or provide a zoomed view of the individual represented by the video feed to highlight the individual's reactions to the presenter, such as facial features. In some implementations, the presentation GUI 1146 can include video feeds of multiple participants associated with a meeting (e.g., a general communication session). In other implementations, the presentation GUI 1146 can be associated with a channel such as a chat channel, an enterprise team channel, etc. Thus, the presentation GUI 1146 can be associated with an external communication session that is different from a general communication session.
[0107] Figure 12 A diagram illustrating example components of an example device 1200 (also referred to herein as "computing device 100" or "system 100") configured to generate data for some of the user interfaces disclosed herein is shown. Device 1200 can generate data that can include one or more portions that can render or include video, images, virtual objects, and / or content for display on display screen 1129. Device 1200 can represent one of the device(s) described herein. Additionally or alternatively, device 1200 can represent one of client computing devices 1106.
[0108] As shown, device 1200 includes one or more data processing units 1202, computer-readable media 1204, and communication interface(s) 1206. The components of device 1200 are operatively connected, for example, via a bus 1209, which may include one or more of a system bus, a data bus, an address bus, a PCI bus, a Mini-PCI bus, and any of various local, peripheral, and / or independent buses.
[0109] As utilized herein, data processing unit(s) (e.g., data processing unit(s) 1202 and / or data processing unit(s) 1192) may represent, for example, a CPU-type data processing unit, a GPU-type data processing unit, a field programmable gate array (“FPGA”), another type of DSP, or other hardware logic component (which, in some instances, may be driven by a CPU). For example, but not limitation, illustrative types of hardware logic components that may be utilized include application specific integrated circuits (“ASICs”), application specific standard products (“ASSPs”), systems on chips (“SOCs”), complex programmable logic devices (“CPLDs”), and the like.
[0110] As used herein, computer-readable media (e.g., computer-readable media 1204 and computer-readable media 1194) can store instructions that can be executed by (multiple) data processing units. The computer-readable media can also store instructions that can be executed by an external data processing unit (e.g., by an external CPU, an external GPU) and / or by an external accelerator (e.g., an FPGA-type accelerator, a DSP-type accelerator, or any other internal or external accelerator). In various examples, at least one CPU, GPU, and / or accelerator is incorporated into the computing device, while in some examples, one or more of the CPU, GPU, and / or accelerator is external to the computing device.
[0111] Computer-readable media (also referred to herein as multiple computer-readable media) may include computer storage media and / or communication media. Computer storage media may include one or more of volatile memory, non-volatile memory, and / or other persistent and / or secondary computer storage media, removable and non-removable computer storage media implemented in any method or technology for storing information (e.g., computer-readable instructions, data structures, program modules, or other data). Thus, computer storage media includes media in tangible and / or physical form included in a device and / or in a hardware component that is part of a device or external to a device, including but not limited to random access memory (“RAM”), static random access memory (“SRAM”), dynamic random access memory (“DRAM”), phase change memory (“PCM”), read-only memory (“ROM”), erasable programmable read-only memory (“EPROM”), electrically erasable programmable read-only memory (“EEPROM”), flash memory, compact disk read-only memory (“CD-ROM”), digital versatile disk (“DVD”), optical cards or other optical storage media, magnetic cassettes, magnetic tape, magnetic disk storage devices, magnetic cards or other magnetic storage devices or media, solid-state memory devices, storage arrays, network attached storage devices, storage area networks, hosted computer storage devices, or any other storage memory, storage devices, and / or storage media that can be used to store and maintain information for access by a computing device.
[0112] In contrast to computer storage media, communication media may embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism. As defined herein, computer storage media does not include communication media. That is, computer storage media does not include communication media consisting solely of a modulated data signal, a carrier wave, or a propagated signal itself.
[0113] The communication interface(s) 1206 may represent, for example, a network interface controller ("NIC") or other type of transceiver device to send and receive communications over a network. Additionally, the communication interface(s) 1206 may include one or more cameras and / or audio devices 1222 to enable the generation of video feeds and / or still images, etc.
[0114] In the example shown, computer-readable medium 1204 includes a data repository 1208. In some examples, data repository 1208 includes a data storage device, such as a database, a data warehouse, or other type of structured or unstructured data storage device. In some examples, data repository 1208 includes a corpus and / or a relational database having one or more tables, indexes, stored procedures, etc. to enable data access, including, for example, one or more of the following: Hypertext Markup Language ("HTML") tables, Resource Description Framework ("RDF") tables, Web Ontology Language ("OWL") tables, and / or Extensible Markup Language ("XML") tables.
[0115] The data repository 1208 can store data for the operation of processes, applications, components, and / or modules stored in the computer-readable medium 1204 and / or executed by the data processing unit(s) 1218 and / or accelerator(s). For example, in some examples, the data repository 1208 can store session data 1210 (e.g., session data 1136), profile data 1212 (e.g., associated with participant profiles), and / or other data. The session data 1210 can include the total number of participants (e.g., users and / or client computing devices) in a communication session, activities occurring in the communication session, a list of invitees for the communication session, and / or other data related to when and how the communication session was conducted or hosted. The data repository 1208 can also include content data 1214, for example, including video, audio, or other content for rendering and display on one or more of the display screens 1129.
[0116] Alternatively, some or all of the data referenced above may be stored on a separate memory 1181 on board one or more data processing units 1202 (e.g., on board a CPU-type processor, a GPU-type processor, an FPGA-type accelerator, a DSP-type accelerator, and / or other accelerator). In this example, the computer-readable medium 1204 also includes an operating system 1218 and (multiple) application programming interfaces 1210 (APIs) configured to expose the functionality and data of the device 1200 to other devices. Additionally, the computer-readable medium 1204 includes one or more modules (e.g., a server module 1230, an output module 1232, and a GUI presentation module 1246), but the number of modules shown is merely an example, and the number may be higher or lower. That is, the functions described herein in association with the modules shown may be performed by a smaller number of modules or a larger number of modules on one device or across multiple devices.
[0117] It should be appreciated that, unless specifically stated otherwise, conditional language used herein (e.g., "can," "could," "might," or "may") is understood within the context to convey that certain examples include and other examples do not include certain features, elements, and / or steps. Thus, such conditional language is generally not intended to imply that one or more examples in any way require certain features, elements, and / or steps, or that one or more examples must include logic for deciding (with or without user input or prompting) whether certain features, elements, and / or steps are to be included or performed in any particular example. Unless specifically stated otherwise, conjunction language such as the phrase "at least one of X, Y, or Z" should be understood to mean that an item, term, etc. may be X, Y, or Z, or a combination thereof.
[0118] It should also be appreciated that various variations and modifications may be made to the examples described above, and that the elements therein should be understood to be among other acceptable examples. All such modifications and variations are intended to be included herein within the scope of the present disclosure and protected by the appended claims. Finally, although various configurations have been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter defined in the accompanying representations is not necessarily limited to the specific features or actions described. Rather, the specific features and actions disclosed serve as example forms of implementing the claimed subject matter.
[0119] The disclosure presented herein also covers the subject matter set forth in the following clauses:
[0120] Clause 1. A method for operation to be performed by a data processing system, the method comprising: causing display of a user interface comprising a video display area, a text entry field, and a text portion; processing video content of video data to generate rendered content for display within the video display area; receiving input text at the text entry field, the input text comprising at least one keyword; identifying a portion of the text data having the at least one keyword; and inserting the portion of the text data having the at least one keyword into the text entry field, the user interface being configured to: display the portion of the text data in a comment portion in response to receiving a confirmation input.
[0121] Clause 2. The method according to Clause 1 further includes: processing audio content associated with the video data to generate text data, the text data including phrases expressed in the audio content; parsing the text data into multiple sentences based on one or more criteria; and selecting sentences having at least one keyword, wherein inserting the portion of the text data into the text typing field includes: inserting the selected sentence into the text typing field.
[0122] Clause 3. The method according to clauses 1 and 2 further includes: analyzing text data to determine individual times of phrases in the text data, the phrases containing at least one keyword; selecting an individual phrase from a plurality of phrases, the individual time of the individual phrase being closer to the selected time marker than the individual time of another phrase in the plurality of phrases that also contains at least one keyword; and inserting the individual phrase as part of the text data to be inserted into a text entry field.
[0123] Clause 4. The method of clauses 1-3, wherein the selected time marker is based on a time indicated in the input text.
[0124] Clause 5. The method of clauses 3-4, wherein the selected time marker is based on a position of a playback cursor associated with the rendered content displayed in the video display area.
[0125] Clause 6. The method according to clauses 1-5 further includes: analyzing the text data to determine a time stamp for a portion of the text data containing at least one keyword; determining that the time stamp is within a predetermined threshold of a current time stamp, the current time stamp being for the rendered content displayed in the video display area; and inserting the portion of the text data into a text entry field in response to determining that the time stamp for the portion of the text data is within a predetermined threshold of the current time stamp, the current time stamp being for the rendered content displayed in the video display area.
[0126] Clause 7. The method according to clauses 1-6 further includes: analyzing audio content associated with the video content to detect at least one of the pitch, inflection point, or volume of a portion of the audio content; and determining a start point and an end point within the text data based on a threshold change of at least one of the pitch, inflection point, or volume, wherein the start point and the end point define the boundary of the portion of the text data.
[0127] Clause 8. The method according to clauses 1-7 further includes: analyzing audio content associated with the video content to detect at least one of the pitch, inflection point, or volume of the audio content; and determining a start point and an end point within the text data based on a threshold change of at least one of the pitch, inflection point, or volume, wherein the start point and the end point define the boundaries of a portion of the text data.
[0128] Clause 9. The method according to clauses 1-8 further includes: analyzing audio content associated with the video content to detect at least one of the pitch, inflection point, or volume of the audio content; determining a threshold level of at least one of the pitch, inflection point, or volume of the audio content; and selecting at least one of the style, arrangement, or appearance of characters of a portion of the text data in response to determining the threshold level of at least one of the pitch, inflection point, or volume of the audio content.
[0129] Clause 10. The method according to clauses 1-9 further includes: analyzing audio content associated with the video content to detect at least one of the pitch or volume of the audio content; determining a threshold degree of change in at least one of the pitch or volume of the audio content; and selecting at least one of the style, arrangement, or appearance of characters of a portion of the text data in response to determining the threshold change in at least one of the pitch or volume of the audio content.
[0130] Clause 11. A system comprising: one or more processing units; and a computer-readable medium having computer-executable instructions encoded thereon, the computer-executable instructions being used to cause the one or more processing units to perform a method comprising: causing display of a user interface comprising a video display area and a typing field; processing video content of video data to generate rendered content within the video display area; receiving input at the typing field; analyzing text data associated with the video data to determine that the input has a threshold level of correlation with a portion of the text data; and inserting a portion of the text data into the typing field in response to determining that the input has a threshold level of correlation with the portion of the text data.
[0131] Clause 12. A method according to clause 11, wherein the method further comprises: receiving a confirmation input indicating acceptance of a portion of the text data; and inserting the portion of the text data into a comment portion of a user interface in response to the confirmation input, wherein the portion of the text data in the comment portion is configured to: cause audio rendering of the audio content on a speaker.
[0132] Clause 13. The system of clauses 11-12, wherein the input and the portion of text data have a threshold level of relevance based on a number of common keywords between the input and the portion of text data.
[0133] Clause 14. The system of clauses 11-13, wherein the input has a threshold level of relevance to the portion of the textual data based on a threshold difference between a time stamp indicated in the input and a time associated with the portion of the textual data.
[0134] Clause 15. A system according to clauses 11-14, wherein the input has a threshold level of correlation with the portion of the text data based on a threshold difference between an identifier referenced in the input and another identifier associated with the portion of the text data, and a threshold difference between a time stamp indicated in the input and a time associated with the portion of the text data.
[0135] Clause 16. A system according to clauses 11-15, wherein the input has a threshold level of relevance to the portion of the text data based on an identifier referenced in the input and another identifier associated with the portion of the text data, and a number of common keywords between the input and the portion of the text data.
[0136] Clause 17. A system comprising: a unit for displaying a user interface, the user interface including a video display area, a typing field, and a comment section; a unit for processing video content of video data to generate rendered content for display within the video display area; a unit for receiving input at the typing field, the input including at least one keyword or a vocal input comprising a series of tones; a unit for selecting a portion of text data having at least one keyword, or a portion of audio content having a threshold level of correlation with a sequence of notes in the series of tones; and a unit for populating the typing field with a representation of the portion of text data having at least one keyword or the portion of audio content, the user interface being configured to display the portion of text data or the representation of the portion of audio content in the comment section in response to receiving a confirmation input.
[0137] Clause 18. The system according to clause 17 further includes: a unit for processing audio content associated with the video data to generate text data, the text data including phrases expressed in the audio content; a unit for parsing the text data into a plurality of sentences based on one or more criteria; and a unit for selecting sentences having at least one keyword, wherein filling in a portion of the text data into a typing field includes: inserting the selected sentence into the typing field.
[0138] Clause 19. The method according to clauses 17-18 further includes: a unit for analyzing text data to determine an individual time of a phrase in the text data, the phrase containing at least one keyword; a unit for selecting an individual phrase from a plurality of phrases, the individual time of the individual phrase being closer to the selected time marker than the individual time of another phrase in the plurality of phrases that also contains at least one keyword; and a unit for inserting the individual phrase as part of the text data to be inserted into a typing field.
[0139] Clause 20. The system according to clauses 17-19 further includes: a unit for analyzing text data to determine a time stamp for a portion of the text data containing at least one keyword; a unit for determining that the time stamp is within a predetermined threshold of a current time stamp, wherein the current time stamp is for the rendered content displayed in the video display area; and a unit for filling a portion of the text data into a typing field in response to determining that the time stamp for the portion of the text data is within a predetermined threshold of the current time stamp, wherein the current time stamp is for the rendered content displayed in the video display area.
Claims
1. A method for operation to be performed by a data processing system, the method comprising: causing display of a user interface comprising a video display area, a text entry field, and a comment section; receiving video data; processing video content of the video data to generate rendered content for display within the video display area; receiving input text for commenting on the video data at the text entry field, the input text including at least one keyword; In response to the input text: identifying, based on the at least one keyword, a portion of text data corresponding to spoken content included in the video content having the at least one keyword; and inserting the identified portion of the text data corresponding to the spoken content included in the video content having the at least one keyword into the text entry field; and The input text and the inserted portion of the text data are caused to be displayed in the comment section in response to receiving a confirmation input.
2. The method according to claim 1, further comprising: processing audio content associated with the video data to generate the text data, the text data including phrases expressed in the audio content; parsing the text data into a plurality of sentences based on one or more criteria; as well as A sentence having the at least one keyword is selected, wherein inserting the portion of the text data into the text entry field comprises inserting the selected sentence into the text entry field.
3. The method according to claim 1, further comprising: analyzing the text data to determine individual times of phrases in the text data, the phrases containing the at least one keyword; selecting an individual phrase from the plurality of phrases, the individual time of the individual phrase being closer to the selected time marker than the individual time of another phrase from the plurality of phrases that also includes the at least one keyword; as well as The individual phrase is inserted as the portion of the text data to be inserted into the text entry field.
4. The method according to claim 3, wherein: The selected time marker is based on a time indicated in the input text.
5. The method according to claim 3, wherein The selected time marker is based on a position of a playback cursor associated with the rendered content displayed in the video display area.
6. The method according to claim 1, further comprising: analyzing the text data to determine a time stamp for the portion of the text data containing the at least one keyword; determining that the time stamp is within a predetermined threshold of a current time stamp for the rendered content displayed in the video display area; as well as The portion of the text data is inserted into the text entry field in response to determining that the time stamp for the portion of the text data is within the predetermined threshold of the current time stamp for the rendered content displayed in the video display area.
7. The method according to claim 1, further comprising: analyzing audio content associated with the video data to detect at least one of a pitch, an inflection point, or a volume of a portion of the audio content; as well as A start point and an end point within the text data are determined based on a threshold change in at least one of the pitch, the inflection point, or the volume, wherein the start point and the end point define one or more boundaries of the inserted portion of the text data.
8. The method according to claim 1, further comprising: analyzing audio content associated with the video data to detect at least one of a pitch, an inflection point, or a volume of the audio content; as well as A start point and an end point within the text data are determined based on a threshold change in at least one of the pitch, the inflection point, or the volume, wherein the start point and the end point define one or more boundaries of the inserted portion of the text data.
9. The method according to claim 1, further comprising: analyzing audio content associated with the video data to detect at least one of a pitch, an inflection point, or a volume of the audio content; determining a threshold level for at least one of the pitch, the inflection point, or the volume of the audio content; and At least one of a style, arrangement, or appearance of characters of the portion of the text data is selected in response to determining the threshold level of at least one of the pitch, the inflection point, or the volume of the audio content.
10. The method according to claim 1, further comprising: analyzing audio content associated with the video data to detect at least one of a pitch or a volume of the audio content; determining a threshold degree of change in at least one of the pitch or the volume of the audio content; as well as At least one of a style, arrangement, or appearance of characters of the portion of the text data is selected in response to determining a threshold degree of change in at least one of the pitch or the volume of the audio content.
11. A system comprising: one or more processing units; as well as A computer-readable medium having encoded thereon computer-executable instructions for causing the one or more processing units to perform a method comprising: causing display of a user interface comprising a video display area and a typing field; receiving video data; processing video content of the video data to generate rendered content within the video display area; receiving input for commenting on the video data at the input field; analyzing textual data associated with the video data to determine that the input has a threshold level of correlation with a portion of the textual data; as well as In response to determining that the input has the threshold level of correlation with the portion of the text data, the portion of the text data is inserted into the entry field.
12. The system according to claim 11, wherein The method further comprises: receiving a confirmation input indicating acceptance of the portion of the text data; and In response to the confirmation input, the portion of the text data is inserted into a comments section of the user interface, wherein the portion of the text data in the comments section is configured to cause audio rendering of audio content associated with the video data on a speaker.
13. The system according to claim 11, wherein: The input has the threshold level of relevance to the portion of the text data based on a number of common keywords between the input and the portion of the text data.
14. The system according to claim 11, wherein: The input has the threshold level of correlation with the portion of the text data based on a threshold difference between a time stamp indicated in the input and a time associated with the portion of the text data.
15. The system according to claim 11, wherein The input has the threshold level of correlation with the portion of the text data based on an identifier referenced in the input and another identifier associated with the portion of the text data, and a threshold difference between a time stamp indicated in the input and a time associated with the portion of the text data.
Citation Information
Patent Citations
Internet search-based television
CN101422041A
Method of and client device for interactive television communication
US20020144273A1
Electronic messaging synchronized to media presentation
US20040098754A1
Searchable annotations-augmented on-line course content
US20170004139A1
Autocomplete suggestions by context-aware key-phrase generation
US20170154125A1