Methods, computing devices, and computer-readable media for performing audio playback

By automatically recognizing and playing back text content on the display using computing devices, the problem of complex menu navigation in existing technologies is solved, and intuitive text-to-speech playback is achieved, improving user experience and resource utilization efficiency.

CN115454370BActive Publication Date: 2026-08-25GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211006728.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-11-14
Publication Date
2026-08-25
Estimated Expiration
2039-11-14

AI Technical Summary

Technical Problem

In existing technologies, selecting text content for text-to-speech playback requires complex menu navigation, resulting in low user efficiency and wasted resources, especially for users unfamiliar with menu navigation. Furthermore, existing methods are not intuitive, increasing the user's learning curve and development costs.

Method used

By detecting user scrolling on the screen using computing devices, the system automatically identifies the location of text content and plays audio back within the playback area, reducing reliance on complex menus and optimizing user preferences and content segmentation using machine learning models.

Benefits of technology

It enables intuitive audio playback of text content, reduces user input, saves device resources, reduces interaction complexity and development costs, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115454370B_ABST
    Figure CN115454370B_ABST
Patent Text Reader

Abstract

A computer-implemented method of performing audio playback, comprising: obtaining data describing one or more content items comprising a plurality of portions of textual content; determining position data indicating respective positions of one or more of the portions of textual content on a display associated with the computing device; receiving data indicating user input modifying the respective positions of the one or more of the portions relative to the display; and in response to receiving the data: determining updated position data indicating respective updated positions of the one or more of the portions; identifying, based at least in part on the updated position data, that a first portion of textual content is positioned within a playback region of the display, wherein the portion is paused for a period of time prior to identifying that the portion is displayed within the playback region for automatic playback; and causing an audio signal comprising speech of at least a portion of the first portion to be automatically played back.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application filed on November 14, 2019, with application number 201980040263.6 and invention title "Automatic Audio Playback of Displayed Text Content". Technical Field

[0002] This disclosure generally relates to audio playback of text content. More specifically, this disclosure relates to systems and methods for providing intuitive audio playback of text content in response to user input actions, such as scrolling a portion of text content on a display. Background Technology

[0003] Typically, selecting a portion of text content to initiate text-to-speech playback requires navigating a complex menu. For users unfamiliar with or unaccustomed to navigating such complex menus, selecting text content for text-to-speech playback can be a significant challenge. This complexity may cause some users to completely ignore the text-to-speech feature. In other cases, it may be necessary to develop detailed tutorial services to guide users through the process of accessing the text-to-speech service for specific portions of the text content.

[0004] Even for those familiar with menu navigation, the repetitive process of selecting text can demonstrate inefficient use of device resources and be a source of user frustration. For example, a user wanting to select paragraphs from an article for text-to-speech playback is forced to repeatedly navigate to the menu for each paragraph they wish to play. As another example, the device must dedicate processing and battery resources to serving each request. Summary of the Invention

[0005] Aspects and advantages of embodiments of this disclosure will be set forth in part in the description which follows, or may be learned from the description or by practice of the embodiments.

[0006] One example aspect of this disclosure relates to a method implemented on a computer. The method includes obtaining data by one or more computing devices describing one or more content items, the one or more content items comprising multiple portions of text content. The method includes determining location data by one or more computing devices, the location data indicating the corresponding positions of one or more portions of the text content on a display associated with one or more computing devices. The method includes receiving data by one or more computing devices indicative of user input, the user input modifying the corresponding positions of one or more portions of the text content relative to the display. The method includes, in response to receiving the data indicative of user input, determining updated location data by one or more computing devices, the updated location data indicating the corresponding updated positions of one or more portions of the text content. The method includes, in response to receiving the data indicative of user input, identifying by one or more computing devices, and at least in part based on the updated location data, that a first portion of the text content is positioned within a playback area of ​​the display. The method includes, in response to receiving the data indicative of user input, causing playback of an audio signal by one or more computing devices, the audio signal comprising speech of at least a portion of the first portion of the text content.

[0007] Another example aspect of this disclosure relates to a method implemented on a computer. The method includes obtaining data by one or more computing devices describing one or more content items comprising multiple portions of text content; determining location data by the one or more computing devices, the location data indicating corresponding positions of one or more portions of the text content on a display associated with the one or more computing devices; receiving data by the one or more computing devices indicating user input, the user input modifying the corresponding positions of the one or more portions of the text content relative to the display; and in response to receiving the data indicating the user input: determining updated location data by the one or more computing devices, the updated location data indicating corresponding updated positions of the one or more portions of the text content; identifying by the one or more computing devices, and at least in part based on the updated location data, that a first portion of the text content is positioned within a playback area of ​​the display, wherein the one or more computing devices pause for a period of time before identifying that the portion of the text content is displayed within the playback area for automatic playback; and causing an audio signal of speech comprising at least a portion of the first portion of the text content to be automatically played back by the one or more computing devices.

[0008] Another example aspect of this disclosure relates to a computing device. The computing device includes one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the computing device to perform the aforementioned methods.

[0009] Another example aspect of this disclosure relates to one or more non-transitory computer-readable media. The one or more non-transitory computer-readable media store instructions that, when executed by one or more processors, cause the one or more processors to perform the aforementioned methods.

[0010] Other aspects of this disclosure relate to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.

[0011] These and other features, aspects, and advantages of the various embodiments of this disclosure will be better understood by referring to the following description and the appended claims. The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments of the disclosure and, together with the description, serve to explain the relevant principles. Attached Figure Description

[0012] Figure 1A A block diagram of an example computing system according to an example embodiment of the present disclosure is depicted.

[0013] Figure 1B A block diagram of an example computing device according to an example embodiment of the present disclosure is depicted.

[0014] Figure 1C A block diagram of an example computing device according to an example embodiment of the present disclosure is depicted.

[0015] Figure 2A-2B The illustration depicts a user interface interaction associated with a first portion of text content being positioned within a playback area, according to an example embodiment of the present disclosure.

[0016] Figure 2C-2D The illustration depicts a user interface interaction associated with a second portion of text content being positioned within a playback area, according to an example embodiment of the present disclosure.

[0017] Figures 3A-3B A method for dividing unstructured content items into multiple parts of text content, according to an example embodiment of the present disclosure, is described.

[0018] Figure 4 A user interface for modifying user playback area settings according to an example embodiment of this disclosure is depicted.

[0019] Figure 5 A flowchart is provided illustrating an example method for performing audio playback of displayed text content according to an example embodiment.

[0020] Figure 6 A flowchart depicting an example method for audio playback of text content that is stopped from being displayed, according to an example embodiment. Detailed Implementation

[0021] Overview

[0022] The exemplary aspects of this disclosure pertain to systems and methods for providing intuitive audio playback of text content in response to user input actions, such as scrolling through portions of text content on a display. Specifically, the systems and methods described herein are for the playback of audio (e.g., text-to-speech audio) comprising text content located at a position on a device display by user input. Thus, as an example, a user can simply scroll through a webpage or other content item to cause the text-to-speech system to perform audio playback of the text content displayed in one or more playback segments of the device's viewport (e.g., without requiring the user to perform additional tapping or gestures to specifically select a portion of the text content). The proposed technology represents a robust adjustment to navigation control for text-to-speech reading. Specifically, by controlling text-to-speech via scrolling instead of tapping or gestures, the proposed system provides more intuitive playback control for novice or technically inexperienced users.

[0023] As an example, a computing device (e.g., a user device such as a smartphone) can receive one or more content items comprising multiple portions of text content (e.g., multiple search results corresponding to a search request). The computing device can also receive user input (e.g., a scrolling operation that moves content downwards on the display) and corresponding position data indicating where each portion of the text content is moved on the display of the computing device. As an example, the computing device can receive position data indicating that a first portion of the text content (e.g., a search result) is moved to the upper half of the display. More specifically, if a portion of the text content is moved to a playback area on the display, the computing device can perform audio playback including at least some of that portion of the text content. As an example, if a search result is moved to the playback area of ​​the display, the computing device can provide playback of an audio signal including a summary of the search result's content. As another example, if a paragraph of a news article is moved to the playback area of ​​the display, the computing device can provide playback of an audio signal including the speech of the entire paragraph. In this way, a user can purposefully move a portion of the text content to a specific location on the display to enable audio signal playback. Therefore, users can initiate playback with a single user input instead of navigating a series of complex menus to enable the service. In some implementations, the method described herein is initiated in response to a voice search.

[0024] This disclosure provides numerous technical effects and benefits. As an example of a technical effect and benefit, the systems and methods of this disclosure enable a significant reduction in the user input required to perform operations using computing devices. By providing automatic playback of text content when a portion of the text content is within the playback area, it is no longer necessary for the user to select a portion of the text content and then perform a series of inputs to initiate playback on the computing device. This elimination of additional user-device interactions can save significant processing, battery, and memory resources associated with processing a series of complex user inputs.

[0025] As another example of technical effect and benefit, the systems and methods of this disclosure significantly reduce the complexity and clutter of device interactions associated with accessing device services. For example, users were previously required to memorize a series of menu interactions that enabled voice playback of text content. This memorization requirement might have necessitated creating tutorial services to teach users how to operate the audio playback menu, which in turn led to increased development costs and user frustration. The systems and methods of this disclosure eliminate these requirements by automatically providing intuitive playback, resulting in reduced device interaction complexity and a decreased need to create user tutorial services.

[0026] In addition to the above descriptions, users can be provided with control over whether and when the systems, programs, or features described herein can collect user information (e.g., information about a user's social networks, social behaviors or activities, occupation, user preferences, or the user's current location) and whether content or communications are sent to the user from a server. Furthermore, data can be processed in one or more ways before being stored or used to remove personally identifiable information. For example, a user's identity can be processed so that no personally identifiable information can be determined for that user, or the user's geographic location can be generalized at the place where location information is obtained (such as city, zip code, or state level) so that the user's specific location cannot be determined. Therefore, users can have control over what information about themselves is collected, how that information is used, and what information is provided to them.

[0027] More specifically, previous text content playback methods have raised numerous issues related to user device interaction. As an example, previous methods require a playback menu item that overlays a portion of the display for user interaction. This menu item can reduce the amount of display space available to the user and may potentially obscure the underlying text content. As another example, previous methods are often perceived as unintuitive by users. For instance, the complex series of steps required to utilize previous playback methods necessitates a level of user device knowledge that is rarely possessed by all users. The proposed system overcomes these problems by eliminating the need for a playback menu item and significantly reducing the learning curve associated with operating a text content playback service.

[0028] One or more computing devices disclosed herein can obtain content items that include portions of text content. In some embodiments, a content item can be one or more search results in response to a search query. As an example, one or more computing devices can obtain multiple content items (e.g., multiple search results) in response to a search query, the multiple content items including multiple portions of text content (e.g., each search result of the multiple search results). As another example, a content item can be a single search result (e.g., a list of best action movies), wherein the single search result includes multiple portions of text content (e.g., each movie in the movie list). As yet another example, a content item can be a web page or web document, a text document (e.g., a PDF or word processing document), an image including text content (e.g., an image captured by a user depicting an advertisement, street sign, etc.), and / or any other content including text.

[0029] In some implementations, multiple portions of text content can be presented in card format or included in card format. Card format can be used to structure portions of text content into text content cards (e.g., cards containing various snippets of that portion of text content). As an example, each search result of a search query can be presented in its own text content card or included in its own text content card. As another example, multiple search results can be summarized in a single text content card. Text content cards can organize portions of text content in a manner that facilitates playback of audio signals. As an example, text content cards can highlight the title and relevant contextual information of the webpage corresponding to the search results. As will be discussed in more detail in another subsection, text content cards can include one or more portions of text content from unstructured content items. For example, multiple paragraphs corresponding to a news article can be formatted into separate text content cards to provide structure to that content item.

[0030] In some implementations, a content item can be a webpage or other form of web document. Multiple sections of the text content can be divided by paragraphs, headlines, captions, or other such divisions of the text content. As an example, a content item (e.g., an archived newspaper page containing multiple news stories) can be structured as multiple text content cards, each representing a portion of the text content from that content item (e.g., the headline of a news story and the first paragraph from that news story). As another example, a webpage containing a single news article can be structured as multiple text content cards associated with sections of text content, each text content card containing an article header or paragraph of the text content.

[0031] In some implementations, content items may be unstructured or otherwise comprise unstructured text. Unstructured content may be, for example, text without clear divisions between headings, articles, paragraphs, etc. In this case, obtaining portions of the text content may include determining multiple portions of the text content from the unstructured text content. As an example, unstructured content items such as web pages and other numbers may be divided into portions of text content at least in part based on the numeric tag elements associated with the content item. As an example, web pages may be divided into portions of text content at least in part based on metadata associated with the web page (e.g., Hypertext Markup Language paragraph tags associated with the web page). In another example, in some implementations, machine intelligence (e.g., machine learning models) may be used to divide unstructured content into portions of text content. For example, based on user feedback, expected user behavior, and other relevant training data, machine intelligence may determine where to divide content items into portions of text content.

[0032] According to another aspect of this disclosure, one or more computing devices can determine the position of one or more portions of text content relative to a display of the computing device. In some embodiments, location data can indicate the position of all portions of the text content relative to the display, regardless of whether those portions of the content are positioned on the display. As an example, location data may indicate that a first portion of the text content is in the upper 20% of the display. As another example, location data may indicate that a second portion of the text content, not currently displayed, is positioned at a distance below one or more currently displayed portions of the text content. Location data can be obtained, for example, from an application that is rendering or otherwise processing the content item (e.g., a browser application), from a separate application monitoring the device display, and / or from other sources of location data.

[0033] According to another aspect of this disclosure, one or more computing devices can receive user input that modifies the location of a portion of text content. The user input can correspond to a physical user action. As an example, a user can use touch input (e.g., by placing a finger or stylus on a touch-sensitive display and swiping up or down) to scroll the display up or down. As another example, a user can use a voice command to scroll the display. As yet another example, a user can use a gesture or physical movement recognized by the computing device to scroll the display.

[0034] User input can modify the position of one or more portions of text content relative to the display. As an example, user scrolling input can move the display so that a previously displayed portion of the text content is now positioned at the top of the display. As another example, user gesture input can move a portion of the text content from the lower half of the display to the upper half.

[0035] After receiving user input, one or more computing devices can determine the updated position of one or more portions of the text content. As an example, a portion of the text content that was not previously positioned on the display (e.g., a text content card associated with a search result) may have an updated position where it is positioned on the display. As another example, a portion of the content that was previously positioned on the display may have an updated position where it is not positioned on the display.

[0036] According to another aspect of this disclosure, one or more computing devices can identify that a first portion of text content is positioned within a playback area of ​​the display. The playback area of ​​the display can be a screen area that enables audio playback associated with the first portion of the text content. As an example, if the first portion of the text content is positioned within the playback area, one or more computing devices can enable audio playback of portions of the text content. In some embodiments, multiple portions of the text content can be identified as being positioned within the playback area. For example, two relatively small text content cards (e.g., relative to the size of the display) can both be simultaneously fitted into the playback area of ​​the display.

[0037] In some embodiments, the playback area can be defined as a rectangular portion of the display. The size and position of the playback area can be determined in a variety of ways. As an example, the size and position of the playback area can be determined by user preferences. A settings menu may exist that allows the user to determine the position and size of the playback area. For example, a user might specify that the playback area should be a small rectangle positioned in the center of the display. Another user might choose to increase the size of the playback area and place it at the very top of the display. As another example, the size and position of the playback area can be determined at least in part by machine intelligence (e.g., a machine learning model). For example, a computing device can use machine intelligence to analyze past user behavior and determine the size and position of the playback area most preferred by the relevant user. Alternatively or additionally, the size and position of the playback area can be predefined and / or static.

[0038] In some embodiments, a user can move the playback area to other parts of the display. User input allows the user to move the playback area of ​​the display to place it over another portion of the text content(s). As an example, if operating a tablet device with a relatively large display capable of displaying multiple portions of text content (e.g., a scanned newspaper article with multiple columns), a user can drag the playback area from one portion of the text content (e.g., a paragraph) to another without moving the position of the underlying content. In this manner, the user can selectively identify the portions of text content(s) they wish to place within the playback area. Therefore, the playback area can have visual effects (e.g., a semi-transparent shadow associated with the area of ​​the playback area) to further enable the user to move the playback area precisely.

[0039] In some embodiments, the playback area may be associated with one or more visual effects. As an example, the playback area may have a semi-transparent shadow to allow user movement within the playback area. As another example, display areas and / or portions of the display other than the playback area may have visual effects. As an example, the playback area may have no visual effects while all other portions of the screen are covered with reduced brightness and / or blurred visual effects to emphasize text positioned within the playback area. In some embodiments, portions of text content positioned within the playback area may also have one or more visual effects. As an example, when individual characters are positioned within the playback area and / or played back by a text-to-speech system, they may be sequentially emphasized (e.g., enlarged, bolded, highlighted, etc.). As another example, all or one portions of text content positioned within the playback area may be emphasized (e.g., enlarged, bolded, highlighted, etc.), while one or more portions of text content positioned in other areas of the display may be de-emphasized (e.g., shrunk, moved off-display, blurred, darkened, etc.). As an example, when the title of the content included in a text content card is included in the playback of an audio signal, it can be highlighted.

[0040] In some embodiments, the computing device may pause for a period of time before recognizing a portion of the text content to be displayed in the playback area. By pausing for a period of time, the computing device can ensure that the user intentionally places that portion of the text content in the playback area (or vice versa). As an example, the period of time can be determined based on user preferences. As another example, the period of time can be determined based on many other factors, such as the number of portions of the text content, the user's previous preferences and actions, user input that moves the text content, etc.

[0041] According to another aspect of this disclosure, one or more computing devices can provide playback of an audio signal comprising a portion of text content. In some embodiments, the audio signal may include contextual information associated with the first portion of the text content. The contextual information may include a summary, a brief description, a title, and / or a list of items included in the first portion of the text content. As an example, a search result containing a list of 20 local restaurants can be summarized by playing only the top three results. As another example, a list of all actors in movie Z can be simply described as "a list of actors in movie Z". As yet another example, an audio signal comprising a search result for a news article may only include the title of the news article. In some embodiments, after playback of the audio signal comprising the first portion of the text content, one or more computing devices may automatically scroll to the next portion of the text content to begin playback of the audio signal comprising the next portion of the text content. In some embodiments, playback of the audio signal depends on initial user input. As an example, one or more computing devices may use a button adjacent to the playback area to request user confirmation, and the user may press the button to confirm playback of the audio signal.

[0042] In some implementations, the audio signal can be determined at least in part by machine intelligence (e.g., a machine learning model). Machine intelligence can be used to determine preferred portions and formats of the text content to be included in the audio signal. As an example, a machine learning model can identify a user's preference for including only a brief description of each portion of the text content in the audio signal. As another example, a machine learning model can determine that a second user prefers to include the entire text content in the audio signal. As yet another example, a machine learning model can determine a third user's preference for including different portions and formats of the text content in the audio signal based on the type of text content to be included.

[0043] In some implementations, the audio signal may further include a language translation of portions of the text content. The audio signal may provide language translation based on user preferences. For example, a user may be browsing a document written in a foreign language. One or more computing devices may selectively provide audio signal translation for portions of the text content located in the playback area. This allows the user to quickly receive a specific translation of a portion of the text content if they know the location of the text content they wish to translate.

[0044] As another example, character recognition can be selectively performed to identify characters in multiple portions of unrecognized text content (e.g., a scanned document without numeric character identifiers) that have moved into the playback window. The computing device can wait to perform character recognition until that portion of the text content is positioned in the playback window, thus saving the time and processing resources associated with performing character recognition on the entire content item. For example, a user can browse a PDF document containing unrecognized text content. The computing device can wait until a portion of the PDF is positioned within the playback area before performing character recognition and including the recognized characters in the speech signal.

[0045] In some implementations, the method described herein is initiated in response to a voice search. As an example, a search initiated by a user touch command will not activate playback of the audio signal, including the text content, located in the first portion of the playback area. As another example, a search initiated by a user voice search will activate playback of the audio signal, including the text content, located in the first portion of the playback area.

[0046] According to another aspect of this disclosure, one or more computing devices can stop playback of an audio signal based on a second user input. More specifically, the computing device can stop playback of the audio signal if the user input moves a portion(s) of the text content currently positioned in the playback area out of the playback area. As an example, a user can perform a scroll input and move a portion of the text content included in the current audio signal out of the playback area. If that portion of the text content is moved out of the playback area, the computing device can stop playback and then determine whether a new portion of the text content is positioned in the playback area. If the new portion of the text content is positioned in the playback area, the computing device can begin playback of the audio signal including that new portion of the text content.

[0047] Any of the machine intelligences (e.g., machine learning models) described in this paper can be trained based on user feedback (e.g., user instructions that the action performed is correct or incorrect). User feedback can be aggregated across multiple users to generate (e.g., retrain) a global model, and / or user-specific feedback can be used for personalization (e.g., retraining) a personalized model. As an example, user feedback might indicate whether a playback area is the wrong size and / or in the wrong location and / or whether content items are correctly / incorrectly divided into multiple parts of text content. Retraining the model based on user feedback can advance improved model performance.

[0048] Therefore, in addition to the intuitive playback of text content in response to user input, the systems and methods of this disclosure provide a variety of other ancillary services, such as text summarization and / or machine intelligence-based user preference management. Thus, the systems and methods of this disclosure can significantly reduce the inherent complexity of navigating complex menus and directing computing devices to enable audio playback of specific segments of content.

[0049] Example embodiments of this disclosure will now be discussed in further detail with reference to the illustrations. Figure 1A A block diagram of an example computing system 100 according to an exemplary embodiment of the present disclosure is depicted. System 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled via a network 180.

[0050] User computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0051] User computing device 102 includes one or more processors 112 and memory 114. The one or more processors 112 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operatively connected. Memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 114 can store data 116 and instructions 118 that are executed by processor 112 to cause user computing device 102 to perform operations.

[0052] User computing device 102 may include an audio playback system 119 that implements the playback experience described herein. Audio playback system 119 may be an application, a plugin (e.g., a plugin for a browser application), or other form of software that may be implemented by one or more processors 112. User computing device 102 may access one or more content items 190 (e.g., via network 180 and / or from local storage 114), and audio playback system 119 may generate an audio playback experience for text content included in one or more content items 190 (e.g., multiple search results, a single search result, a text document, a web page, a web document, a PDF, an image, etc.).

[0053] In some implementations, user computing device 102 (e.g., audio playback system 119) may store or include one or more machine learning models 120. For example, machine learning model 120 may be, or may otherwise include, various machine learning models such as neural networks (e.g., deep neural networks) or other types of machine learning models, including nonlinear and / or linear models. Neural networks may include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks.

[0054] In some implementations, one or more machine learning models 120 may be received from server computing system 130 via network 180, stored in user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, user computing device 102 may implement multiple parallel instances of a single machine learning model 120.

[0055] Additionally or alternatively, one or more machine learning models 140 may be included in, or otherwise stored and implemented by, server computing system 130, which communicates with user computing device 102 according to a client-server relationship. For example, machine learning model 140 may be implemented by server computing system 140 as part of a web service (e.g., a semantic entity recognition service). Thus, one or more models 120 may be stored and implemented at user computing device 102, and / or one or more models 140 may be stored and implemented at server computing system 130.

[0056] The audio playback system may also include a TTS system 121. The TTS system can execute any number of TTS techniques to generate audio signals that include text-to-speech (e.g., Figure 2B The audio signal 212). The voice playback system may also include a playback area processor 123. The playback area processor 123 may determine that one or more portions of the text content are located within the playback area. The playback area processor 123 may also determine the position, size, and sensitivity of the playback area based on multiple factors, such as user settings, previously analyzed user behavior, machine learning analysis of global user behavior, etc.

[0057] User computing device 102 may also include one or more user input components 122 for receiving user input. For example, user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). The touch-sensitive component may be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other means by which the user can provide input.

[0058] User computing device 102 may also include one or more speakers 124. Speakers 124 may be physically connected to device 102 or not physically connected to device 102. Speakers 124 may include stand-alone speakers, earphones, etc.

[0059] Server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operatively connected. Memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 134 can store data 136 and instructions 138 that are executed by processor 132 to cause server computing system 130 to perform operations.

[0060] In some implementations, server computing system 130 includes one or more server computing devices, or is otherwise implemented by one or more server computing devices. In instances where server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or a combination thereof.

[0061] As described above, server computing system 130 may store or otherwise include one or more machine learning models 140. For example, model 140 may be, or may otherwise include, various machine learning models. Example machine learning models include neural networks or other multi-layered nonlinear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks.

[0062] User computing device 102 and / or server computing system 130 can train models 120 and / or 140 via interaction with training computing system 150, wherein training computing system 150 is communicatively coupled via network 180. Training computing system 150 may be separate from server computing system 130, or may be part of server computing system 130.

[0063] The training computing system 150 includes one or more processors 152 and memory 154. The one or more processors 152 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operatively connected. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 154 can store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some embodiments, the training computing system 150 includes one or more server computing devices, or is otherwise implemented by one or more server computing devices.

[0064] Training computing system 150 may include model trainer 160, which uses various training or learning techniques, such as, for example, backpropagation of error, to train machine learning models 120 and / or 140 stored on user computing device 102 and / or server computing system 130. In some embodiments, performing backpropagation of error may include performing backpropagation through time truncation. Model trainer 160 may perform multiple generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.

[0065] Specifically, model trainer 160 can train machine learning models 120 and / or 140 based on a set of training data 162. Training data 162 may include, for example, user feedback or data describing user actions performed in response to various playback experience settings or operations.

[0066] In some implementations, training examples can be provided by the user computing device 102 if the user provides consent. Therefore, in such an implementation, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 on user-specific data received from the user computing device 102. In some instances, this process may be referred to as model personalization.

[0067] Model trainer 160 includes computer logic for providing the desired functionality. Model trainer 160 may be implemented as hardware, firmware, and / or software controlling a general-purpose processor. For example, in some embodiments, model trainer 160 includes a program file stored on a storage device, loaded into memory, and executed by one or more processors. In other embodiments, model trainer 160 includes one or more sets of computer-executable instructions stored in a tangible computer-readable storage medium, such as RAM, a hard disk, or optical or magnetic media.

[0068] Network 180 can be any type of communication network, such as a local area network (e.g., intranet), a wide area network (e.g., the Internet), or a combination thereof, and can include any number of wired or wireless links. Generally, communication on Network 180 can be conducted via any type of wired and / or wireless connection using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).

[0069] Figure 1A An example computing system that can be used to implement this disclosure is shown. Other computing systems may also be used. For example, in some embodiments, user computing device 102 may include a model trainer 160 and a training dataset 162. In such embodiments, model 120 can be trained and used locally at user computing device 102. In some of such embodiments, user computing device 102 may implement model trainer 160 based on user-specific data to personalize model 120.

[0070] Figure 1B A block diagram depicts an example computing device 10 implemented according to an exemplary embodiment of the present disclosure. The computing device 10 may be a user computing device or a server computing device.

[0071] The computing device 10 includes multiple applications (e.g., applications 1 to N). Each application contains its own machine learning library and machine learning model. For example, each application may include a machine learning model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.

[0072] like Figure 1B As shown, each application can communicate with multiple other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is application-specific.

[0073] Figure 1C A block diagram depicts an example computing device 50 implemented according to an example embodiment of the present disclosure. The computing device 50 may be a user computing device or a server computing device.

[0074] Computing device 50 includes multiple applications (e.g., applications 1 to N). Each application communicates with a central intelligence layer. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc. In some implementations, each application may communicate with the central intelligence layer (and the multiple models stored therein) using an API (e.g., a common API across all applications).

[0075] The central intelligence layer comprises multiple machine learning models. For example, such as... Figure 1C As shown, a corresponding machine learning model (e.g., a model) can be provided for each application and managed by a central intelligent layer. In other embodiments, two or more applications can share a single machine learning model. For example, in some embodiments, the central intelligent layer can provide a single model (e.g., a single model) for all applications. In some embodiments, the central intelligent layer is included within or otherwise implemented by the operating system of computing device 50.

[0076] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized repository for data used by computing device 50. For example... Figure 1C As shown, the central device data layer can communicate with other components of many computing devices, such as, for example, one or more sensors, context managers, device status components, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0077] Figure 2A-2B The illustration depicts a user interface interaction associated with a first portion of text content being positioned in a playback area, according to an example embodiment of this disclosure. Figure 2A At this location, the computing device display 200 includes multiple display elements. These display elements include a search query 202, portions of text content 204a-204e (e.g., search results), and a playback area 206. In some embodiments, all text results associated with the search query 202 may be referred to as content items. As illustrated in FIG2, a portion of text content 204 may be multiple search results. In other embodiments, a portion of text content 204 may be a single search result, a paragraph in a web document, a column in a news article, etc. Figures 3A-3B As explained, unstructured content items can be divided into text content 204 parts.

[0078] A portion of text content 204 may be presented in a card format or included in a card format. A card format can be used to structure a portion of text content into text content cards (e.g., cards containing various fragments of that portion of text content). As an example, each portion of text content 204 associated with search query 202 is presented in its own text content card. As another example, a portion of text content 204 may be summarized in a single text content card. The text content cards may organize the portions of text content in a manner that facilitates playback of an audio signal (e.g., audio signal 212). As an example, each portion of text content 204 includes associated URL information (e.g., “movie-database > current-movies-in-production” included in portion 204b of text content) and a selected portion of the text content from the associated search results (e.g., the top three movies included in portion 204d of text content).

[0079] Playback area 206 may be a screen area enabling playback of an audio signal 212 comprising at least a portion of a first portion of text content 204b. As an example, the first portion 204b of the text content is positioned within playback area 206. Multiple computing devices 200 may enable playback of the audio signal 212 comprising at least a portion of the first portion 204b of the text content. The audio signal may include converting a portion of the first portion 204b of the text content into speech. The audio signal may also include a translation of the first portion 204b of the text content from a first language to a second language. The audio signal may also include contextual information associated with the first portion 204b of the text content. Figure 4 The audio signal described may be based at least in part on user preferences. The audio signal may also be based at least in part on machine intelligence (e.g., machine learning type 120). In some embodiments, multiple portions of the text content (e.g., 204b and 204c) may be identified as being located in the playback area and may be included in the audio signal 212.

[0080] In some embodiments, playback of the audio signal 212 may further include determining that a first portion of the text content includes unrecognized text content. As an example, the first portion 204b of the text content may include unrecognized text content (e.g., a PDF scanned from a newspaper without numeric character identifiers). If the first portion 204b of the text content contains unrecognized text content, the computing device(s) 200 may perform character recognition to identify characters included in the unrecognized text content. Before performing character recognition and including the identified characters in the audio signal, the computing device 200 may wait until the portion of the text content including the unrecognized text is located in the playback area.

[0081] The playback area 206 can be defined as a rectangular portion of the display. The size and position of the playback area 206 can be determined in several ways. As an example, the size and position of the playback area 206 can be determined by user preferences, such as... Figure 4 As depicted. As another example, the size and location of playback area 206 may be determined at least in part by machine intelligence (e.g., machine learning model 120). For example, computing devices 200 may use machine intelligence to analyze past user behavior and determine the size and location of playback area 206 most preferred by the relevant user. Alternatively or additionally, the size and location of playback area 206 may be predefined and / or static.

[0082] Playback area 206 can be moved by the user to other parts of the display. User input can move playback area 206 of the display to place it on another part of text content 204. As an example, the user can drag the playback area from part 204b of text content to part 204c of text content without moving the position of the underlying part of text content 204. In this way, the user can selectively identify the part of text content 204 they wish to place within playback area 206. Therefore, playback area 206 can have visual effects (e.g., a semi-transparent shadow associated with the area of ​​playback area) to further enable the user to move the playback area accurately. However, visual effects are not necessary for playback area 206 in all embodiments.

[0083] Multiple computing devices 200 can pause for a period of time before recognizing a portion of the text content 204 and displaying it in the playback area. By pausing for a period of time, the multiple computing devices 200 can ensure that the user intentionally positions a portion of the text content 204 in the playback area (or vice versa). As an example, such as Figure 4The time period to be described can be determined based on user preferences. As another example, the time period can be determined based on several other factors, such as the number of portions of text content, the user's previous preferences and actions, user input that moves the text content, etc. For example, computing device 200 can pause for 1.5 seconds before determining that portion 204b of the text content is positioned in playback area 206.

[0084] exist Figure 2B At this location, the computing device display 200 includes several visual effects (e.g., 208, 210, 214) associated with determining that a portion of text content 204 is positioned within the playback area 206. When it is determined that a portion of text content 204 is positioned within the playback area 206, a deemphasis effect 208 can be used to deemphasize non-playback segments of the display. The deemphasis effect 208 may include a shrinking effect, moving non-playback segments off the display, a blurring effect, a darkening effect, or any other effect designed to indicate that a portion of the display is not within the playback area.

[0085] Emphasis effect 210 can be used to emphasize playback area 206. In some embodiments, emphasis effect 210 can be used when it is determined that a portion of the text content is positioned within playback area 206. In other embodiments, emphasis effect 210 can be used even when it is not determined that a portion of the text content is positioned within playback area 206. In this way, emphasis effect 210 can help the user move playback area 206 to another location on the display.

[0086] The emphasis effect 210 can be a semi-transparent shadow effect, bolding of text, or any other effect that can emphasize the position and size of the playback area 206. Furthermore, portions of text content positioned within the playback area can have a text emphasis effect 214. The text emphasis effect 214 can be the sequential emphasis of individual characters positioned within the playback area 206 and / or during text-to-speech system playback. This sequential emphasis can include text enlargement, bolding, highlighting, etc. For example, as... Figure 2B As depicted, a portion of the text content is bolded and italicized using text emphasis effect 214. As another example, text emphasis effect 214 can emphasize an entire portion of the text content 204 positioned in playback area 206 (e.g., enlarge, bold, highlight, etc.).

[0087] Figure 2C-2D The illustration depicts a user interface interaction associated with a second portion of text content being positioned in a playback area, according to an example embodiment of the present disclosure. Figure 2C Depicting Figure 2AThe same display layout is depicted, including search results 202 and text content sections 204a-e (e.g., search results) positioned in the same relative locations. Moving input 216 changes... Figure 2D The relative positions of portions of the text content depicted. As an example, before receiving the movement input 216, the first portion 204b of the text content is positioned within the playback area 206. After receiving the movement input 216, the first portion 204b of the text content has been positioned outside the playback area 206, while the second portion 204c of the text content has been positioned within the playback area 206.

[0088] In some embodiments, motion input 216 can be user input. User input can correspond to physical user actions. As an example, a user can use touch input to scroll the display up or down (e.g., by placing a finger or stylus on the touch-sensitive display and swiping up or down). As another example, a user can use voice commands to scroll the display. As yet another example, a user can use gestures or physical movements recognized by the computing device to scroll the display.

[0089] In some embodiments, motion input 216 may be computer-generated input. As an example, machine intelligence (e.g., machine learning model 120) may automatically move portions of text content. Automatic movement may be based on prior user behavior and / or a global user behavior model. As another example, automatic movement of portions of text content may be based on user preferences. Users can utilize an "auto-scroll" feature that allows scrolling through content items at a speed determined by the user.

[0090] Moving the first portion 204b of the text content out of the playback area 206 stops the playback of the audio signal of the first portion 204b of the text content. As an example, a user can perform a movement input 216 (e.g., scroll input) and move the first portion 204b of the text content currently being played out of the playback area 206. If the first portion 204b of the text content is moved out of the playback area 206, the computing device(s) 200 can stop playback and then determine whether the second portion 204c of the text content is positioned in the playback area 206. If the second portion 204c of the text content is positioned in the playback area 206, the computing device(s) 200 can begin playback of the audio signal including the second portion 204c of the text content.

[0091] Figures 3A-3B A method for dividing unstructured content items into multiple parts of text content, according to an example embodiment of the present disclosure, is described. Figure 3AThis includes heading 302 and unstructured content item 304. Heading 302 can be the title of unstructured content item 304 or its source. Content items (e.g., unstructured content item 304) can be unstructured, or otherwise include unstructured text. For example, unstructured content can be text without clear divisions between headings, articles, paragraphs, etc. In this case, the portion from which the text content is obtained can include multiple parts of the text content determined from the unstructured text content.

[0092] Figure 3B Unstructured content items 304 are depicted as multiple parts (e.g., 308a, 308b, 308c, 308d, 308e) that have been divided into text content 308. Unstructured content items can be divided into multiple parts of text content in many ways. As an example, unstructured content items of numbers can be divided into parts of text content at least in part based on the numeric tag elements associated with the content items. For example, unstructured content item 304 includes numeric tag elements associated with paragraph indentations 306 (e.g., 306a, 306b, 306c, 306d, 306e).

[0093] Despite Figure 3A The text shows a paragraph indentation of 306. Numerical markup elements not actively displayed can also be used to segment unstructured content items. For example, metadata associated with a webpage (e.g., a Hypertext Markup Language header tag associated with a header) can be used to segment unstructured content items into multiple parts of the content. In another example, in some implementations, machine intelligence (e.g., machine learning model 120) can be used to segment unstructured content into parts of text content. For example, based on user feedback, expected user behavior, and other relevant training data, machine intelligence can determine where to segment content items into parts of text content.

[0094] Figure 4 A user interface for modifying user playback area settings according to an example embodiment of the present disclosure is depicted. The playback area settings interface 400 includes a playback area position setting 402, a playback area size setting 404, and a playback area start time setting 406. In some embodiments, the playback area settings interface 400 may include a playback area playback speed setting. The playback speed setting may at least partially determine the speed at which portions of text content positioned within the playback area are read. For example, changing the playback area speed setting to a value of 1.25 may increase the playback speed of the audio signal by 25% or some other amount. As another example, changing the playback area speed setting to a value of 0.5 may decrease the playback speed of the audio signal by 50% or some other amount.

[0095] The location of the playback area can be determined based on the user-specified playback area location setting 402. As an example, the playback area location setting 402 allows the user to determine the location of the playback area corresponding to the display of the computing device. For instance, the user can specify that the playback area should be a small rectangle positioned in the center of the display. Another user can choose to increase the size of the playback area and position it at the very top of the display.

[0096] The size of the playback area can be determined based on user-specified playback area size setting 404. As an example, the user can specify that the size and shape of the playback area should be a square providing sufficient space to include both parts of the text content. As another example, a second user can specify that the size and shape of the playback area should be a rectangle with sufficient space to include a newspaper column. As yet another example, a third user can specify that the size and shape of the playback area should be a thin rectangle providing sufficient space to fit only one part of the text content.

[0097] The playback start time of a playback area can be determined based on the user-specified playback area start time setting 406. Setting a specific start time (e.g., pause time) can be used to ensure that the user intentionally positions a portion of the text content within the playback area (or vice versa). As an example, a user can specify that playback should begin ½ second after determining that a portion of the text content is within the playback area. As another example, a second user can specify that playback should begin 3 seconds after determining that a portion of the text content is within the playback area.

[0098] Playback region settings 400 (e.g., playback region location settings 402, playback region size settings 404, playback region start time settings 406) can be adjusted at least in part by machine intelligence (e.g., machine learning model 120). Machine intelligence can adjust user-specified playback region settings 400 based at least in part on detected inefficiencies. Machine learning model 120 can be trained at least in part on previous and current user behavior, global user behavior, and / or other factors. For example, machine intelligence can determine that a user who has specified a certain playback region location setting 402 exhibits behavior indicating that playback region location setting 402 could be located in a location preferred by the user. In some embodiments, machine intelligence can determine to automatically adjust the associated playback region settings 400 without notifying the user. In other embodiments, machine intelligence can determine to notify the user before adjusting the playback region settings 400.

[0099] Figure 5 A flowchart is depicted illustrating an example method 500 for performing audio playback of displayed text content according to an example embodiment. Method 500 can be implemented, for example, using the computing devices(s) of FIG1. Figure 5The steps are described in a specific order for illustrative and discussion purposes. Those skilled in the art using the disclosure provided herein will understand that the individual steps of any of the methods described herein may be omitted, rearranged, performed concurrently, extended, and / or modified in various ways without departing from the scope of this disclosure.

[0100] At 502, the method may include obtaining data describing one or more content items comprising multiple parts of text content. For example, a content item may be one or more search results in response to a search query. As an example, one or more computing devices may obtain (e.g., multiple search results) content items in response to a search query, wherein the content items comprise multiple parts of text content (e.g., each of the multiple search results). As another example, a content item may be a single search result (e.g., a list of best action movies), wherein the single search result comprises multiple parts of text content (e.g., each movie in the movie list).

[0101] A portion of text content can be a division of content items that include the text content. As an example, a content item can be a webpage or web document that can be divided into multiple parts of text content, or each part of the text content can be a paragraph of a webpage or web document. The portions of text content can be presented in a card format or included in a card format. A card format can be used to structure portions of text content into text content cards (e.g., cards containing various fragments of text content). As an example, each search result in a search result query can be presented in its own text content card or included in its own text content card. As another example, multiple search results can be summarized in a single text content card. Text content cards can organize a portion of text content in a way that facilitates playback of audio signals. As an example, text content cards can emphasize the title and relevant contextual information of the webpage corresponding to the search result.

[0102] At 504, the method may include determining location data indicating the corresponding positions of one or more portions of text content on a display associated with one or more computing devices. The location data may indicate the positions of all portions of the text content relative to the display, regardless of whether those portions of the content are positioned on the display. As an example, the location data may indicate that a first portion of the text content is in the upper 20% of the display. As another example, the location data may indicate that a second, not currently displayed portion of the text content is positioned at a distance below one or more currently displayed portions of the text content. The location data may be obtained, for example, from an application that is rendering or otherwise processing the content item (e.g., a browser application), from a separate application monitoring the device's display, and / or from other sources of location data.

[0103] At 506, the method may include receiving data instructing a user to input that modifies one or more portions of text content at a corresponding position relative to the display. The user input may correspond to a physical user action. As an example, a user may use touch input (e.g., by placing a finger or stylus on a touch-sensitive display and swiping up or down) to scroll the display up or down. As another example, a user may use a voice command to scroll the display. As yet another example, a user may use a gesture or physical movement recognized by the computing device to scroll the display.

[0104] User input can modify the position of one or more portions of text content relative to the display. As an example, user scrolling input can move the display so that a previously displayed portion of the text content is now positioned at the top of the display. As another example, user gesture input can move a portion of the text content from the lower half of the display to the upper half.

[0105] At 508, the method may include update position data that determines corresponding updated positions of one or more portions of text content in response to receiving data indicating user input. After one or more portions of the text content have been moved, the update position data may indicate their new corresponding positions relative to the screen. As an example, a portion of text content that was not previously positioned on the display (e.g., a text content card associated with a search result) may have an updated position that is now positioned on the display. As another example, a portion of content that was previously positioned on the display may have an updated position that is not now positioned on the display.

[0106] At 510, the method may include recognizing, in response to receiving data instructing user input and at least in part based on updated location data, that a first portion of the text content is positioned within the playback area of ​​the display. The playback area of ​​the display may be a screen area that enables audio playback associated with the first portion of the text content. As an example, if the first portion of the text content is positioned within the playback area, the one or more computing devices may enable audio playback including that portion of the text content. In some embodiments, multiple portions of the text content may be identified as being positioned within the playback area. For example, two relatively small text content cards (e.g., relative to the size of the display) may both fit within the playback area of ​​the display simultaneously.

[0107] In some embodiments, the playback area can be defined as a rectangular portion of the display. The size and position of the playback area can be determined in a variety of ways. As an example, the size and position of the playback area can be determined by user preferences. A settings menu may exist that allows the user to determine the position and size of the playback area. For example, a user might specify that the playback area should be a small rectangle positioned in the center of the display. Another user might choose to increase the size of the playback area and position it at the very top of the display. As another example, the size and position of the playback area can be determined at least in part by machine intelligence (e.g., a machine learning model). For example, a computing device can use machine intelligence to analyze past user behavior and determine the size and position of the playback area most preferred by the relevant user. Alternatively or additionally, the size and position of the playback area can be predefined and / or static.

[0108] At 512, the method may include playback of an audio signal of speech comprising at least a portion of a first portion of the text content in response to receiving data instructing user input. In some embodiments, the audio signal may include contextual information associated with the first portion of the text content. The contextual information may include a summary, a brief description, a title, and / or a list of items included in the first portion of the text content. As an example, search results containing a list of 20 local restaurants can be summarized by playing only the top three results.

[0109] In some implementations, the audio signal can be determined at least in part by machine intelligence (e.g., a machine learning model). Machine intelligence can be used to determine preferred portions and formats of the text content to be included in the audio signal. As an example, a machine learning model can identify a user's preference for including only a brief description of each portion of the text content in the audio signal. As another example, a machine learning model can determine that a second user prefers to include the entire portion of the text content in the audio signal.

[0110] In some implementations, the audio signal may also include a translation of portions of the text content from one language to another. The audio signal can provide language translation based on user preferences. For example, a user may be browsing a document written in a foreign language. One or more computing devices can selectively provide audio signal translation for portions of the text content located in the playback area. This allows the user to quickly receive a specific translation of a portion of the text content if they know the location of the text content they wish to translate.

[0111] Figure 6 A flowchart depicts an example method 600 for audio playback of stopped text content according to an example embodiment. Method 600 can be implemented, for example, using the computing devices(s) of FIG1. Figure 6The steps are described in a specific order for illustrative and discussion purposes. Those skilled in the art using the disclosure provided herein will understand that the individual steps of any of the methods described herein may be omitted, rearranged, performed concurrently, extended, and / or modified in various ways without departing from the scope of this disclosure.

[0112] At 602, method 600 may include receiving data instructing a second user to input that modifies one or more portions of content relative to a corresponding position on the display. The user input may correspond to a physical user action. As an example, a user may use touch input (e.g., by placing a finger or stylus on a touch-sensitive display and swiping up or down) to scroll the display up or down. Figure 5 Other examples of user input are described in the document.

[0113] At 604, method 600 may include determining update position data in response to receiving data instructing a second user input, the update position data indicating a corresponding second update position of one or more portions of the text content. After one or more portions of the text content have been moved, the update position data may indicate their new corresponding positions relative to the screen. As an example, a portion of the text content that was not previously positioned on the display (e.g., a text content card associated with a search result) may have an updated position positioned on the display.

[0114] At 606, method 600 may include recognizing, in response to receiving data instructing a second user input and at least in part based on second updated position data, that a first portion of the text content is positioned outside the playback area of ​​the display. For example, the first portion of the text content, included in a currently playing audio signal, may be moved out of the playback area in response to user input while the audio signal is still playing.

[0115] At 608, method 600 may include stopping the playback of an audio signal in response to receiving data instructing a second user input. If the currently playing audio signal includes a portion of text content that has been positioned outside the playback area by user input, the computing device may stop the playback of the audio signal. The stopping of the audio signal may be delayed to allow the user to reposition the first portion of the text content back into the playback area if desired. As an example, the computing device may pause for 3 seconds before stopping the audio signal including a portion of text content that has been moved outside the playback area. This allows a user who has submitted unexpected input (e.g., an accidental touch of the screen) to correct the error before playback is affected. In some embodiments, the pause may be determined at least partially by user settings. In other embodiments, the pause may be determined at least partially by machine intelligence (e.g., machine learning model 120).

[0116] Additional Disclosure

[0117] This paper discusses technologies related to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functions between and within components. For example, the processes discussed in this paper can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can work sequentially or in parallel.

[0118] While the subject matter has been described in detail with reference to various specific example embodiments, each example is provided for explanation and not for limitation of this disclosure. Those skilled in the art, upon gaining an understanding of the foregoing, can readily generate modifications, variations, and equivalents to these embodiments. Therefore, this disclosure does not exclude such modifications, variations, and / or additions to the subject matter that would be apparent to those of ordinary skill in the art. For example, features illustrated or described as part of an embodiment may be used with another embodiment to produce further embodiments. Therefore, this disclosure is intended to cover such modifications, variations, and equivalents.

Claims

1. A computer-implemented method for performing audio playback of displayed text content, the method comprising: Data describing one or more content items, comprising multiple parts of text content, is obtained from one or more computing devices; Location data is determined by the one or more computing devices, the location data indicating the corresponding position of one or more portions of text content on a display associated with the one or more computing devices; The one or more computing devices receive data instructing a user to input, wherein the user input modifies the position of one or more portions of the text content relative to the display; and In response to receiving the data indicating user input: The one or more computing devices determine update location data, which indicates the corresponding update location of the one or more portions of the text content; The first portion of the text content, identified by the one or more computing devices and at least in part based on the updated location data, is positioned within the playback area of ​​the display. The one or more computing devices pause for a period of time before the portion of the identified text content is displayed in the playback area for automatic playback; as well as The one or more computing devices cause the audio signal of speech, including at least a portion of the first part of the text content, to be automatically played back.

2. The computer-implemented method according to claim 1, wherein, The user input includes scrolling input performed by the user.

3. The computer-implemented method according to any of the preceding claims further includes: One or more visual effects are applied on the display by the one or more computing devices to one or both of the following: a first portion of text content positioned within the playback area, or one or more portions of the display other than the playback area.

4. The computer-implemented method according to claim 1 or 2, wherein, The content items include the search results page generated in response to a search query.

5. The computer-implemented method according to claim 4, wherein, The multiple portions of the text content include text content cards that are at least partially based on one or more search results in response to the search query.

6. The computer-implemented method according to claim 1 or 2, wherein, The audio signal includes text-to-speech conversion of one or more portions of the text content.

7. The computer-implemented method according to claim 1 or 2, wherein, Audio signals include translations from first language to second language.

8. The computer-implemented method according to claim 1 or 2, wherein, The audio signal includes contextual information associated with the first part of the text content.

9. The computer-implemented method according to claim 1 or 2, wherein, The speech of at least a portion of the first part of the text content is at least partially based on user preferences.

10. The computer-implemented method according to claim 1 or 2, wherein, The speech of at least a portion of the first part of the text content is determined at least partially by a machine learning model.

11. The computer-implemented method according to claim 1 or 2, wherein, The playback of an audio signal of speech comprising at least a portion of a first part of text content, caused by the one or more computing devices, further includes: The first portion of the text content determined by the one or more computing devices includes unrecognized text content; Characters included in the unrecognized text content are identified by the one or more computing devices.

12. The computer-implemented method according to claim 1 or 2, wherein, The data obtained by the one or more computing devices describing one or more content items comprising multiple portions of text content also includes: Unstructured text content is obtained from the one or more computing devices; The one or more computing devices determine multiple portions of the text content from the unstructured text content.

13. The computer-implemented method according to claim 12, wherein, The determination of multiple portions of the unstructured text content by the one or more computing devices is at least partially achieved through a machine learning model.

14. A computing device, comprising: One or more processors; and One or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the computing device to perform the method of any one of claims 1-13.

15. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1-13.

Citation Information

Patent Citations

  • Facilitating creation and playback of user-recorded audio

    CN108228132A

  • Controls for dictated text navigation

    CN109983432A