Selecting the primary source of text-to-speech based on posture
By generating parallel audio streams and applying voice attributes and gesture monitoring, the problem of users having to listen to multiple text content items one by one in existing technologies is solved, and voice output that efficiently scans and focuses on content of interest is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-09
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies, when converting multiple text content items into speech, require users to listen to each item individually to find the topic of interest, which is inefficient and makes it impossible to quickly scan and focus on specific content items.
It generates parallel audio streams of multiple text segments and applies different speech properties to distinguish the position and characteristics of each audio selection in a multidimensional sound space. It also incorporates user posture monitoring to adjust output characteristics, simulating a multidimensional sound space so that users can efficiently scan and focus on content of interest.
It improves the scanning and focusing efficiency when users convert multiple text content items into speech. By locating and adjusting audio selection in a multi-dimensional sound space, users can quickly identify topics of interest and listen to them in a focused manner.
Smart Images

Figure CN115698935B_ABST
Abstract
Description
Background Technology Technical Field
[0002] One or more embodiments of the present invention relate generally to data processing, and more specifically to gesture-based selection of primary sources of text-to-speech.
[0003] Description of related technologies
[0004] Many computing devices support text-to-speech (TTS) technology, which translates text content into speech output through speakers or other audio output devices. When a user opens a website containing multiple text content items aggregated from multiple sources (such as news websites or search engine results) and selects to use TTS to listen to the text content from that website, the TTS converts the text into speech, starting with individual text strings within the website. Requiring users to listen to the speech converted from each item of text content visible on the website one by one to find topics of interest is less efficient than visually scanning the website, which allows for quick scanning of keywords in different headings or other text of multiple text content items displayed simultaneously on a page, and quickly determining which text content items to focus on and fully read.
[0005] U.S. Patent 6,199,076 to Logan et al. describes an audio program player that "automatically plays audio program segments according to a predetermined schedule" and allows users to "listen to audio voice notifications describing each segment as needed, then play it again, and skip forward or backward to the next advance announcement, thereby immediately obtaining the information needed to determine whether a given section is of interest." Furthermore, Logan et al. describe features that identify one or more discrete segments within any program segment as "highlights," and the player, including devices operating within it, is placed in a "play highlights" mode to skip those that are not highlights, allowing listeners to review only the key points of the presentation or more quickly locate specific sections of interest within the main body of a particular program. Additionally, Logan et al. describe that "information available in text form from news sources, libraries, etc., can be converted into compressed audio form by a human reader or through conventional speech synthesis. If speech synthesis is used, the text-to-speech conversion is preferably performed by the player at client station 103." "However, Logan et al.'s audio program player still requires users to listen to audio voice notifications about segments one by one, or to listen to highlights from segments converted from text to speech. This is less efficient than users being able to quickly and intuitively browse multiple text content items displayed on a single page and select specific text content items for further inspection."
[0006] Craig Jin's "Real-time Head-Tracking 3D Audio with Infinite Simultaneous Sounds" describes "a novel approach for playing an infinite number of simultaneous sound sources and 3D audio with real-time head tracking. An infinite number of simultaneous sound sources can be spatially rendered by performing 3D audio processing offline. Smooth playback with real-time head tracking is achieved by presenting a set of binaural sound tracks with fixed and closed intervals for head orientation." However, Jin's 3D audio playback for an infinite number of simultaneous sound sources is limited to preloading existing audio sources as movie soundtracks and does not identify any content sources or convert any content sources from text to speech. Furthermore, Jin's 3D audio playback for an infinite number of simultaneous sound sources does not provide a tuning source to allow users to quickly scan all simultaneous sound sources and then select to focus on one source.
[0007] As the number of online content aggregators continues to increase, there is a need for methods, systems, and computer programs to support improved efficiency in enabling users to scan multiple text content items from content aggregators when text content items are converted from text to speech. Furthermore, with the increasing prevalence of stereo headphones and speakers for listening to multidimensional audio, there is a need for methods, systems, and computer program products to support improved source tuning to output in a multidimensional sound space to facilitate simultaneous listening to multiple sound sources, identifying topics of interest from the sound sources, and then selectively adjusting the output characteristics of specific sound sources to focus listening to those specific sources while others remain. Summary of the Invention
[0008] In view of the foregoing, one or more embodiments of the present invention provide a parallel stream of multiple audio selections generated from multiple text segments converted to speech, wherein different speech attributes are applied to each audio selection to audibly distinguish each audio selection from other audio selections, and they are played in parallel to provide an audio output that creates a space filled with different speech, which facilitates a user's efficient and audible scanning of the content of multiple text content items converted to speech. One or more embodiments of the present invention provide a parallel stream of multiple audio selections generated from multiple text segments converted to speech, wherein different speech attributes are applied to each audio selection to audibly position each audio selection at different locations within a multidimensional sound space, so that a user can scan the content of multiple audio selections presented in parallel at different audible locations, simulating a space filled with different speakers, allowing the user to identify the subject selections of interest from the audio. Furthermore, in view of the foregoing, one or more embodiments of the present invention provide monitoring of user gesture movements relative to the position of an audio selection in a multidimensional sound space as an indicator of user preference to focus on a specific audio selection converted from a text content source, and selectively adjusting the output characteristics of the specific audio selection to facilitate focused listening to the specific audio selection, while also continuing to play back the other audio selections in the multidimensional sound space, providing an audible environment for the output of aggregated text content converted into speech, similar to the visual environment for outputting aggregated text content converted into ingestible speech, allowing the user to control focused listening using gesture movements.
[0009] In one or more embodiments, the method involves a computer system identifying multiple content sources, each comprising text content, wherein the multiple content sources are ordered according to a user's interests. The method involves the computer system converting the text content of each of the multiple content sources into speech comprising individual audio selections among multiple audio choices. The method involves the computer system applying one or more speech attributes to each audio selection to specify a corresponding position of the corresponding audio selection from multiple positions in a multidimensional sound space and audibly distinguishing one or more characteristics of the corresponding audio selection from the audio attributes of other audio selections among the multiple audio selections, wherein the corresponding position of the corresponding audio selection reflects the ranking of the corresponding audio selection according to interest. The method involves the computer system outputting a parallel stream of multiple audio selections to an audio output device for stereo playback to a user of the multiple audio selections within multiple positions in a multidimensional sound space, wherein the multiple positions reflect multiple content sources ordered according to interest. One advantage of applying one or more audio attributes to each audio selection to specify the corresponding position of the corresponding audio selection from multiple locations within a multidimensional sound space, and to audibly distinguish one or more characteristics of the corresponding audio selection from other audio selections, is that the parallel streaming of audio selections supports the user in scanning the content of multiple audio selections presented in parallel at different audible locations detectable by the user, and in identifying topics of interest from the audio selections. Another advantage of identifying multiple content sources sorted according to user interests and specifying audio attributes to select the corresponding position of the audio selection reflecting the ranking of the corresponding audio selections according to interest is that the parallel streaming of audio selections provides audio output that positions the content sources that the user is most likely to be most interested in at one or more primary locations. This improves the predictability of the user's scanning of audio locations, indicating that the user will be most interested in the audio selections at one or more primary locations.
[0010] Furthermore, in one or more embodiments, the method for identifying multiple content sources by a computer system, each content source comprising text content, wherein the multiple content sources are ordered according to a user's interests, can be monitored, and the computer system accesses content accessed by the user to indicate the user's selection of content. The method may further involve, in response to detecting an indication of a user's selection of content, the computer system recording the selection indication and content in a content history log. The method may further involve the computer system analyzing the content history log for the frequency of at least one type of content. The method may further involve the computer system generating at least one topic preference for the user based on at least one type of content and a corresponding frequency. The method may further involve the computer system training a cognitive model using at least one topic preference and a corresponding frequency to identify other content including at least one topic. The method may further involve, in response to a user accessing a website comprising multiple text paragraphs, the computer system identifying multiple content sources from selections of multiple text paragraphs including at least one topic using a cognitive model. The method may further involve the computer system ordering the multiple content sources according to a corresponding priority identified by the cognitive model, based on the corresponding frequency of at least one topic present in the respective content sources. One advantage of training a cognitive model with topic preferences and frequencies based on monitored access content and user activity in response to that content, and then having the trained cognitive model selectively identify and rank content sources, is that ranking the selection of content sources played in parallel in parallel streams at audible locations according to the likelihood of interest to a particular user improves the efficiency of scanning multiple audio selections transmitted through parallel streams and finding at least one audible selection that interests the user.
[0011] Furthermore, in one or more embodiments, the method for identifying multiple content sources by a computer system, each content source comprising text content, wherein the multiple content sources are ordered according to user interests, can be used to identify multiple text segments available from a content aggregator provided by the computer system. The method may further involve the computer system dividing the multiple text segments from the content aggregator into separate content sources based on an output interface setting that specifies at least one graphical characteristic for identifying boundaries between text segments in the interface. An advantage of identifying text segments available from the content aggregator and then dividing the text segments into content sources based on an output interface setting that specifies at least one graphical feature for identifying boundaries between text segments in the interface is that the identification of content sources is independent of any particular content aggregator interface, in order to generate parallel streams of content sources from multiple types of content aggregators, and regardless of whether the content aggregator outputs text segments in the interface for conversion to speech.
[0012] Additionally, in one or more embodiments, the method involves a computer system applying one or more speech attributes to each audio selection to specify a corresponding position of the corresponding audio selection from multiple positions within a multidimensional sound space and to audibly distinguish one or more features of the corresponding audio selection from the audio attributes of other audio selections among the multiple audio selections. The corresponding position of the corresponding audio selection reflects the ranking of the corresponding audio selections sorted by interest, which can be applied by the computer system. For each audio selection, the one or more speech attributes specify the audio attributes of a specific position among the multiple positions in the multidimensional sound space of a specific audio selection converted from a specific content source with the highest ranking in the order of interest, in the primary position in front of the user. One advantage of applying the one or more speech attributes to specify the audio attributes of a specific audio selection is providing the user with audio output of multiple audio selections at different positions, sorted by interest in the multidimensional sound space according to the primary position in front of the user. The audio selection that the user is most likely to be interested in is predictably located in the position in front of the user, improving the efficiency of the user in determining the order in which to scan and focus on each audio selection among those played in parallel.
[0013] Furthermore, in one or more embodiments, the method may involve a computer system monitoring input indicating a change in user posture. The method may also involve, in response to detecting input indicating a change in user posture, the computer system analyzing the direction of movement of the user posture change relative to the positioning of multiple audio selections in a multidimensional sound space, including a 3D sound space. The method may further involve the computer system identifying a specific audio selection from the plurality of audio selections selected by the user based on the direction of movement of the user posture change relative to the positioning of the plurality of audio selections in the 3D sound space. The method may also involve the computer system applying one or more updated speech attributes to each audio selection to specify audio attributes that select a corresponding position of the corresponding audio selection from a plurality of positions in the 3D sound space and audibly distinguish one or more characteristics of the corresponding audio selection from other audio selections in the plurality of audio selections, wherein the specific audio selection is moved to a primary position in front of the user in the 3D sound space, and the volume of the specific audio selection is increased relative to the other audio selections in the plurality of audio selections. The advantage of monitoring input that indicates a change in user posture and detecting the direction of movement relative to the location of an audio selection in a multidimensional sound space is that the user provides input focused on a specific audio selection with minimal movement. This minimal movement allows the user to select a specific audio selection among multiple audio selections in a parallel stream in a hands-free manner without the need for precise movement.
[0014] In one or more embodiments, a computer system includes one or more processors, one or more computer-readable memories, one or more computer-readable storage devices, and program instructions stored on at least one of the one or more storage devices for execution by at least one of the one or more processors via at least one of the one or more memories. The stored program instructions include program instructions for identifying a plurality of content sources, each content source including text content, wherein the plurality of content sources are ordered according to a user's interests. The stored program instructions include program instructions for converting each of the plurality of content sources from text content into speech comprising individual audio selections among a plurality of audio selections. The stored program instructions include program instructions for applying one or more speech attributes to each audio selection to specify audio attributes, the audio attributes selecting a corresponding position of the corresponding audio selection from a plurality of positions in a multidimensional sound space and making one or more characteristics of the corresponding audio selection audibly distinguishable from other audio selections of the plurality of audio selections, wherein the corresponding position of the corresponding audio selection reflects a ranking of the corresponding audio selection ordered according to interests. The stored program instructions include instructions for outputting a parallel stream of the plurality of audio selections to an audio output device for stereo playback of the plurality of audio selections within the plurality of locations in the multidimensional sound space, wherein the plurality of locations reflect the plurality of content sources ordered according to interest. One advantage of applying one or more speech attributes to each audio selection to specify the corresponding location of the corresponding audio selection from the plurality of locations in the multidimensional sound space and to audio attributes that audibly distinguish one or more characteristics of the corresponding audio selection from other audio selections is that the parallel stream of audio selections supports the user in scanning the content of the plurality of audio selections presented in parallel at different audible locations detectable by the user and identifying topics of interest from the audio selections. One advantage of identifying the plurality of content sources ordered according to user interest and specifying audio attributes to select the corresponding location of the audio selection reflecting the ranking of the corresponding audio selection according to interest is that the parallel stream of audio selections provides audio output that positions the content sources that the user is most likely to be most interested in at one or more primary locations. This improves the predictability of the user scanning audio locations, indicating that the user will be most interested in the audio selections at one or more primary locations.
[0015] Furthermore, in one or more embodiments, stored program instructions for identifying multiple content sources, each comprising text content, can be directed to stored program instructions to monitor content accessed by a user for the user to select content. The computer system may further include stored program instructions for recording the selection instruction and content in a content history log in response to detecting a user's selection instruction for content. The computer system may further include stored program instructions for analyzing the content history log for the frequency of at least one type of content. The computer system may further include stored program instructions for generating at least one topic preference for the user based on at least one type of content and a corresponding frequency. The computer system may further include stored program instructions for training a cognitive model using at least one topic preference and a corresponding frequency to identify other content including at least one topic. The computer system may further include stored program instructions for identifying multiple content sources from selections of multiple text paragraphs including at least one topic via a cognitive model in response to a user accessing a website comprising multiple text paragraphs. The computer system may further include stored program instructions for ranking the multiple content sources according to a corresponding priority identified by the cognitive model based on the corresponding frequency of at least one topic present in the respective content sources. One advantage of training a cognitive model with topic preferences and frequencies based on monitored access content and user activity in response to that content, and then having the trained cognitive model selectively identify and rank content sources, is that ranking the selection of content sources played in parallel in parallel streams at audible locations according to the likelihood of interest to a particular user improves the efficiency of scanning multiple audio selections transmitted through parallel streams and finding at least one audible selection that interests the user.
[0016] Furthermore, in one or more embodiments, stored program instructions for identifying multiple content sources, each comprising text content, may point to stored program instructions for identifying an interface comprising multiple text segments obtainable from a content aggregator, wherein the multiple content sources are ordered according to the user's interests. The computer system may further include stored program instructions for dividing multiple text segments from the content aggregator into individual content sources based on an output interface setting specifying at least one graphical characteristic for identifying boundaries between text segments in the interface. An advantage of identifying text segments obtainable from the content aggregator and then dividing the text segments into content sources based on an output interface setting specifying at least one graphical characteristic for identifying boundaries between text segments in the interface is that identifying content sources is independent of any particular content aggregator interface, enabling the generation of parallel streams for content sources from multiple types of content aggregators and independent of whether the content aggregator outputs text segments in an interface for conversion to speech.
[0017] Furthermore, in one or more embodiments, the stored program instructions are used to apply one or more speech attributes to each audio selection to specify a corresponding position of the corresponding audio selection from multiple positions in a multidimensional sound space and to audibly distinguish one or more features of the corresponding audio selection from other audio attributes of the multiple audio selections, wherein the corresponding position of the corresponding audio selection reflects the ranking of the corresponding audio selection sorted according to interest, which can be directed to the stored program instructions to be applied. For each audio selection, the one or more speech attributes specify the audio attributes of a specific position among the multiple positions in the main position in front of the user within the multidimensional sound space for a specific audio selection converted from a specific content source with the highest ranking in the order of interest. One advantage of applying the one or more speech attributes to specify the audio attributes of a specific audio selection is to provide the user with audio output of multiple audio selections in different positions, the audio selections being sorted according to interest in the main position in front of the user within the multidimensional sound space, wherein the audio selection that the user is most likely to be interested in is predictably located in the position in front of the user, thereby improving the efficiency of the user in determining the order in which to scan and focus on each audio selection among the audio selections played in parallel.
[0018] Furthermore, in one or more embodiments, the computer system may include stored program instructions to monitor input indicating a change in user posture. The method may also involve, in response to detecting input indicating a change in user posture, the computer system analyzing the direction of movement of the user posture change relative to the positioning of multiple audio selections in a multidimensional sound space including a 3D sound space. The computer system may further include stored program instructions for identifying a specific audio selection from a plurality of audio selections selected by the user based on the direction of movement of the user posture change relative to the positioning of the multiple audio selections in the 3D sound space. The computer system may further include stored program instructions for applying one or more updated speech attributes to each audio selection to specify audio attributes that select a corresponding position of the corresponding audio selection from a plurality of positions in the 3D sound space and audibly distinguish one or more characteristics of the corresponding audio selection from other audio selections in the plurality of audio selections, wherein the specific audio selection is moved to a primary position in front of the user in the 3D sound space, and the volume of the specific audio selection is increased relative to other audio selections in the plurality of audio selections. The advantage of monitoring input that indicates a change in user posture and detecting the direction of movement relative to the location of an audio selection in a multidimensional sound space is that the user provides input focused on a specific audio selection with minimal movement. This minimal movement allows the user to select a specific audio selection among multiple audio selections in a parallel stream in a hands-free manner without requiring precise movement.
[0019] In one or more embodiments, the computer program product includes a computer-readable storage medium having program instructions embodied therein. The computer-readable storage medium itself is not a transient signal. The program instructions are executable by a computer to cause the computer to identify, respectively, a plurality of content sources comprising text content, wherein the plurality of content sources are ordered according to a user's interests. The program instructions are executable by a computer to cause the computer to: convert each of the plurality of content sources from the text content into speech comprising an individual audio selection among a plurality of audio selections. The program instructions are executable by a computer to cause the computer to: apply one or more speech attributes to each audio selection to specify a corresponding position of the corresponding audio selection from a plurality of positions in a multidimensional sound space and to distinguish one or more characteristics of the corresponding audio selection from the other audio selections of the plurality of audio selections audibly, wherein the corresponding position of the corresponding audio selection reflects the ranking of the corresponding audio selection according to interest. The program instructions are executable by a computer to cause the computer to output a parallel stream of the plurality of audio selections to an audio output device for stereo playback of the plurality of audio selections to the user within the plurality of positions in the multidimensional sound space, wherein the plurality of positions reflect the plurality of content sources ordered according to interest. One advantage of applying one or more audio attributes to each audio selection to specify the corresponding position of the corresponding audio selection from multiple locations within a multidimensional sound space, and to audibly distinguish one or more characteristics of the corresponding audio selection from other audio selections, is that the parallel streaming of audio selections supports the user in scanning the content of multiple audio selections presented in parallel at different audible locations detectable by the user, and in identifying topics of interest from the audio selections. Another advantage of identifying multiple content sources sorted according to user interests and specifying audio attributes to select the corresponding position of the audio selection reflecting the ranking of the corresponding audio selections according to interest is that the parallel streaming of audio selections provides audio output that positions the content sources that the user is most likely to be most interested in at one or more primary locations. This improves the predictability of the user's scanning of audio locations, indicating that the user will be most interested in the audio selections at one or more primary locations.
[0020] Furthermore, in one or more embodiments, the computer program product may include computer-executable program instructions to identify multiple content sources, each comprising text content, which are ordered according to the user's interests and may point to stored program instructions to monitor content accessed by the user for the user to select content. The computer program product may include computer-executable program instructions for recording the selection instruction and content in a content history log in response to detecting a user's selection instruction. The computer program product may include computer-executable program instructions to analyze the content history log by the computer for the frequency of at least one type of content. The computer program product may include computer-executable program instructions to generate at least one topic preference for the user based on the at least one type of content and the corresponding frequency. The computer program product may include computer-executable program instructions to train a cognitive model by the computer using the at least one topic preference and the corresponding frequency to identify other content including the at least one topic. The computer program product may include computer-executable program instructions for identifying multiple content sources from the selection of multiple text paragraphs including at least one topic by the computer in response to a user accessing a website comprising multiple text paragraphs. The computer program product may include program instructions executable by a computer to rank the plurality of content sources according to a priority identified by the cognitive model, based on the corresponding frequency of at least one topic present in the respective content sources. One advantage of training a cognitive model with topic preferences and frequencies based on monitored access to and responses to content, and then selectively identifying and ranking content sources by the trained cognitive model, is that ranking the selection of content sources played in parallel in a parallel stream at audible locations based on the likelihood of interest to a particular user, improves the efficiency of scanning multiple audio choices transmitted via parallel streams and finding at least one audible choice of interest to the user.
[0021] Furthermore, in one or more embodiments, the computer program product may include computer-executable program instructions to identify multiple content sources, each comprising text content, via a computer, wherein the multiple content sources are ordered according to user interests, and stored program instructions may be used to identify an interface comprising multiple text segments obtainable from a content aggregator. Additionally, in one or more embodiments, the computer program product may include computer-executable program instructions to divide multiple text segments from a content aggregator into individual content sources among multiple content sources based on output interface settings specifying at least one graphical characteristic for identifying boundaries between text segments in the interface. One advantage of identifying text segments obtainable from a content aggregator and then dividing the text segments into content sources based on output interface settings specifying at least one graphical characteristic for identifying boundaries between text segments in the interface is that identifying content sources is independent of any particular content aggregator interface, enabling the generation of parallel streams for content sources from multiple types of content aggregators and independent of whether the content aggregator outputs text segments in an interface for conversion to speech.
[0022] Furthermore, in one or more embodiments, the computer program product may include computer-executable program instructions to apply one or more speech attributes to each audio selection via a computer, specifying audio attributes for specific positions among multiple primary positions in front of the user within a multidimensional sound space for a specific audio selection derived from a specific content source having a highest ranking sorted by interest. One advantage of applying the one or more speech attributes to specify audio attributes for a specific audio selection is providing the user with audio output having multiple audio selections at different positions, sorted by interest in the primary position in front of the user within a multidimensional sound space, wherein the audio selection most likely of interest to the user is predictably positioned in the position in front of the user, thereby improving the efficiency of the user in determining the order in which to scan and focus on each audio selection among those played in parallel. Attached Figure Description
[0023] The appended claims set forth novel features that are considered to be characteristic of one or more embodiments of the invention. However, one or more embodiments of the invention itself will be best understood by referring to the following detailed description of illustrative embodiments, when read in conjunction with the accompanying drawings, wherein:
[0024] Figure 1 A block diagram is shown illustrating an example of a listening service implemented with a simultaneous voice playback controller, which is guided by user gesture information to position audio selection in a primary position. The audio selection is converted from text-to-speech playback of a specific content source among multiple content sources provided by a content aggregator into multiple content sources that are played in parallel from text to speech.
[0025] Figure 2 A block diagram of an example computing device implementing a simultaneous voice playback controller is shown, wherein the simultaneous voice playback controller is guided by user gesture information from a gyroscope to tune to an audio selection positioned in the primary location among multiple content sources that have been converted from text to speech and played in parallel.
[0026] Figure 3 A block diagram showing an example of the components of a simultaneous voice playback controller;
[0027] Figure 4 A block diagram illustrating an example of a computing device implementing a simultaneous voice playback controller, which is guided by user gesture information detected by a visual monitor to tune to a primary audio selection among multiple content sources that have been converted from text to speech and played in parallel.
[0028] Figure 5 A block diagram illustrates an example of adjusting the position of an audio selection in a parallel stream of audio selections derived from multiple text content sources in response to detecting user gesture information indicating the selection of a specific audio selection output within a 3D sound space.
[0029] Figure 6 An example block diagram of an interface content selection identified from a content aggregator interface is shown for a listening service that manages parallel streams of text source content converted to speech and played in parallel.
[0030] Figure 7 An example of a computer system that can implement an embodiment of the present invention is shown;
[0031] Figure 8 A high-level logic flowchart of a computer program is shown for the process and computer program used by a listening service to track user topic preferences to determine and prioritize content sources for conversion from text to speech and parallel playback in parallel streams.
[0032] Figure 9 A high-level logic flowchart of the process and computer program for managing a listening service implemented by a voice playback controller guided by user gesture information is shown.
[0033] Figure 10 An example of a high-level logic flowchart for managing a simultaneous voice playback controller is shown, whereby the controller is guided by user gesture information to position audio selection in a primary location. Audio selection is converted from text-to-speech playback of a specific content source among multiple content sources provided by a content aggregator into multiple content sources that are played in parallel from text to speech.
[0034] Figure 11 An example of a high-level logic flowchart is shown for the process and procedure of setting user preferences, user gesture movement, and gesture-based adjustment settings for a voice playback controller guided by user gesture information. Detailed Implementation
[0035] In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of the invention. However, it will be apparent to those skilled in the art that the invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form to avoid unnecessarily obscuring the invention.
[0036] Furthermore, in the following description, numerous systems are described for illustrative purposes. It is important to note, and will be apparent to those skilled in the art, that the present invention can be implemented in a variety of systems, including various computer systems and electronic devices operating any number of different types of operating systems.
[0037] Figure 1 A block diagram is shown illustrating an example of a listening service implemented with a simultaneous voice playback controller, which is guided by user gesture information to position audio selection in a primary position. The audio selection is converted from text-to-speech playback of a specific content source among multiple content sources provided by a content aggregator into multiple content sources that are played in parallel from text to speech.
[0038] In one example, listening service 110 provides a service for controlling audio output streams from one or more content sources available from content aggregator 114. In one example, content aggregator 114 includes text content that can represent multiple content sources, such as multiple articles published on a news website or in a news feed, multiple social media entries published on a social media website, multiple search results aggregated on a search results website, multiple entries in an email, multiple entries in a text chain, and other aggregations of content that can be identified as individual content items accessible to a user from an application, website, or other service. In one example, content aggregator 114 includes text content. In another example, content aggregator 114 includes graphic and audio content in addition to text content.
[0039] In one example Figure 1The embodiments described herein provide the advantage of the listening service 110, which is useful for controlling the audio output streams of one or more content sources available from the content aggregator 114, regardless of whether the content aggregator 114 intends to output the content in an audio output stream format. Advantageously, the listening service 110 includes a simultaneous audio playback controller 112 for evaluating web pages or other interfaces provided by the content aggregator 114, evaluating the presence of one or more types of depictions among individual text content items within the web page or other interface, and identifying the content source based on the depicted text content items.
[0040] In one example Figure 1 The embodiments described herein offer the advantage that the voice playback controller 112 also manages the output of a parallel stream 116 of multiple audio selections generated from multiple content sources available from the content aggregator 114, allowing the user to efficiently review the content from multiple content sources in parallel. Advantageously, the voice playback controller 112 outputs the parallel stream 116, which simultaneously includes audio tracks of each of audio selections 132, 134, and 136, each audio track being converted into audible speech from a different content source available from the content aggregator 114. In other or alternative embodiments, the parallel stream 116 includes additional or alternative audio selections.
[0041] In this example, the synchronous speech playback controller 112 converts a text-based content source from the content aggregator 114 into speech as an audible option before being output in the synchronous stream 116. Simultaneously, the speech playback controller 112 can utilize one or more selectable conversion settings to implement one or more types of text-to-speech converters for converting text-based items into speech.
[0042] also, Figure 1The embodiments described herein provide the advantage that the simultaneous voice playback controller 112 manages the output of the parallel stream 116 as stereo, wherein three-dimensional (3D) positional audio adjustment is applied to each of a plurality of audio selections to create the effect of each audio selection located in different regions of a 3D sound space 130, so that the user can effectively and audibly distinguish the content of multiple content sources in parallel. In this example, the 3D sound space 130 represents the audible output of the perceived location of sound detectable by the user listening to the parallel stream 116, such as the output of a stereo output device (such as stereo headphones worn by the user or stereo speakers broadcasting in the area where the user is present). In one example, the stereo output device represents an audio output device that simultaneously reproduces different sounds from different individual speaker units to support a surround sound effect of audio perceived as being located in different locations within a 3D sound space around the user. In another embodiment, the parallel stream 116 is output by an audio output device that supports audio output in other dimensions (including but not limited to one-dimensional or two-dimensional space).
[0043] For example, the voice playback controller 112 can simultaneously manage the output of the parallel stream 116, wherein 3D positional audio adjustment for audio selection 132 is perceived to the user's left side, 3D positional audio adjustment for audio selection 134 is perceived to the user's center, and 3D positional audio adjustment for audio selection 136 is perceived to the user's right side. In another or alternative example, the 3D positional audio adjustment can place the audio selection within the 3D sound space 130 at a position above the user, below the user, or at another perceptible position within the user's perceptible 3D area.
[0044] In one example Figure 1 The embodiments described herein provide the advantage that the simultaneous voice playback controller 112 manages additional voice attribute tuning to distinguish the characteristics of content sources that are simultaneously output as audio selections 132, 134, and 136. For example, the simultaneous voice playback controller 112 adjusts voice attributes that distinguish characteristics such as speed, dubbing, pitch, and volume to differentiate content sources that are converted into outputs as audio selections 132, 134, and 136.
[0045] Advantageously, the voice playback controller 112 simultaneously generates parallel streaming audio selection with multiple voices in a 3D sound space 130. When presented with audio including multiple voices with different characteristics and speaking from different locations, the user can effectively and efficiently navigate the multiple voices in the 3D sound space 130, acting as a result of the user's auditory attention and selective hearing in the brain, which can also be referred to as the cocktail party effect. In one example, based on a user's ability to focus an individual's auditory attention and engage in selective hearing, a person can tune to a single speaker or stimulus while simultaneously tuning out a group of speakers or filtering out a range of other stimuli, such as when a user can focus on a single conversation in a noisy room. Based on a person's ability to engage in selective hearing, a person has the ability to isolate different speakers into different streams and determine which stream is most relevant to the user, thus allowing the person to tune in to one voice or tune out of others. The human brain can have the ability to separate different auditory stimuli into different streams and also determine which stream is most relevant. Meanwhile, the advantage of the voice playback controller 112 adjusting the position and characteristics of different audio selections within the 3D sound space 130 is that the output of the parallel stream 116 simulates a crowd of speakers with different voice types, volumes, voice rates, and speaking from different positions to support the user to listen to all voices in parallel and then selectively focus on specific voices for enhanced cognitive focus, while filtering out other voices.
[0046] In addition to outputting parallel stream 116 with the characteristics and position of each audio selection adjusted to facilitate user engagement in selective listening to tune to a specific audio selection within parallel stream 116, advantageously, the voice playback controller 112 simultaneously adjusts the characteristics and position of each audio selection in response to user posture associated with the audio selection in 3D sound space 130. For example, the voice playback controller 112 simultaneously receives input of user posture information 120, analyzes the localization of the audio selection in 3D sound space 130 associated with the user posture information 120, determines the specific audio selection targeted by the user posture, and adjusts the characteristics and position of each audio selection in parallel stream 116 to control that specific audio selection as the primary source in the primary location. In one embodiment, user posture information 120 represents head movement detected by a gyroscope or other type of motion detector in a stereo headset device worn by the user. Advantageously, the voice playback controller 112 simultaneously receives user posture information 120 indicating the direction of movement, used to indicate the user's selection of an audio selection, but without requiring any specific, precise movement to guide the selection of a specific audio selection to be positioned in the primary location within 3D sound space 130.
[0047] In one embodiment, the listening service 110 evaluates the direction of movement of the user's head toward a specific audio selection within the 3D sound space 130, as identified in the user posture information 120, as input to trigger the simultaneous selection by the voice playback controller 112 of which audio selection is located at a predominant position within the 3D sound space 130. In another embodiment, the listening service 110 may evaluate other postures or posture changes indicated by the user in the user posture information 120 as input to guide the listening service 110 to switch to different content aggregators, wherein the listening service 110 may manage the playback of parallel streams from multiple content aggregators.
[0048] The advantage of the listening service 110, which supports speech-to-text conversion from multiple content sources identified by the content aggregator 114, is that the user receives audio of multiple voices converted from multiple text segments in parallel, providing efficient audio-based browsing of multiple content sources provided by the content aggregator 114. The listening service 110 manages adjustments for specific audio selections in response to user gesture information 120, playing them at a prominent position and higher volume within the 3D sound space 130. This also supports the user's selective listening process with minimal effort required to turn their head in the direction of content of interest, allowing the user to control selective listening in a hands-free environment.
[0049] Figure 2 A block diagram of an example computing device implementing a simultaneous speech playback controller is shown, which is guided by user gesture information from a gyroscope to tune to and play audio selections of the primary position among multiple content sources that are converted from text to speech in parallel.
[0050] In one example, computing device 200 includes display interface 210, where content sources from content aggregator 114 are displayed as selectable content items, such as content item A 212, content item B 214, content item C 216, and content item D 218. In one implementation, each of content item A 212, content item B 214, content item C 216, and content item D 218 represents an independent content item that can be identified as a content source by the simultaneous voice playback controller 112. In one example, each of content item A 212, content item B 214, content item C 216, and content item D 218 is displayed as text content aggregated in display interface 210, which can be individually selected by the user to visually read the text content in its entirety. Alternatively, each of content item A 212, content item B 214, content item C 216, and content item D 218 can also be individually selected by the user to listen to speech converted from the text content. In another embodiment, one or more of content items A 212, B 214, C 216, and D 218 represent audio-based content items that are selectable within the display interface 210 and do not require an initial text-to-speech conversion for output in an audible format.
[0051] In one embodiment, a user can select to read or play each selectable content item in the display interface 210 by selecting an item within the display interface 210. Furthermore, according to an advantage of one implementation, a user can select to simultaneously play each identifiable content item in the display interface 210 by selecting the audio playback selector 220 in the display interface 210 to trigger the simultaneous voice playback controller 112 of the listening service 110, or by providing another type of input to trigger the simultaneous voice playback controller 112 of the listening service 110. The advantage of the user selecting to trigger the simultaneous voice playback controller 112 to simultaneously play each identifiable content item in the display interface 210 is that the user can choose to review each content item in parallel auditorily in a similarly efficient manner. The user can visually scan each content item within the display interface 210 to select for further tuning to a specific content item.
[0052] In one example, computing device 200 controls the output of audio streams via input / output (I / O) interface 222 to a user-accessible stereo audio output device that supports 3D audio streaming. Figure 2In one embodiment shown, computing device 200 controls the output of an audio stream via I / O interface 222 to a stereo headset 240 worn on a user's head 244. In one example, I / O interface 222 can output the audio stream via a wired connection between I / O interface 222 and stereo headset 240 or via a wireless connection between I / O interface 222 and stereo headset 240.
[0053] In one example, the simultaneous voice playback controller 112 initially outputs parallel streams to a 3D sound space 230, where content item A 212 is output as audio selection 232 at "position #1", content item B 214 as audio selection 234 at "position #2", content item C 216 as audio selection 236 at "position #3", and content item D 218 as audio selection 238 at "position #4". In this example, "position #3" is considered the primary position in front of the user, with "position #1" and "position #2" to the left of the center position and "position #4" to the right of the center position. Advantageously, the simultaneous voice playback controller 112 selects the position of each content source when converting it into an output as audio selection based on the ranking order of the user's interest, positioning the audio selections that the user is most likely to be interested in at the primary position in front of the user, to support the user's application of selective hearing to effectively listen to all audio selections in parallel, where the audio selections that are likely to be most interested in are positioned at an audible position in front of the user. In one example, the order of content items within display interface 210 differs from the order of user interest determined by the simultaneous voice playback controller 112 and applied to the position selection of content items, wherein content item C216 is determined to be the highest-ranking content source and is initially located at position #3 in the primary position. Alternatively, the order of content items within display interface 210 may reflect the order of user interest.
[0054] In this example, I / O interface 222 supports an interface for detecting output from stereo headphones 240 or other output devices that detect changes in user posture and broadcast or transmit these changes as user posture information 120. In one embodiment, stereo headset 240 incorporates a gyroscope 246 that detects changes in motion of the user's head 244 by measuring rotational motion and reports these changes as user posture information 120. In other embodiments, other sensors, whether worn or not by the user, may monitor movement of the user's head 244 or other changes in the user's posture and report these changes as user posture information 120. In one embodiment, stereo headphones 240 include a gyroscope 246 for detecting changes in the orientation of the user's head to dynamically adjust relative to the user's head position to a selected actual location for positioning audio within the 3D sound space 230, while the voice playback controller 112 also receives the detected changes in motion as user posture information 120.
[0055] In one example, gyroscope 246 detects rotational motion of a user's head 244 turned to the left, as indicated by the direction of rotational motion at reference numeral 242. In one example, gyroscope 246 includes one or more sensors, each measuring angular velocity, for example, in degrees per second or revolutions per second, by monitoring rotation about one or more axes, including but not limited to three axes. In one example, each sensor in gyroscope 246 has a resonant mass that shifts with changes in angular velocity, and this resonant mass motion is converted into an electrical signal indicating the change in degrees per second or revolutions per second.
[0056] In this embodiment, in response to I / O interface 222 detecting the output from gyroscope 246 as user posture information 120, voice playback controller 112 simultaneously evaluates the user posture information 120 relative to the position of the audio selection in 3D sound space 230 to determine how to adjust the position of the audio selection in 3D sound space 230 based on user input. In one example, audio selection 236 at "position #3" is perceived by the user at the main center position within 3D sound space 230, and simultaneously voice playback controller 112 analyzes the user's leftward head movement to indicate the user's selection of audio selection 234. In response to the analysis of the user's selection of audio selection 234, voice playback controller 112 simultaneously adjusts the audio attributes of each audio selection to effectively reposition audio selection 234 to the center position within 3D sound space and moves audio selection 236 to an alternative position within 3D sound space 230. Advantageously, the voice playback controller 112 adjusts the position of each content source when converting the audio selection into an output based on the user's posture information 120 indicating a preference for a particular audio selection, so as to position the user's selected audio selection in a prominent position in front of the user, to support the user's application of selective hearing to effectively listen to a particular audio selection, while all other audio selections continue to play in the background, allowing the user to still turn their attention to tuning to other audio selections.
[0057] Figure 3 A block diagram showing an example of the components of a simultaneous voice playback controller.
[0058] In this example, the simultaneous audio playback controller 112 includes a content monitor 350. The content monitor 350 monitors content accessed by a user from one or more online sources and stores records of content access in a content history log 352. While monitoring content accessed by a user, the content monitor 350 can also analyze user preferences regarding which content they access, such as a user choosing to listen to a specific audio selection among multiple audio options in parallel stream 116, and update or filter the content history log 352 based on these user preferences.
[0059] In this example, the voice playback controller 112 also includes a topic selector 360. The topic selector 360 analyzes the content history log 352 to determine the types of topics the user is interested in and the frequency of each type. The topic selector 360 generates a topic preference 362 that includes interested topics and their frequencies, providing an indicator of the priority of interested topics. In one example, the topic selector 360 can apply cognitive analytics functions that receive a large amount of content from the content history log 352 to train and periodically update a cognitive model that learns the user's topic preference 362, and then be able to analyze new content and quickly determine whether the new content selection matches the user's topic preference 362.
[0060] In one embodiment, when a user chooses to access the listening service 110 through one or more interfaces, the topic interface controller 370 may initially request or access and also filter available text content from the content aggregator 114 for output by the topic interface controller 370. For example, the topic interface controller 370 may request current content from the content aggregator 114 based on topic preferences 362 to determine interface content selection 374 for output in the parallel stream 116. The topic interface controller 370 may further filter the content received from the content aggregator 114 using a cognitive model of the topic selector 360 to select only the content sources received from the content aggregator 114 as interface content selection 374, and sort the content sources according to the probability of user frequency of interest and other priority information available in the topic preferences 362.
[0061] In another embodiment, output interface setting 372 instructs topic interface controller 370 to divide the homepage of content aggregator 114 into content sources based on one or more graphical features that identify the boundaries of text paragraphs, such as dividing a news website into multiple content sources based on the layout of a news website with a title graphic element or title code indicating the beginning of a text paragraph boundary. Topic interface controller 370 then analyzes each content source by applying topic preferences 362 to sort and rank each content source according to a specific user's interest in the topics present in each content source.
[0062] In this example, the simultaneous voice playback controller 112 includes a text-to-speech converter 310 for converting text into speech. In this example, the theme interface controller 370 selects multiple content sources in an interface content selection 374, and the text-to-speech converter 310 converts each text content source into speech. Each speech is specified by one or more optional features for adjusting audio properties of the speech conversion to differentiate between different content sources for parallel streaming, thereby simulating an environment where a user hears multiple different speakers in parallel and supports the user selectively tuning to one speaker while others continue playing in the background. In one embodiment, the multiple optional features for adjusting the audio properties of the speech conversion include speech location and one or more characteristics, such as, but not limited to, speech type, speech volume, and speech rate.
[0063] In this example, the speech type selector 312 selects the speech type for the application via the text-to-speech converter 310, specifying the speech conversion characteristics of the content source from the interface content selector 374. The speech type may include one or more dubbed speech types in one or more languages.
[0064] In this example, the voice volume selector 314 selects the voice volume for the application via the text-to-speech converter 310 to specify the voice conversion characteristics of the content source from the interface content selector 374. The voice volume may include one or more types of volume characteristics, including (but not limited to) pitch, duration, timbre, and loudness.
[0065] In this example, the speech rate selector 316 selects a speech rate for the text-to-speech converter 310 to apply as a feature to the speech conversion of a specified content source from the interface content selector 374. The speech rate may include one or more increments that slow down or speed up the word rate per minute in the speech converted from text.
[0066] In this example, the speech source location selector 318 selects a spatial location for the application via a text-to-speech converter 310 to specify the speech conversion of the content source from the interface content selection 374. The spatial location can represent the cognitive placement of sound within the 3D source space 130, including placing the source in multiple planes at a specific perceptual distance from the user.
[0067] In one example, the audio stream mixer 330 initially applies attribute preference settings 320 to mix the speech attributes from the content source 374 and outputs an audio stream 332, such as parallel stream 116, that simultaneously plays the speech from each converted content selection. In this example, attribute preference settings 320 specify one or more preferences for specifying the speech type selector 320, speech volume selector 314, speech speed selector 316, and speech source location selector 318.
[0068] In the example, attribute preference setting 320 can specify preferences based on topic objects, topic themes, topic priorities, and other features that facilitate user differentiation between voices from different content sources during parallel streaming. For example, a user can specify attribute preference setting 320 with different voice types by topic, attribute preference setting 320 with different voice volumes by weight of interest in the topic and by the position or font size of content on the content aggregator website, attribute preference setting 320 with different voice speeds by topic of the topic, and attribute preference setting 320 with different voice source positions by weight of interest, where the most interesting content is in the main position.
[0069] In this example, the user gesture interpreter 340 detects user gesture information 120 representing user movement and determines the user's selection based on the user movement relative to the voice source position setting of the audio selection output in the audio stream 332. The audio stream mixer 330 applies gesture-based adjustment settings 342 to adjust one or more attributes applied by the voice type selector 312, voice volume selector 314, voice speed selector 316, and voice source position selector 318, reflecting the user's selection of the audio selection based on the user movement relative to the voice source position setting of the audio selection output in the audio stream 332. In one example, the gesture-based adjustment settings 342 specify that in response to the user making a movement toward the center right or left towards the audio selection, the audio stream mixer 330 moves the position applied by the voice source position selector 316 to each content transition of the voice to reposition the user's moving audio selection to a central position within the 3D sound space 230. Furthermore, the gesture-based adjustment settings 342 may include settings for adjusting the speech type, speech volume, and speech rate of audio selections in the audio mixed in the audio stream 332 to characterize the selected audio selection in a way that supports the user's ability to further selectively focus on the audio selection located at the center of the 3D source space 230. Further, the gesture-based adjustment settings 342 may include settings for adjusting the speech type, speech volume, and speech rate of audio selections in the audio mixed in the audio stream 332 to increase focus on the center-located audio selection over time, such as by increasing the volume of the center-located audio selection over time and decreasing the volume of other located audio selections over time, while all audio selections continue to play in parallel within the audio stream 332.
[0070] In one example, attribute setting controller 322 manages the user interface so that a user selects one or more preferences in each of attribute preference setting 320, gesture-based adjustment setting 342, and output interface setting 372. In one example, in addition to user preference selection, attribute setting controller 322 may prompt the user to indicate whether the application of attribute preference setting 320, gesture-based adjustment setting 342, and output interface setting 372 results in the user's desired output, or monitor user input indicating whether the application of attribute preference setting 320, gesture-based adjustment setting 342, and output interface setting 372 results in the user's desired output. In one embodiment, attribute setting controller 322 evaluates user responses and applies machine learning to train a model that manages one or more of attribute preference setting 320, gesture-based adjustment setting 342, and output interface setting 372. Advantageously, by supporting an interface for users to specify settings preferences and by continuing to train preferences based on user responses, the attribute setting controller 322 supports the adjustment of the position and characteristics of audio attributes selected in the audio output in a parallel stream, so as to customize the parallel output of multiple content sources in the 3D sound space 130 in a way that effectively supports the preferences of a particular user for participating in selective hearing, thereby generating a space filled with multiple voices.
[0071] Figure 4 A block diagram illustrating an example of a computing device implementing a simultaneous audio playback controller, which is guided by user gesture information detected by a visual monitor to tune to a primary audio selection among multiple content sources that have been converted from text to speech and played in parallel.
[0072] In one example, a listening service 110 runs on a computing device 400 and generates an audio stream that streams multiple content selections, such as audio selection 432, audio selection 434, audio selection 436, and audio selection 438, in parallel from text. In this example, an I / O interface 422 outputs a parallel stream 416 generated by the listening service 110 to a stereo speaker system 410. The stereo speaker system 410 includes multiple speaker units controlled to output audio channels in the parallel stream 416 within a 3D sound space 430, which can be audibly detected by a user present within that space.
[0073] In one embodiment, if a user is listening to a parallel stream 416 via a stereo speaker system 410, a visual monitor 412 detects movement of the user's head 444. In this example, the visual monitor 412 is positioned close to the user's head 444 to detect movement of the user's head 444 or other posture information indicating selected movement, such as the user rotating their head 444 to the left, as indicated by reference numeral 442. For example, the visual monitor 412 represents one or more types of cameras or motion detectors used to detect a specific type of movement within a recording area. In this example, the visual monitor 412 is positioned close to a component of the stereo speaker system 410 and outputs user posture information 414 to an I / O interface 422 of the computing device 400 for application via the listening service 110. Alternatively, the visual monitor 412 may be positioned independently of a component of the stereo speaker system 410 and may be a component of the computing device 400 for outputting user posture information 414 for application via the listening service 110.
[0074] Figure 5 A block diagram illustrates an example of adjusting the position of an audio selection in a parallel stream of audio selections derived from multiple text content sources in response to detecting user gesture information indicating the selection of a specific audio selection output within a 3D sound space.
[0075] In one example, 3D sound space 530 represents the sound space of a parallel stream 116 output by listening service 110, including "Audio Selection #1", "Audio Selection #2", "Audio Selection #3", and "Audio Selection #4". In this example, as shown by reference numeral 532, "Audio Selection #3" is initially located at "Position #1" within 3D sound space 530. As shown by reference numeral 534, "Audio Selection #2" is initially located at "Position #2" within 3D sound space 530. As shown by reference numeral 536, "Audio Selection #1" is initially located at "Position #3" within 3D sound space 530. As shown by reference numeral 538, "Audio Selection #4" is initially located at "Position #4" within 3D sound space 530. In addition to each audio selection being distributed at a different location within 3D sound space 530, each audio selection can be output as having audio attributes that audibly distinguish one or more characteristics of the corresponding audio selection.
[0076] In this example, the listening service 110 detects user posture information 120 reflecting the movement of the user's head rotated to the right, as indicated by reference numeral 510. Advantageously, in response to the user's head movement rotating to the right towards "Audio Selection #4" at reference numeral 538, the listening service 110 selects to reposition "Audio Section #4" to the primary position of "Position #1", as indicated by reference numeral 540, to support the user in applying selective listening to further tune to "Audio Selection #4" by positioning "Audio Selection #4" in the primary position in front of the user. Furthermore, the listening service 110 selects to reposition "Audio Section #1" to the open position of "Position #4", as indicated by reference numeral 542. In addition to repositioning "Audio Selection #1" and "Audio Selection #4", the listening service 110 advantageously adjusts the audio attributes of one or more features of each audio selection to audibly distinguish "Audio Selection #4" from other audio selections in the primary position, in order to support the user's application of selective listening to further tune to "Audio Selection #4" by increasing the audible focus feature on "Audio Selection #4" and decreasing the audible focus feature on other audio selections.
[0077] Figure 6 A block diagram illustrates an example of interface content selection identified from a content aggregator interface by a listening service that manages parallel streams of text source content converted to speech and played in parallel.
[0078] In one example, the content aggregator interface 602 includes multiple types of text content. The theme interface controller 370 applies an output interface setting 372, which includes at least one graphical property that identifies the boundaries between text sections to evaluate the graphical properties of the text content in the content aggregator interface 602, and identifies one or more content sources in the interface content selection 630 for converting text to speech and playing them in parallel in a parallel stream.
[0079] In this example, the content aggregator interface 602 includes four separate paragraphs of text content, shown as "Text Content #1" 612, "Text Content #2" 616, "Text Content #3A" 622, and "Text Content #3B" 626. In this example, the theme interface controller 370 evaluates the graphical characteristics identified by the output interface settings 372 within the content aggregator interface 602. The theme interface controller 370 recognizes that the heading 610 has a code style reflecting the boundary of the beginning "Text Content #1" 612, the line 614 is a graphical element reflecting the boundary of the beginning "Text Content #2" 616, the heading 620 has a code style reflecting the boundary of the beginning "Text Content #3A" 622, and the advertisement 624 is a graphical element that graphically separates "Text Content #3A" 622 and "Text Content #3B" 626, but is not a boundary between different text paragraphs.
[0080] Advantageously, by evaluating the graphical characteristics within the content aggregator 602 to identify boundaries between text paragraphs, the theme interface controller 370 is able to segment text paragraphs from different content aggregators into separate content sources. In this example, the theme interface controller 370 divides the content aggregator interface 602 into three content sources in the interface content selection 630 based on the boundaries identified between text paragraphs. In this example, the theme interface controller 370 divides the content aggregator interface 602 into a first content selection 632 from heading 610 and “Text Content #1” 612, a second content selection 634 from heading 620 and “Text Content #2” 616, and a third content selection 636 from “Text Content #3A” and “Text Content #3B”.
[0081] Figure 7 A block diagram of an example computer system that can implement one embodiment of the present invention is shown. The present invention can be implemented in various systems and combinations of systems composed of functional components, such as those described with reference to computer system 700, and can be communicatively connected to a network, such as network 702.
[0082] Computer system 700 includes a bus 722 or other communication device for transmitting information within computer system 700, and at least one hardware processing device, such as processor 712, coupled to bus 722 for processing information. Bus 722 preferably includes low-latency and high-latency paths connected by bridges and adapters, and is controlled within computer system 700 by multiple bus controllers. When implemented as a server or node, computer system 700 may include multiple processors designed to improve network service power.
[0083] Processor 712 may be at least one general-purpose processor that processes data under the control of software 750 during normal operation. Software 750 may include at least one of application software, operating system, middleware, and other code and computer-executable programs accessible from dynamic storage devices (such as random access memory (RAM) 714), static storage devices (such as read-only memory (ROM) 716), data storage devices (such as mass storage device 718), or other data storage media. Software 750 may include, but is not limited to, code, applications, protocols, interfaces, and processes for controlling one or more systems within a network, including but not limited to adapters, switches, servers, cluster systems, and mesh environments.
[0084] Computer system 700 can communicate with remote computers (such as server 740) or remote clients. In one example, server 740 can be connected to computer system 700 via any type of network (such as network 702), via a communication interface (such as network interface 732), or via a network link that can be connected to, for example, network 702.
[0085] In the example, multiple systems within a network environment can be communicatively connected via network 702, which is a medium used to provide communication links between different devices and computer systems connected via network 702. Network 702 may include permanent connections such as wired or fiber optic cables and temporary connections via telephone and wireless transmission connections, and may include routers, switches, gateways, and other hardware to provide communication channels between systems connected via network 702. Network 702 may represent one or more of the following: packet-switched networks, telephone-based networks, broadcast television networks, local area networks (LANs) and wired LANs, public networks, and restricted networks.
[0086] Network 702 and systems communicatively connected to computer 700 via network 702 may implement one or more layers of one or more types of network protocol stacks, which may include one or more of the following: physical layer, data link layer, network layer, transport layer, presentation layer, and application layer. For example, network 702 may implement one or more of the Transmission Control Protocol / Internet Protocol (TCP / IP) stack or the Open Systems Interconnection (OSI) stack. Furthermore, for example, network 702 may represent a global collection of networks and gateways that communicate with each other using the TCP / IP protocol suite. Network 702 may implement a secure HTTP protocol layer or other security protocols to protect communication between systems.
[0087] In this example, network interface 732 includes adapter 734 for connecting computer system 700 to network 702 via a link, and for communicatively connecting computer system 700 to server 740 or other computing systems via network 702. Although not depicted, network interface 732 may include additional software such as device drivers, additional hardware, and other controllers enabling communication. When implemented as a server, computer system 700 may include multiple communication interfaces accessible, for example, via multiple Peripheral Component Interconnect (PCI) bus bridges connected to input / output controllers. In this way, computer system 700 allows connections to multiple clients via multiple individual ports, and each port may also support multiple connections to multiple clients.
[0088] In one embodiment, the operation performed by the processor 712 can control Figure 8-11The flowchart describes the operation and other operations described herein. Operations performed by processor 712 may be requested by software 750 or other code, or steps in one embodiment of the invention may be performed by a specific hardware component containing hard-wired logic for performing these steps, or by any combination of programmed computer components and custom hardware components. In one embodiment, one or more components of computer system 700, or other components that may be integrated into one or more components of computer system 700, may include components for performing... Figure 8-11 The hard-wired logic for the operations in the flowchart.
[0089] Furthermore, the computer system 700 may include multiple peripheral components that facilitate input and output. These peripheral components are connected to multiple controllers, adapters, and expansion slots, such as input / output (I / O) interfaces 726, coupled to one of the multi-level buses 722. For example, input devices 724 may include, for example, a microphone, video capture device, image scanning system, keyboard, mouse, or other input peripherals enabled to communicate on the bus 722 via the I / O interface 726 for controlling input. Additionally, output devices 720, enabled to communicate on the bus 722 via the I / O interface 726 for controlling output, may include, for example, one or more graphics display devices, audio speakers, and haptic-sensorable output interfaces, but may also include other output interfaces. In alternative embodiments of the invention, additional or alternative input and output peripheral components may be added.
[0090] about Figure 7 The present invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention.
[0091] A computer-readable storage medium can be a tangible means for retaining and storing instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital universal disc (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards, or protrusions in recesses having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0092] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the corresponding computing / processing device.
[0093] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++, etc.) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as a standalone software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)) or may be connected to an external computer (e.g., via the Internet using an Internet service provider)... In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may be personalized to execute computer-readable program instructions by utilizing state information from the computer-readable program instructions in order to perform aspects of this invention.
[0094] The present invention will now be described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0095] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner, such that the computer-readable storage medium storing the instructions includes an article of manufacture containing instructions that implement aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.
[0096] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce computer-implemented processing, such that the instructions executed on the computer, other programmable apparatus, or other device perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than indicated in the figures. For example, depending on the functions involved, two consecutively shown blocks may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0098] Those skilled in the art will understand that Figure 7 The hardware depicted may vary. Furthermore, those skilled in the art will understand that the examples depicted are not intended to imply any architectural limitations regarding the invention.
[0099] Figure 8 A high-level logic flowchart is shown for the process and computer program used by the listening service to track user topic preferences to determine and prioritize content sources for conversion from text to speech and for parallel playback in parallel streams.
[0100] In one example, the process and computer program begin at box 800 and continue to box 802. Box 802 shows the monitoring of user instructions for content selection. Next, box 804 shows the determination of whether a user instruction for content selection has been detected. At box 804, if a user instruction for content selection has been detected, the process proceeds to box 806. Box 806 shows the logging of the user instruction and content for content selection in the content history log. Next, box 808 shows the analysis of the content history log to obtain the type and frequency of user content selections. Subsequently, box 810 shows the generation of user topic preferences based on topics among the content types frequently selected by the user. Next, box 812 shows the training of a topic preference model using the topic preferences, and the process ends.
[0101] Figure 9 A high-level logic flowchart of the process and computer program for managing the listening service implemented by a voice playback controller guided by user gesture information is shown.
[0102] In one example, the process and computer program begin at box 900 and continue to box 902. Box 902 shows determining whether the user has selected to listen to text-to-speech content. At box 902, if the user has selected to listen to text-to-speech content, the process proceeds to box 904. Box 904 shows identifying a set of potential selections of text segments that can be listened to from the content aggregator. Next, box 906 shows dividing the text segments from the content aggregator into separate content sources, these separate content sources being identified by boundaries based on graphical characteristics identified in the output interface settings applied to the output interface of the content aggregator, wherein each content source reflects separate text segments from that set. Subsequently, box 908 shows identifying selections of content sources that have content that satisfies the user's topic preferences. Next, box 910 shows sorting the selections of content sources according to the priority identified by each topic preference. Subsequently, box 912 shows applying attribute preference settings to specify the audio attributes of the speech conversion for the selection of content sources, wherein the audible location in the 3D sound space is selected according to the sorted priority. Next, box 914 illustrates the parallel streaming of text-to-speech conversion of the content source selection based on the applied audio attributes into a single stereo output stream, wherein each text segment from the content source selection is output as an audio selection associated with a different location in the 3D sound space according to the sorted priority, and this process is passed to box 916.
[0103] Box 916 illustrates monitoring user gesture information input. Next, Box 918 illustrates determining whether user gesture information input has been received. In Box 918, if user gesture information input is received, the process proceeds to Box 920. Box 920 illustrates the evaluation that triggers user gesture information input, and the process proceeds to Box 922. Returning to Box 918, if no user gesture information input is received, the process proceeds to Box 922.
[0104] Box 922 illustrates the process of determining whether user input indicates a halt to listening to text-to-speech content. At box 922, if user input to halt listening to text-to-speech content is received, the process terminates. Otherwise, at box 922, if no user input to halt listening to text-to-speech content is received, the process returns to box 916.
[0105] Figure 10 A high-level logic flowchart is shown for the process and computer program used to manage a simultaneous voice playback controller, which is guided by user gesture information to position audio selection in a primary position. The audio selection is converted from text to speech of a specific content source among multiple content sources provided by a content aggregator, and the audio selection is converted from text to speech of a specific content source among multiple content sources that are converted from text to speech and played in parallel.
[0106] In one example, the process and computer program begin at box 1000 and proceed to box 1002. Box 1002 shows the determination of whether a user gesture information input evaluation is triggered. In box 1002, if the user gesture information input evaluation is triggered, the process proceeds to box 1004. Box 1004 shows the analysis of the direction of movement of the user gesture information input relative to the location of the audio selection in the 3D sound space to identify a specific audio selection associated with the direction of movement in the 3D sound space. Next, box 1006 shows the adjustment of the storage of the content source selection to guide the location of the specific audio selection in a primary position within the 3D sound space. Subsequently, box 1008 shows the application of gesture-based adjustment settings to the adjusted selection of the content source to adjust the audio properties of the speech-transformation of the content source selection to audibly adjust the location of the audio selection in the 3D sound space and adjust the audio characteristics of the audio selection to increase the user's ability to selectively focus on a specific audio selection. Next, box 1010 shows how the parallel stream of text-to-speech conversion from the selected content source is adjusted into a single stereo output stream based on the adjusted audio attributes, wherein each of the selected text segments from the content source is output acoustically as an audio selection, and the specific audio selection is characterized by its dominant position in the 3D sound space, and the processing ends.
[0107] Figure 11 This diagram illustrates a high-level logic flowchart of the process and computer program by which a voice playback controller, guided by user gesture information, sets user preferences for the application, including attribute preference settings, user gesture movement, and gesture-based adjustment settings.
[0108] In one example, the process and computer program begin at box 1100 and proceed to box 1102. Box 1102 shows the determination of whether the user has selected to specify a text-to-speech preference. At box 1102, if the user has selected to specify a text-to-speech preference, the process proceeds to box 1104. Box 1104 shows a prompt for the user to select a sorting preference for audio priority positions within the 3D sound space, including selecting the primary position among multiple positions. Next, box 1106 shows an identification of the selection of a topic from the user's topic preferences. Next, box 1108 shows a prompt for the user to select one or more speech type, speech volume, speech position, and speech speed attributes according to the topic, subject, and presentation characteristics of the content. Next, box 1110 shows the storage of the user's selection as the user's attribute preference settings, and the process proceeds to box 1112.
[0109] Box 1112 illustrates prompting the user to perform one or more gestural movements to sample the user's choices of movements for selecting different locations within the 3D audio space. Next, Box 1114 illustrates training a cognitive model to perform user gestural movement interpretation based on the sampled user gestural movements. Next, Box 1116 illustrates prompting the user to select one or more speech type, speech volume, speech location, and speech velocity attributes associated with the user's gestural movements at different locations within the 3D audio space. Following this, Box 1118 illustrates storing the user's selections as the user's gestural-based adjustment settings, and the process concludes.
[0110] In one or more embodiments, a method, computer system, and computer program product involve identifying a plurality of content sources, each including text content, wherein the plurality of content sources are ordered according to a user's interests. The method, computer system, and computer program product involve converting each of the plurality of content sources from text content into speech comprising an individual audio selection from a plurality of audio choices. The method, computer system, and computer program product involve applying one or more speech attributes to each audio selection to specify an audio attribute that selects a corresponding position from a plurality of positions in a multidimensional sound space and distinguishes one or more characteristics of the corresponding audio selection audibly from other audio selections of the plurality of audio choices, wherein the corresponding position of the corresponding audio selection reflects a ranking of the corresponding audio selection ordered according to interest. The method, computer system, and computer program product involve outputting a parallel stream of the plurality of audio selections to an audio output device for stereo playback to a user of the plurality of audio selections within a plurality of positions in a multidimensional sound space, wherein the plurality of positions reflect the plurality of content sources ordered according to interest. One advantage of applying one or more audio attributes to each audio selection to specify the corresponding position of the corresponding audio selection from multiple locations within a multidimensional sound space, and to audibly distinguish one or more characteristics of the corresponding audio selection from other audio selections, is that the parallel streaming of audio selections supports users in scanning the content of multiple audio selections presented in parallel at different audible locations detectable by the user, and identifies the topics of interest to the user from the audio selections. Another advantage of identifying multiple content sources sorted according to user interests and specifying audio attributes to select the corresponding position of the audio selection reflecting the ranking of the corresponding audio selections according to interest is that the parallel streaming of audio selections provides audio output that positions the content sources that the user is most likely to be most interested in at one or more primary positions. This improves the predictability of the user's scanning of audio positions, with audio selections at one or more primary positions being of the highest interest to the user.
[0111] Furthermore, in one or more embodiments, methods, computer systems, and computer program products involving the identification of multiple content sources, each comprising text content, may also involve monitoring content accessed by a user for the user to select content, wherein the multiple content sources are ordered according to the user's interests. The methods, computer systems, and computer program products may further involve recording the selection instruction and content in a content history log in response to detecting a user's indication of content selection. The methods, computer systems, and computer program products may also further involve analyzing the content history log to determine the frequency of at least one type of content. The methods, computer systems, and computer program products may further involve generating at least one topic preference for the user based on the at least one type of content and the corresponding frequency. The methods, computer systems, and computer program products may further involve training a cognitive model using the at least one topic preference and the corresponding frequency to identify other content including the at least one topic. The methods, computer systems, and computer program products may further involve: in response to a user accessing a website comprising multiple text paragraphs, identifying the multiple content sources from selections of the multiple text paragraphs including the at least one topic using a cognitive model. The method, computer system, and computer program product may further involve ranking the plurality of content sources based on the corresponding frequency of the at least one topic present in the respective content sources, according to a corresponding priority identified by the cognitive model. One advantage of training a cognitive model with topic preferences and frequencies based on monitored access to and responses to content, and then selectively identifying and ranking content sources by the trained cognitive model, is that ranking the selection of content sources played in parallel in a parallel stream at audible locations based on the likelihood of interest to a particular user, improves the efficiency of scanning multiple audio selections transmitted through parallel streams and finding at least one audible selection of interest to the user.
[0112] Furthermore, in one or more embodiments, the method, computer system, and computer program product relating to identifying multiple content sources, each comprising text content, may also relate to identifying an interface comprising multiple text segments obtainable from a content aggregator, wherein the multiple content sources are ordered according to the user's interests. The method, computer system, and computer program product may further relate to dividing the multiple text segments from the content aggregator into individual content sources among multiple content sources based on an output interface setting that specifies at least one graphical characteristic for identifying boundaries between text segments in the interface. One advantage of identifying text segments obtainable from a content aggregator and then dividing the text segments into content sources based on an output interface setting that specifies at least one graphical characteristic for identifying boundaries between text segments in the interface is that identifying content sources is independent of any particular content aggregator interface, enabling the generation of parallel streams for content sources from multiple types of content aggregators and independent of whether the content aggregator outputs text segments in an interface for conversion to speech.
[0113] Furthermore, in one or more embodiments, the method, computer system, and computer program product involve applying one or more voice attributes to each audio selection by a computer system to specify a corresponding position of the corresponding audio selection from multiple positions within a multidimensional sound space and to audibly distinguish one or more features of the corresponding audio selection from other audio attributes of the multiple audio selections, wherein the corresponding position of the corresponding audio selection reflects the ranking of the corresponding audio selection sorted according to interest and may also point to the audio attributes applied to each audio selection, the one or more voice attributes being used to specify the audio attributes of a specific position among the multiple positions for a specific audio selection converted from a specific content source having the highest ranking sorted according to interest in the multidimensional sound space in front of the user. One advantage of applying the one or more voice attributes to specify the audio attributes of a specific audio selection is to provide the user with audio output having multiple audio selections in different positions, the audio selections being sorted according to interest in the multidimensional sound space in front of the user, wherein the audio selection that the user is most likely to be interested in is predictably located in the position in front of the user, to improve the efficiency of the user determining the order in which to scan and focus on each audio selection among the parallel-playing audio selections.
[0114] Furthermore, in one or more embodiments, the method, computer system, and computer program product may also involve monitoring input indicating a change in user posture. The method, computer system, and computer program product may also involve: in response to detecting input indicating a change in user posture, analyzing the direction of movement of the user posture change relative to the location of a plurality of audio selections in a multidimensional sound space including a 3D sound space. The method, computer system, and computer program product may also involve identifying a specific audio selection from a plurality of audio selections selected by the user based on the direction of movement of the user posture change relative to the location of the plurality of audio selections in the 3D sound space. The method, computer system, and computer program product may also involve applying one or more updated speech attributes to each audio selection to specify an audio attribute that selects a corresponding position of the corresponding audio selection from a plurality of positions within the 3D sound space and audibly distinguishes one or more features of the corresponding audio selection from other audio selections within the plurality of audio selections, wherein the specific audio selection is moved to a primary position in front of the user within the 3D sound space, and the volume of the specific audio selection is increased relative to other audio selections within the plurality of audio selections. The advantage of monitoring input that indicates a change in user posture and detecting the direction of movement relative to the location of an audio selection in a multidimensional sound space is that the user provides input focused on a specific audio selection with minimal movement. This minimal movement allows the user to select a specific audio selection among multiple audio selections in a parallel stream in a hands-free manner without requiring precise movement.
[0115] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used herein, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” are intended to also include the plural forms. It will be further understood that when the terms “comprises” and / or “comprising” are used in this specification, they specify the presence of the stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or combinations thereof.
[0116] All means or steps in the following claims, plus corresponding structures, materials, actions, and equivalents of the functional elements, are intended to include any structure, material, or action for performing the function in conjunction with other claimed elements as specifically claimed. Descriptions of one or more embodiments of the invention have been presented for purposes of illustration and description, but are not intended to be exhaustive or limited to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the invention. The embodiments were chosen and described in order to best explain the principles and practical application of the invention and to enable others skilled in the art to understand various embodiments of the invention with various modifications suitable for the intended particular purpose.
[0117] Although the invention has been specifically shown and described with reference to one or more embodiments, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the invention.
Claims
1. A computer-implemented method comprising: identifying, by a computer system, a plurality of content sources each comprising textual content, wherein the plurality of content sources are ordered according to interests of a user; converting, by the computer system, each of the plurality of content sources from textual content to speech comprising an individual audio selection of a plurality of audio selections; applying, by the computer system, one or more speech properties to each audio selection to specify a respective position of the respective audio selection selected from a plurality of positions within a multidimensional sound space and an audio property that audibly distinguishes one or more characteristics of the respective audio selection from other audio selections of the plurality of audio selections, wherein the respective position of the respective audio selection reflects a ranking of the respective audio selection according to the ordering of interests; and outputting, by the computer system, a concurrent stream of the plurality of audio selections to an audio output device for stereophonic playback of the plurality of audio selections within a plurality of positions within the multidimensional sound space to the user, wherein the plurality of positions reflect the plurality of content sources ordered according to the interests of the user; wherein identifying, by the computer system, the plurality of content sources each comprising textual content, wherein the plurality of content sources are ordered according to the interests of the user further comprises: monitoring, by the computer system, content accessed by the user for an indication of a selection of the content by the user; responsive to detecting the indication of the selection of the content by the user, recording, by the computer system, the indication of the selection and the content in a content history log; analyzing, by the computer system, the content history log for a frequency of at least one type of content; generating, by the computer system, at least one topic preference for the user based on the at least one type of content and a corresponding frequency; training, by the computer system, a cognitive model with the at least one topic preference and the corresponding frequency to identify other content comprising the at least one topic; and responsive to the user accessing a website comprising a plurality of text passages, identifying, by the computer system, a plurality of content sources from a selection of the plurality of text passages comprising the at least one topic by the cognitive model; and ordering, by the computer system, the plurality of content sources according to a respective priority identified by the cognitive model based on a respective frequency of the at least one topic present in a respective content source of the plurality of content sources.
2. The method of claim 1, wherein, identifying, by the computer system, the plurality of content sources each comprising textual content, wherein the plurality of content sources are ordered according to the interests of the user further comprises: identifying, by the computer system, an interface comprising a plurality of text passages available from a content aggregator; and partitioning, by the computer system, the plurality of text passages from the content aggregator into individual content sources of a plurality of content sources based on an output interface setting specifying at least one graphical characteristic used to identify boundaries between text passages in the interface.
3. The method of claim 1, wherein, applying, by the computer system, the one or more voice properties to each audio selection to specify audio properties that select a respective position of the respective audio selection from among a plurality of positions within a multidimensional sound space and audibly distinguish the one or more characteristics of the respective audio selection from other audio selections of the plurality of audio selections, wherein the respective position of the respective audio selection reflects a ranking of the respective audio selection according to the interest ordering, further comprising: applying, by the computer system, the one or more voice properties to each audio selection to distinguish the one or more characteristics, the one or more voice properties including one or more of one or more voice type properties, one or more voice volume properties, and one or more voice speed properties.
4. The method of claim 3, wherein, applying, by the computer system, the one or more voice properties including one or more of the one or more voice type properties, the one or more voice volume properties, one or more voice position properties, and the one or more voice speed properties to each audio selection further comprises: selecting, by the computer system, for each audio selection, one or more voice properties based on a topic, a graphical presentation characteristic of the respective content source, and a property preference setting specified by the topic.
5. The method of claim 1, wherein, applying, by the computer system, the one or more voice properties to each audio selection to specify audio properties that select a respective position of the respective audio selection from among a plurality of positions within a multidimensional sound space and audibly distinguish the one or more characteristics of the respective audio selection from other audio selections of the plurality of audio selections, wherein the respective position of the respective audio selection reflects a ranking of the respective audio selection according to the interest ordering, further comprising: applying, by the computer system, one or more voice properties to each audio selection to specify that the audio properties for a particular position of a plurality of positions in a primary position in front of the user within the multidimensional sound space are selected from a particular audio that transitions from a particular content source having a highest ranking according to the interest ordering.
6. The method of claim 1, further comprising: monitoring, by the computer system, for input indicative of a user posture change; in response to detecting the input indicative of the user posture change, analyzing, by the computer system, a direction of movement of the user posture change relative to a positioning of a plurality of audio selections in a multidimensional sound space including a 3D sound space; identifying, by the computer system, a particular audio selection from the plurality of audio selections selected by the user based on the direction of movement of the user posture change relative to the positioning of the plurality of audio selections in the 3D sound space; and applying, by the computer system, one or more updated audio properties to each audio selection to specify an audio property that selects a respective position of the respective audio selection from among a plurality of positions within the 3D sound space and audibly distinguishes the one or more characteristics of the respective audio selection from other audio selections of the plurality of audio selections, wherein a particular audio selection is moved to a primary position in front of the user within the 3D sound space and a volume of the particular audio selection is increased relative to other audio selections of the plurality of audio selections.
7. The method of claim 6, wherein, monitoring, by the computer system, for an input indicative of a change in posture of the user further comprises: receiving, by the computer system, output from a gyroscope that measures angular velocity, the angular velocity measuring rotational movement of a head of a user listening to the concurrent stream of audio output.
8. A computer system comprising one or more processors, one or more computer- readable memories, one or more computer-readable storage devices, and program instructions stored on at least one of the one or more storage devices for execution by at least one of the one or more processors with at least one of the one or more memories, the stored program instructions comprising: program instructions to identify a plurality of content sources each comprising textual content, wherein the plurality of content sources are ordered according to interests of a user; program instructions to convert each content source of the plurality of content sources from the textual content to speech comprising an individual audio selection of a plurality of audio selections; program instructions to apply one or more audio properties to each audio selection to specify an audio property that selects a respective position of the respective audio selection from among a plurality of positions within a multidimensional sound space and audibly distinguishes the one or more characteristics of the respective audio selection from other audio selections of the plurality of audio selections, wherein the respective position of the respective audio selection reflects a ranking of the respective audio selection according to the ordering of interests; and program instructions to output a concurrent stream of the plurality of audio selections to an audio output device for stereophonic playback of the plurality of audio selections within the plurality of positions within the multidimensional sound space to the user, wherein the plurality of positions reflect the plurality of content sources ordered according to the interests; wherein the stored program instructions to identify the plurality of content sources each comprising textual content further comprise: program instructions to monitor content accessed by a user for an indication of a selection of the content by the user; program instructions to record, in response to detecting the indication of the selection of the content by the user, the indication of the selection and the content in a content history log; program instructions to analyze the content history log for a frequency of at least one type of content; program instructions to generate, for the user, at least one topic preference based on the at least one type of content and a corresponding frequency; program instructions to train a cognitive model with the at least one topic preference and the corresponding frequency to identify other content comprising the at least one topic; and program instructions to output the other content to the user. In response to the user accessing a website comprising a plurality of text passages, program instructions to identify, by a cognitive model, a plurality of content sources from a selection of the plurality of text passages comprising the at least one topic; and Program instructions to order the plurality of content sources according to respective priorities identified by the cognitive model based on respective frequencies of the at least one topic in respective ones of the plurality of content sources.
9. The computer system of claim 8, wherein, The stored program instructions to identify the plurality of content sources each comprising text content further comprise: Program instructions to identify an interface comprising a plurality of text passages available from a content aggregator; and Program instructions to divide the plurality of text passages from the content aggregator into separate ones of the plurality of content sources based on output interface settings specifying at least one graphical characteristic used to identify boundaries between text passages in the interface.
10. The computer system of claim 8, wherein, The stored program instructions apply the one or more voice properties to each audio selection to specify a respective position of the respective audio selection from among a plurality of positions within a multidimensional sound space and audio properties that audibly distinguish the one or more characteristics of the respective audio selection from other ones of the plurality of audio selections, wherein the respective position of the respective audio selection reflects a ranking of the respective audio selection according to interest ordering, the program instructions further comprising: Program instructions to apply the one or more voice properties to each audio selection to distinguish the one or more characteristics, the one or more voice properties comprising one or more of one or more voice type properties, one or more voice volume properties, and one or more voice speed properties.
11. The computer system of claim 10, wherein the stored program instructions that apply the one or more voice attributes to the each audio selection further comprise: One or more of the one or more voice type properties, the one or more voice volume properties, the one or more voice position properties, and the one or more voice speed properties: Program instructions to select the one or more voice properties for each audio selection based on properties preferences settings specified by topic, by graphical presentation characteristics of respective content sources, and by topic.
12. The computer system of claim 8, wherein, The stored program instructions apply the one or more voice properties to each audio selection to specify a respective position of the respective audio selection from among a plurality of positions within a multidimensional sound space and audio properties that audibly distinguish the one or more characteristics of the respective audio selection from other ones of the plurality of audio selections, wherein the respective position of the respective audio selection reflects a ranking of the respective audio selection according to interest ordering, the program instructions further comprising: Program instructions to apply the one or more voice properties to each audio selection to specify an audio property for a particular position in a plurality of positions in a primary position in front of the user within the multidimensional sound space selected from a particular audio having a highest ranking according to interest ordering from a particular content source.
13. The computer system of claim 8, wherein, The stored program instructions further comprise: Program instructions to monitor input indicative of a change in user posture; program instructions for analyzing a movement direction of a user posture change relative to a position of a plurality of audio selections in a multi-dimensional sound space comprising a 3D sound space in response to detecting an input indicative of the user posture change; program instructions for identifying a particular audio selection from a plurality of audio selections selected by a user based on a movement direction of a user posture change relative to a position of the plurality of audio selections in a 3D sound space; and program instructions for applying one or more voice properties to each audio selection to specify audio properties that select a respective position of a respective audio selection from a plurality of positions within the 3D sound space and audibly distinguish the one or more characteristics of the respective audio selection from other audio selections of the plurality of audio selections, wherein the particular audio selection is moved to a primary position in front of the user within the 3D sound space and a volume of the particular audio selection is increased relative to other audio selections of the plurality of audio selections.
14. The computer system of claim 13, wherein, The stored program instructions for monitoring the input indicative of the user posture change further comprise: program instructions for receiving an output from a gyroscope that measures an angular velocity that measures rotational movement of a head of a user listening to the parallel stream of audio outputs.
15. A computer program product, comprising a computer readable storage medium having program instructions embodied therewith, wherein, The computer-readable storage medium is not a transitory signal per se, and the program instructions are executable by a computer to cause the computer to: identify, by the computer, a plurality of content sources each comprising textual content, wherein the plurality of content sources are ordered by interest to a user; convert, by the computer, each content source of the plurality of content sources from textual content to speech comprising an individual audio selection of a plurality of audio selections; apply, by the computer, one or more voice properties to each audio selection to specify audio properties that select a respective position of a respective audio selection from a plurality of positions within a multi-dimensional sound space and audibly distinguish one or more characteristics of the respective audio selection from other audio selections of the plurality of audio selections, wherein the respective position of the respective audio selection reflects a ranking of the respective audio selection ordered by interest; and output, by the computer, a parallel stream of the plurality of audio selections to an audio output device for stereophonic playback of the plurality of audio selections within the plurality of positions within the multi-dimensional sound space to the user, wherein the plurality of positions reflect the plurality of content sources ordered by interest; further comprising program instructions executed by the computer to cause the computer to: monitor, by the computer, content accessed by a user for an indication of a selection of the content by the user; in response to detecting the indication of the selection of the content by the user, record, by the computer, the indication of the selection and the content in a content history log; analyze, by the computer, the content history log for a frequency of at least one type of content; generate, by the computer, at least one topic preference for the user based on the at least one type of content and the respective frequency; train, by the computer, a cognitive model with the at least one topic preference and the respective frequency to identify other content comprising the at least one topic; and identify, by the computer, the other content comprising the at least one topic. in response to the user accessing a website comprising a plurality of text passages, identifying, by the computer, a plurality of content sources from a selection of the plurality of text passages comprising the at least one topic by a cognitive model; and ordering, by the computer, the plurality of content sources according to respective priorities identified by a cognitive model based on respective frequencies of the at least one topic present in respective content sources of the plurality of content sources.
16. The computer program product of claim 15, further comprising program instructions for execution by a computer to cause the computer to: identify, by the computer, an interface comprising a plurality of text passages available from a content aggregator; and divide, by the computer, the plurality of text passages from the content aggregator into separate content sources in the plurality of content sources based on an output interface setting specifying at least one graphical characteristic used to identify boundaries between text passages in the interface.
17. The computer program product of claim 15, further comprising program instructions for execution by a computer to cause the computer to: apply, by the computer, the one or more speech properties to each audio selection to distinguish the one or more characteristics, the one or more speech properties comprising one or more of one or more speech type properties, one or more speech volume properties, and one or more speech speed properties.
Citation Information
Patent Citations
Audio program player including a dynamic program selection controller
US6199076B1
Interactive audio system
US20020103554A1
Spatialized audio enhanced text communication and methods
US20070027691A1