Selecting a primary source for pose-based text-to-speech conversion
By generating concurrent audio streams with distinct attributes in a multidimensional space and adjusting based on user posture, the method enhances the efficiency of identifying topics of interest in multiple text content items, mimicking visual scanning.
Patent Information
- Application Number
- JP2022573559
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-24
- Filing Date
- 2021-06-09
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-06-09
AI Technical Summary
Existing text-to-speech technologies require users to listen to each segment individually, which is inefficient for quickly identifying topics of interest in multiple text content items, and 3D audio playback does not support conversion from text to speech or adjust sound sources for focused viewing.
Generate multiple concurrent audio streams from text passages with distinct audio attributes in a multidimensional acoustic space, allowing users to auditorily scan and focus on specific content using postural movements.
Facilitates efficient auditory scanning and focused viewing of multiple text content items by positioning audio selections based on user interest and adjusting characteristics dynamically, mimicking a visual scanning experience.
Smart Images

Figure 0007714303000001 
Figure 0007714303000002 
Figure 0007714303000003
Abstract
Description
Technical Field
[0001] One or more embodiments of the present invention generally relate to data processing, and more particularly to selecting a primary source of text-to-speech conversion based on posture.
Background Art
[0002] Many computing devices support text-to-speech technology that converts text content into speech output via a speaker or other audio output device. When a user opens a website containing multiple items of text content aggregated from multiple sources, such as a news website or search engine results, and selects to listen to the text content from that website converted using text-to-speech, the text-to-speech converter begins converting the text from a single text string within the website into speech. Requiring the user to individually listen to the speech converted from each item of text content visible within the website in order to find topics of interest is less efficient than the visual scanning that allows the user to quickly scan across keywords in the text of various headings or other text content items simultaneously displayed on one page, or quickly decide which item of text content to focus on and read in detail.
[0003] U.S. Patent No. 6,199,076 to Logan et al. describes an audio program player that "automatically plays back a predetermined schedule of audio program segments" and supports the user's "desired ability to listen to audio voice announcements describing each segment before it is played, and to skip forward or backward to the next or previous announcement, thereby immediately obtaining the information necessary to determine whether a given segment is of interest." Additionally, Logan et al. describe a feature that identifies one or more distinct passages within any program segment as "highlighted," and the player incorporates means for operating the player in a "play highlights" mode that skips over those portions of content that are not highlighted, thereby enabling the viewer to review only key points of the presentation or more quickly identify particular passages of interest within the body of a particular program. Additionally, Logan et al. explain that "information available in text format from news sources, libraries, etc. may be converted into a condensed audio format by a human reader or by conventional speech synthesis. If speech synthesis is used, the conversion of text to speech is preferably performed by the player at the client station 103." However, the audio program player of Logan et al. still requires the user to listen to an audio voice announcement for each segment or highlights from the text-to-speech converted segments, which is not efficient for allowing a user to quickly visually scan through a display of multiple text content items displayed simultaneously on a single page and select a particular text content item for further review.
[0004] Craig Jin's paper, "Real-Time Head-Tracking 3D Audio with Unlimited Concurrent Sound," describes how "a novel method using real-time head tracking for 3D audio playback for an unlimited number of simultaneous sound sources is proposed. By performing 3D audio processing offline, an unlimited number of simultaneous sound sources can be spatially rendered. Smooth playback with real-time head tracking is achieved by rendering binaural audio tracks for a set of fixed, closely spaced head poses." However, Jin's 3D audio playback for an unlimited number of simultaneous sound sources is limited to preloaded existing sound sources, typically movie soundtracks, and does not identify or convert any content sources from text to speech. In addition, Jin's 3D audio playback for an unlimited number of simultaneous sound sources does not provide for adjusting the sound sources, allowing the user to quickly scan all simultaneous sound sources and then select one to focus on.
[0005] As the number of online aggregated content providers continues to grow, there is a need for methods, systems, and computer programs to support improved efficiency for users navigating multiple textual content items from content aggregators when the textual content items are converted from text to speech. Additionally, as stereo headphones and speakers for listening to multidimensional audio become more common, there is a need for methods, systems, and computer program products that support improved tuning of audio sources for output in a multidimensional acoustic space, facilitating users to listen to multiple audio sources simultaneously, identify topics of interest to the user from the multiple audio sources, and then selectively adjust the output characteristics of specific audio sources to facilitate focused viewing while other audio sources are still presented. Summary of the Invention
[0006] In light of the above, one or more embodiments of the present invention provide for generating multiple concurrent streams of audio selections generated from multiple text passages converted to speech and played in parallel, while applying different audio attributes to each audio selection to auditorily distinguish each audio selection from the other audio selections, thereby providing an audio output that creates a space filled with different voices that facilitates a user to efficiently auditorily scan through the content of the multiple text content items converted to speech.One or more embodiments of the present invention provide for generating multiple concurrent streams of audio selections generated from multiple text passages converted to speech, while applying different audio attributes to each audio selection to auditorily locate each audio selection at a different location within a multidimensional acoustic space, thereby facilitating a user to scan through the content of the multiple audio selections presented in parallel at different auditory locations, thereby simulating a space filled with different speakers available to a user to identify topics of interest to the user from among the audio selections. Additionally, in light of the above, one or more embodiments of the present invention provide for monitoring user postural movements relative to the location of an audio selection within a multi-dimensional acoustic space as an indication of a user preference to focus on a particular audio selection converted from a source of textual content, and selectively adjusting output characteristics of the particular audio selection to facilitate focused viewing on the particular audio selection while continuing to play other audio selections in the multi-dimensional acoustic space, thereby providing an auditory environment for output of the aggregated textual content converted to speech that is as efficient and easy to understand as the visual environment for outputting the aggregated textual content, while allowing the user to control focused viewing using postural movements.
[0007] In one or more embodiments, the method is directed to steps by a computer system of identifying a plurality of content sources, each including text content, where the plurality of content sources are ordered by interest to the user. The method is directed to steps by a computer system of converting each of the plurality of content sources from text content into audio including distinct audio selections of a plurality of audio selections. The method is directed to steps by a computer system of applying one or more audio attributes to each audio selection to specify an audio attribute that selects a respective position of each audio selection from among a plurality of positions in a multi-dimensional acoustic space and aurally differentiates one or more characteristics of each audio selection from other audio selections of the plurality of audio selections, where the respective position of each audio selection reflects the rank of the respective audio selection ordered by interest. The method is directed to steps by a computer system of outputting a simultaneous stream of the plurality of audio selections to an audio output device for stereophonic playback of the plurality of audio selections to the user at a plurality of positions in a multi-dimensional acoustic space that reflects the plurality of content sources ordered by interest. One advantage of applying one or more audio attributes to each audio selection to specify an audio attribute that selects a respective position of each audio selection from among a plurality of positions in a multi-dimensional acoustic space and aurally differentiates one or more characteristics of each audio selection from other audio selections of the plurality of audio selections is to provide an audio output that supports a user scanning through the content of a plurality of audio selections presented in parallel at different audible positions detectable to the user and identifying a topic of interest to the user from among the audio selections.One advantage of identifying a plurality of content sources ordered by user interest, specifying audio attributes, and selecting each position of an audio selection that reflects the rank of each audio selection ordered by interest is that a simultaneous stream of audio selections provides an audio output that places the content source that the user is most likely to be interested in at one or more primary positions, where these one or more primary positions improve the predictability for the user that the audio selection at the one or more primary positions will have the highest interest for the user when scanning the auditory positions.
[0008] In addition, in one or more embodiments, a method, by a computer system, of identifying a plurality of content sources each including text content, the plurality of content sources being ordered by interest to a user, the method of identifying may be directed to the step of, by the computer system, monitoring content accessed by a user as an indication of the user's selection of content. The method may further be directed to the step of, in response to detecting an indication of the user's selection of content, by the computer system, recording the indication of selection and the content in a content history log. The method may further be directed to the step of, by the computer system, analyzing the content history log for the frequency of at least one type of content. The method may further be directed to the step of, by the computer system, generating at least one topic preference for the user based on at least one type of content and its respective frequency. The method may be directed to the step of, by the computer system, training a cognitive model using at least one topic preference and its respective frequency to identify other content including at least one topic. The method may be directed to the step of, in response to the user accessing a website including a plurality of text messages, by the computer system, identifying a plurality of content sources via the cognitive model from the selection of a plurality of text messages including at least one topic. The method may further be directed to the step of, by the computer system, ordering the plurality of content sources according to respective priority levels identified by the cognitive model based on the respective frequency of at least one topic present in each of the plurality of content sources of the plurality of content sources.One advantage of training a cognitive model using topic preferences and frequencies based on the activities of monitored users who access and respond to content, and then selectively identifying and ordering content sources by the trained cognitive model, is that the selection of content sources played in parallel in a simultaneous stream is ordered in an auditory position according to the likelihood of interest to a particular user, improving the efficiency for a particular user to scan through multiple audio selections streamed in parallel and find at least one audible selection of interest to the user.
[0009] Additionally, in one or more embodiments, the method steps for identifying, by a computer system, a plurality of content sources each containing text content, the plurality of content sources being ordered by interest to a user, may be directed to identifying, by the computer system, an interface containing a plurality of text passages available from a content aggregator. The method may further be directed to dividing, by the computer system, the plurality of text passages from the content aggregator into distinct content sources of the plurality of content sources based on an output interface configuration specifying at least one graphical characteristic for identifying boundaries between the text passages in the interface. One advantage of identifying multiple text passages available from a content aggregator and dividing the multiple text passages into content sources based on an output interface configuration that specifies at least one graphical characteristic to identify boundaries between the text passages in the interface is that the content sources are identified independently of any particular content aggregator interface, allowing simultaneous streams to be generated for content sources from multiple types of content aggregators and independently of whether the content aggregator outputs the text passages in its interface for conversion to speech.
[0010] In addition, in one or more embodiments, a computer system applies one or more audio attributes to each audio selection to specify an audio attribute that selects a respective position of each audio selection from among a plurality of positions in a multi-channel audio space, and to aurally distinguish one or more characteristics of each audio selection from other audio selections of the plurality of audio selections, wherein the respective position of each audio selection reflects the rank of each audio selection ordered by interest. The step of the method of distinguishing may be targeted at the step of a computer system applying one or more audio attributes to each audio selection to specify an audio attribute for a particular position among a plurality of positions at a primary position in front of the user in a multi-channel audio space for a particular audio selection converted from a particular content source having the highest rank ordered by interest. One advantage of applying one or more audio attributes to specify an audio attribute for a particular audio selection having the highest rank ordered by interest at a primary position in front of the user in a multi-channel audio space is to improve the efficiency when the user determines an order to scan and focus on each of the audio selections that are played in parallel while the audio selection of most interest to the user is predictably positioned at a position in front of the user along with a plurality of audio selections at different positions.
[0011] Additionally, in one or more embodiments, the method may be directed to the computer system monitoring for an input indicative of a user posture change. The method may also be directed to the computer system analyzing, in response to detecting an input indicative of a user posture change, a direction of movement of the user posture change relative to the positioning of a plurality of audio selections within a multi-dimensional audio space including a 3D audio space. The method may also be directed to the computer system identifying a particular audio selection from among the plurality of audio selections selected by the user based on the direction of movement of the user posture change relative to the positioning of the plurality of audio selections in the 3D audio space. The method may also be directed to the computer system specifying an audio attribute that applies one or more updated audio attributes to each audio selection to select a respective position of each audio selection from among a plurality of positions within the 3D audio space and aurally differentiating one or more characteristics of each audio selection from other audio selections of the plurality of audio selections, wherein the particular audio selection moves to a primary position in front of the user within the 3D audio space and the volume of the particular audio selection is increased relative to other audio selections of the plurality of audio selections. One advantage of monitoring for an input indicative of a user posture change and detecting a direction of movement in the user posture change relative to the positioning of a plurality of audio selections in a multi-dimensional audio space is to provide an input for the user to focus on a particular audio selection with a minimal movement that enables the user to select and focus on a particular audio selection from among a plurality of audio selections within a simultaneous stream without hands and without requiring precise movement.
[0012] In one or more embodiments, a computer system includes one or more processors, one or more computer-readable memories, one or more computer-readable storage devices, and program instructions stored in at least one of the one or more computer-readable storage devices and executable by at least one of the one or more processors via at least one of the one or more memories. The stored program instructions include program instructions that identify a plurality of content sources, each including text content, where the plurality of content sources are ordered by interest to the user. The stored program instructions include program instructions that convert each of the plurality of content sources from text content into audio including separate audio selections of a plurality of audio selections. The stored program instructions include program instructions that specify audio attributes that apply one or more audio attributes to each audio selection to select a respective position of each audio selection from among a plurality of positions in a multi-dimensional acoustic space and auditorily distinguish one or more characteristics of each audio selection from other audio selections of the plurality of audio selections, where the respective position of each audio selection reflects the rank of each audio selection ordered by interest. The stored program instructions include program instructions that output a simultaneous stream of the plurality of audio selections to an audio output device for stereophonic playback of the plurality of audio selections to the user at a plurality of positions in a multi-dimensional acoustic space that reflects the plurality of content sources ordered by interest.Specify audio attributes that apply one or more audio attributes to each audio selection to select a respective position of each audio selection from among a plurality of positions in a multi-channel audio space, and one advantage of aurally distinguishing one or more characteristics of each audio selection from other audio selections of the plurality of audio selections is that a simultaneous stream of audio selections provides an audio output that supports a user scanning through the content of a plurality of audio selections presented in parallel at different audible positions that are detectable to the user and identifying topics of interest to the user from among the audio selections. One advantage of identifying a plurality of content sources ordered by a user's interest and specifying audio attributes to select a respective position of each audio selection that reflects the rank of each audio selection ordered by interest is that a simultaneous stream of audio selections provides an audio output that positions the content source that the user is most likely to be interested in at one or more primary positions, where this one or more primary positions improves predictability for the user in that audio selections at the one or more primary positions will have the highest interest for the user when scanning audible positions.
[0013] Additionally, in one or more embodiments, the stored program instructions may be directed to identifying multiple content sources, each containing textual content, the multiple content sources being ordered by user interest. The stored program instructions may be directed to monitoring content accessed by a user for an indication of a content selection by the user. The computer system may further include stored program instructions for recording the indication of selection and the content in a content history log in response to detecting an indication of content selection by the user. The computer system may further include stored program instructions for analyzing the content history log for a frequency of at least one type of content. The computer system may further include stored program instructions for generating at least one topic preference for the user based on the at least one type of content and its respective frequency. The computer system may further include stored program instructions for training a cognitive model using the at least one topic preference and its respective frequency to identify other content containing the at least one topic. The computer system may further include stored program instructions for, in response to a user accessing a website including a plurality of text passages, identifying, via a cognitive model, a plurality of content sources from a selection of the plurality of text passages including at least one topic. The computer system may further include program instructions for ordering the plurality of content sources according to respective priority levels identified by the cognitive model based on a respective frequency of the at least one topic present in each of the plurality of content sources.One advantage of training a cognitive model using topic preferences and frequencies based on the monitored user activities that access and respond to content, and then selectively identifying and ordering content sources by the trained cognitive model, is that the selection of content sources that are played in parallel in a simultaneous stream is ordered in an auditory position according to the likelihood of interest to a particular user, improving the efficiency for a particular user to scan through multiple audio selections streamed in parallel and find at least one audible selection of interest to the user.
[0014] In addition, in one or more embodiments, the stored program instructions may be directed to program instructions that identify a plurality of content sources each including text content, where the plurality of content sources are ordered by interest to a particular user, and identify an interface including a plurality of text messages available from a content aggregator. The computer system may further include stored program instructions to split a plurality of text messages from a content aggregator into separate content sources of a plurality of content sources based on an output interface setting that specifies at least one graphical characteristic to identify a boundary between text messages in the interface. One advantage of identifying a plurality of text messages available from a content aggregator and splitting the plurality of text messages into content sources based on an output interface setting that specifies at least one graphical characteristic to identify a boundary between text messages in the interface is that the content sources are identified independently of any particular content aggregator interface and the simultaneous stream is generated for content sources from multiple types of content aggregators and independently of whether the content aggregator outputs text messages in the interface for conversion to speech.
[0015] Additionally, in one or more embodiments, the stored program instructions may be directed to applying one or more audio attributes to each audio selection to assign audio attributes to select a respective position of each audio selection from among a plurality of positions in the multidimensional acoustic space, audibly distinguishing one or more characteristics of each audio selection from other audio selections of the plurality of audio selections, each position of each audio selection reflecting a rank of each audio selection ordered by interest. The stored program instructions may also be directed to applying one or more audio attributes to each audio selection to assign audio attributes to a particular position of a plurality of positions in a prime position in front of a user in the multidimensional acoustic space for a particular audio selection converted from a particular content source having the highest rank ordered by interest. One advantage of applying one or more audio attributes to specify the audio attributes for a particular audio selection that has the highest rank, ordered by interest, at a primary location in front of the user in a multi-dimensional acoustic space is that the audio selections that are likely to be of most interest to the user are predictably positioned in a location in front of the user, while the user is provided with an audio output with multiple audio selections at different locations, improving the efficiency of the user in determining the order in which to scan and focus each of the audio selections that are played in parallel.
[0016] Additionally, in one or more embodiments, the computer system may further include stored program instructions that monitor input indicative of a user posture change. In response to detecting the input indicative of a user posture change, the computer system may further include stored program instructions that analyze a direction of movement of the user posture change relative to the positioning of the multiple audio selections within a multi-dimensional acoustic space, including a 3D acoustic space. The computer system may further include stored program instructions that identify a particular audio selection from the multiple audio selections selected by the user based on a direction of movement of the user posture change relative to the positioning of the multiple audio selections in the 3D acoustic space. The computer system may further include stored program instructions for applying one or more updated audio attributes to each audio selection, specifying audio attributes that select a respective location of the respective audio selection from among a plurality of locations within the 3D acoustic space, and audibly distinguishing one or more characteristics of each audio selection from other audio selections of the plurality of audio selections, wherein the particular audio selection is moved to a primary location in front of the user within the 3D acoustic space and a volume of the particular audio selection is increased relative to other audio selections of the plurality of audio selections. One advantage of monitoring input indicative of user posture changes and detecting direction of movement in user posture changes relative to positioning of multiple audio selections in a multi-dimensional acoustic space is that it provides an input for focusing on a particular audio selection with minimal movement, allowing a user to select and focus on a particular audio selection from multiple audio selections in simultaneous streams hands-free and without requiring precise movement.
[0017] In one or more embodiments, a computer program product includes a computer-readable storage medium having program instructions embodied therein. The computer-readable storage medium is not a transient signal per se. The program instructions are executable by a computer to cause the computer to identify a plurality of content sources each including text content, the plurality of content sources being ordered by interest to the user. The program instructions are executable by a computer to cause the computer to convert each of the plurality of content sources from text content to an audio including separate audio selections of a plurality of audio selections. The program instructions are executable by a computer to cause the computer to apply one or more audio attributes to each audio selection to specify an audio attribute that selects a respective position of each audio selection from among a plurality of positions in a multi-channel audio space, and to aurally distinguish one or more characteristics of each audio selection from other audio selections of the plurality of audio selections, wherein the respective position of each audio selection reflects the rank of each audio selection ordered by interest. The program instructions are executable by a computer to cause the computer to output a simultaneous stream of the plurality of audio selections to an audio output device for stereo playback of the plurality of audio selections to the user at the plurality of positions in the multi-channel audio space that reflect the plurality of content sources ordered by interest.One advantage of applying one or more audio attributes to each audio selection to specify audio attributes that select a respective location of each audio selection from among multiple locations in a multidimensional acoustic space and auditorily distinguish one or more characteristics of each audio selection from others of the multiple audio selections is that the concurrent stream of audio selections provides an audio output that supports a user scanning through the content of the multiple audio selections presented in parallel at different auditory locations detectable to the user and identifying topics of interest to the user from among the audio selections. One advantage of identifying multiple content sources ordered by user interest and specifying audio attributes to select respective locations of the audio selections that reflect the rank of each audio selection ordered by interest is that the concurrent stream of audio selections provides an audio output that places content sources that are most likely to be of interest to the user in one or more primary locations, where the one or more primary locations improve the user's predictability that the audio selections in the one or more primary locations will be of greatest interest to the user as they scan the auditory locations.
[0018] In addition, in one or more embodiments, a computer program product causes a computer to identify a plurality of content sources each including text content, the plurality of content sources being ordered by interest to a user, and causes the computer to monitor, by the computer, content accessed by a user as an indication of the user's selection of content. The computer program product may be targeted at program instructions that are executable by a computer and cause the computer to monitor, by the computer, content accessed by a user as an indication of the user's selection of content. The computer program product may include program instructions that are executable by a computer and cause the computer to record, in response to detecting an indication of a user's selection of content, the indication of the selection and the content in a content history log. The computer program product may include program instructions that are executable by a computer and cause the computer to analyze, by the computer, the content history log for the frequency of at least one type of content. The computer program product may include program instructions that are executable by a computer and cause the computer to generate, by the computer, at least one topic preference for a user based on at least one type of content and its respective frequency. The computer program product may include program instructions that are executable by a computer and cause the computer to train, by the computer, a cognitive model using at least one topic preference and its respective frequency to identify other content including at least one topic. The computer program product may include program instructions that are executable by a computer and cause the computer to identify, via the cognitive model, a plurality of content sources from a selection of a plurality of text passages including at least one topic in response to a user accessing a website including a plurality of text passages.A computer program product may include computer-executable program instructions that, when executed by a computer, cause the computer to order a plurality of content sources according to respective priority levels identified by a cognitive model based on the respective frequencies of at least one topic present in each of the plurality of content sources. One advantage of training a cognitive model using topic preferences and frequencies based on the monitored user activities of accessing and responding to content and then selectively identifying and ordering content sources by the trained cognitive model is that the selection of content sources to be played in parallel in a simultaneous stream is ordered in an auditory position according to the likelihood of interest to a particular user, improving the efficiency for a particular user to scan through a plurality of audio selections streamed in parallel and find at least one audible selection of interest to the user.
[0019] Additionally, in one or more embodiments, the computer program product may include program instructions executable by a computer to cause a computer to identify a plurality of content sources each containing text content, the plurality of content sources being ordered by interest to a user, and may be directed to stored program instructions that, when executed by a computer, cause the computer to identify an interface containing a plurality of text passages available from a content aggregator. Additionally, in one or more embodiments, the computer program product may include program instructions executable by a computer to cause the computer to split the plurality of text passages from the content aggregator into separate content sources of the plurality of content sources based on an output interface configuration that specifies at least one graphical characteristic to identify boundaries between the text passages in the interface. One advantage of identifying multiple text passages available from a content aggregator and dividing the multiple text passages into content sources based on an output interface configuration that specifies at least one graphical characteristic to identify boundaries between the text passages in the interface is that the content sources are identified independently of any particular content aggregator interface, allowing simultaneous streams to be generated for content sources from multiple types of content aggregators and independently of whether the content aggregator outputs the text passages in its interface for conversion to speech.
[0020] In addition, in one or more embodiments, a computer program product is executable by a computer and includes stored program instructions that, by the computer, apply one or more audio attributes to each audio selection to specify audio attributes for a particular position among a plurality of positions at a primary position in front of the user in a multi-channel acoustic space for a particular audio selection converted from a particular content source having the highest rank ordered by interest. One advantage of applying one or more audio attributes to specify audio attributes for a particular audio selection having the highest rank ordered by interest at a primary position in front of the user in a multi-channel acoustic space is that it improves the efficiency when the user determines the order to scan and focus on each of the audio selections that are provided as audio output and played in parallel with a plurality of audio selections at different positions while the audio selection that is most likely to be of most interest to the user is predictably positioned at a position in front of the user.
[0021] The novel features believed to be characteristic of one or more embodiments of the invention are set forth in the appended claims. The invention itself, however, as well as other objects and advantages thereof, will be best understood by reference to the following detailed description of illustrative embodiments when read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0022]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
DETAILED DESCRIPTION OF THE INVENTION
[0023] In the following description, for the sake of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be apparent to one of ordinary skill in the art that the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form in order to avoid unnecessarily obscuring the present invention.
[0024] In addition, in the following description, for the sake of convenience, numerous systems are described. It should be noted that the present invention may be implemented in various computer systems, including computer systems and electronic devices that operate any number of different types of operating systems, and this will be apparent to one of ordinary skill in the art.
[0025] FIG. 1 shows a block diagram of an example listening service that implements a simultaneous audio playback controller guided by user posture information to place a text-to-speech converted audio selection in a primary location from among multiple content sources provided by a content aggregator and audio playback of a specific content source from among multiple content sources that are converted to text-to-speech and played in parallel.
[0026] In one example, listening service 110 provides a service for controlling the audio output stream of one or more content sources available from content aggregator 114. In one example, content aggregator 114 includes textual content representing multiple content sources, such as articles published on a news website or in a news feed, social media entries published on a social media website, search results aggregated on a search results website, entries in an email, entries in a text chain, and other content aggregations that may be identified as independent content items accessible to a user from an application, website, or other service. In one example, content aggregator 114 includes textual content. In another example, content aggregator 114 includes graphical and audio content in addition to textual content.
[0027] In one example, the embodiment of FIG. 1 provides an advantage of the listening service 110, which is useful for controlling the audio output stream of one or more content sources available from the content aggregator 114, regardless of whether the output content is intended by the content aggregator 114 to be output in the form of an audio output stream. Advantageously, the listening service 110 evaluates the web page or other interface provided by the content aggregator 114, evaluates the presence of one or more types of depictions between individual textual content items within the web page or other interface, and includes a simultaneous audio playback controller 112 for identifying the content source from the depicted textual content item.
[0028] In one example, the embodiment of FIG. 1 provides an advantage that the simultaneous audio playback controller 112 also manages the output of a simultaneous stream 116 of multiple audio selections generated from multiple content sources available from the content aggregator 114, facilitating the user to efficiently review the content of multiple content sources in parallel. Advantageously, the simultaneous audio playback controller 112 outputs a simultaneous stream 116 that simultaneously includes each audio track of audio selections 132, 134, and 136, each of which is converted from a different content source available from the content aggregator 114 into audible audio. In additional or alternative embodiments, the simultaneous stream 116 includes additional or alternative audio selections.
[0029] In this example, before outputting on the simultaneous audio stream 116, the simultaneous audio playback controller 112 converts a text-based content source from the content aggregator 114 into audio to be output as an audible selection. The simultaneous audio playback controller 112 may implement one or more types of text-to-speech convertors that use one or more selectable conversion settings for converting text-based items into audio.
[0030] In addition, the embodiment of FIG. 1 provides the advantage that the simultaneous audio playback controller 112 manages the output of the simultaneous stream 116 as stereo audio using three-dimensional (3D) positional audio adjustment for each of the multiple audio selections, creating the effect that each audio selection is placed in a different region of the 3D acoustic space 130 and facilitating the user to audibly distinguish the content of multiple content sources in parallel. In this example, the 3D acoustic space 130 represents an audible output of the recognized position of sound that can be detected by a user listening to the simultaneous stream 116 as an output by a stereo output device such as stereo headphones worn by the user or stereo speakers broadcasting within the region where the user is located. In one example, the stereo output device represents an audio output device that supports an audio surround sound effect that can be recognized as being placed in different locations within the 3D acoustic space around the user by simultaneously playing different voices from different individual speaker units. In another embodiment, the simultaneous stream 116 is output by an audio output device that supports audio output in other dimensions, including but not limited to one-dimensional space or two-dimensional space.
[0031] For example, the simultaneous audio playback controller 112 may manage the output of the simultaneous stream 116 with 3D position audio adjustment for the audio selection 132 to be recognized to the left of the user, 3D position audio adjustment for the audio selection 134 to be recognized at the center of the user, and 3D position audio adjustment for the audio selection 136 to be recognized to the right of the user. In an additional or alternative example, the 3D position audio adjustment may place the audio selection at a position above the user, below the user, or at other recognizable positions within a 3D region recognizable by the user in the 3D acoustic space 130.
[0032] In one example, the embodiment of FIG. 1 provides the advantage that the simultaneous audio playback controller 112 manages additional audio attribute tuning to distinguish the characteristics of the content sources output simultaneously as the audio selection 132, the audio selection 134, and the audio selection 136. For example, the simultaneous audio playback controller 112 adjusts audio attributes that distinguish characteristics such as speed, dub, tone, and volume to distinguish the content sources output as the audio selection 132, the audio selection 134, and the audio selection 136.
[0033] Advantageously, when presented with audio including multiple voices with different characteristics and speaking from different locations, the simultaneous audio playback controller 112 generates simultaneously streamed audio selections using multiple voices within the 3D acoustic space 130 that can be effectively and efficiently navigated by a user due to the user's focused auditory attention and the effect of selective hearing in the brain, also known as the cocktail party effect. In one example, based on a user's ability to focus their auditory attention and selectively listen, a person can tune into a single speaker or stimulus while ignoring a crowd of speakers or filtering out various other stimuli, such as when a user can focus on a single conversation in a noisy room. Based on a person's ability to selectively listen, a person has the ability to separate different speakers into different streams and determine which stream is most relevant to the user, thereby allowing the person to tune into one voice or tune out other voices. A person's brain may have the ability to separate different auditory stimuli into different streams and determine which stream is most relevant. The benefit of the simultaneous voice playback controller 112 coordinating the position and characteristics of different audio selections within the 3D acoustic space 130 is that the output of the simultaneous streams 116 simulates a crowd of speakers with different voice types, volumes, speaking rates, and speaking from different positions, supporting the user to listen to all the voices in parallel and then, through selective listening, focus on a particular voice for enhanced cognitive focus while tuned out other voices.
[0034] In addition to outputting the simultaneous stream 116 at the characteristics and positions of each audio selection adjusted to facilitate the user to take selective listening and listen attentively to a specific audio selection within the simultaneous stream 116, advantageously, the simultaneous audio reproduction controller 112 dynamically adjusts the characteristics and positions of each audio selection in response to the user's posture related to the audio selection within the 3D acoustic space 130. For example, the simultaneous audio reproduction controller 112 receives the input of the user posture information 120, analyzes the positioning of the audio selection within the 3D acoustic space 130 in relation to the user posture information 120, determines the specific audio selection towards which the user posture is directed, and adjusts the characteristics and positions of each audio selection within the simultaneous stream 116 to control the specific audio selection as the primary source to the primary position. In one example, the user posture information 120 represents the movement of the head detected by a gyroscope or other type of motion sensor within a stereo headphone set worn by the user. Advantageously, the simultaneous audio reproduction controller 112 receives the user posture information 120 indicating the direction of movement for indicating the user selection of the audio selection, but does not require specific precise movements for directly selecting a specific audio selection to be placed at the primary position within the 3D acoustic space 130.
[0035] In one embodiment, the listening service 110 evaluates the direction of the user's head movement identified in the user posture information 120 towards a specific audio selection within the 3D acoustic space 130 as an input that triggers the simultaneous audio reproduction controller 112 to select which audio selection is placed at the primary position within the 3D acoustic space 130. In another embodiment, the listening service 110 may evaluate other gestures or posture changes by the user indicated in the user posture information 120 as an input for causing the listening service 110 to switch to different content aggregators, where the listening service 110 may manage the reproduction of simultaneous streams from multiple content aggregators.
[0036] The advantage that the listening service 110 supports the conversion of text to audio from a plurality of content sources identified from the content aggregator 114 is that the user can receive the audio of multiple voices of the audio converted from a plurality of text passages in parallel, thereby providing efficient audio-based browsing of the plurality of content sources provided by the content aggregator 114. The advantage that the listening service 110 manages the adjustment of specific audio selections for playback at a higher volume at the main position within the 3D acoustic space 130 in response to the user posture information 120 is that selective listening processing of the user is supported in response to the minimum effort required by the user to turn their head in the direction of the content of interest to the user, thereby enabling the user to control selective viewing in a hands-free environment.
[0037] FIG. 2 shows a block diagram of an example of a computing device that implements a simultaneous audio playback controller guided by user posture information from a gyroscope so that attention can be directed to the audio selection placed at the main position among a plurality of content sources converted from text to audio and played in parallel.
[0038] In one example, computing device 200 includes a display interface 210 in which content sources from content aggregator 114 are displayed as selectable content items, such as content item A 212, content item B 214, content item C 216, and content item D 218. In one embodiment, content item A 212, content item B 214, content item C 216, and content item D 218 each represent an independent content item identifiable as a content source by simultaneous audio playback controller 112. In one example, content item A 212, content item B 214, content item C 216, and content item D 218 are displayed as aggregated text content in display interface 210 that is individually selectable by a user to visually read the text content in its entirety. Alternatively, content item A 212, content item B 214, content item C 216, and content item D 218 may each be individually selectable by a user to hear audio converted from the text content. In another embodiment, one or more of content item A212, content item B214, content item C216, and content item D218 are selectable within display interface 210 and represent audio-based content items that do not require initial conversion from text to speech in order to be output in an audible format.
[0039] In one embodiment, the user may select each of the selectable content items within the display interface 210 to read or play by selecting an item within the display interface 210. Additionally, according to the advantages of one embodiment, the user may select the audio playback selector 220 within the display interface 210 to trigger the simultaneous audio playback controller 112 of the listening service 110, or provide another type of input to trigger the simultaneous audio playback controller 112 of the listening service 110, to select to simultaneously play each of the distinguishable content items within the display interface 210. The advantage of the user selecting to simultaneously play each distinguishable content item within the display interface 210 by selecting to trigger the simultaneous audio playback controller 112 is that the user can visually review each content item within the display interface 210 with the same efficiency, and enables the user to auditorily review each content item in parallel and select to pay more attention to a specific content item and listen to it.
[0040] In one example, the computing device 200 controls the output of an audio stream to a user-accessible stereo audio output device that supports the output of a 3D audio stream via the input / output (I / O) interface 222. In one embodiment shown in FIG. 2, the computing device 200 controls the output of an audio stream to the stereo headphones 240 worn on the user's user head 244 via the I / O interface 222. In one example, the I / O interface 222 may output the audio stream through a wired connection between the I / O interface 222 and the stereo headphones 240, or through a wireless connection between the I / O interface 222 and the stereo headphones 240.
[0041] In one example, the simultaneous audio playback controller 112 initially outputs simultaneous streams to the 3D acoustic space 230 with content item A 212 output as audio selection 232 at "Position #1," content item B 214 output as audio selection 234 at "Position #2," content item C 216 output as audio selection 236 at "Position #3," and content item D 218 output as audio selection 238 at "Position #4." In this example, "Position #3" is considered to be the primary position in front of the user, with "Position #1" and "Position #2" to the left of the center position, and "Position #4" to the right of the center position. Advantageously, the simultaneous audio playback controller 112 selects the location of each content source when converted to output as an audio selection based on a ranked order of user interest, placing the audio selections likely to be of most interest to the user in a primary position in front of the user, and supporting the user to apply selective listening to efficiently listen to all audio selections in parallel while placing the audio selections likely to be of most interest in a listening position in front of the user. In one example, the ordering of content items in the display interface 210 is determined by the simultaneous audio playback controller 112, and content item C 216 is determined as the highest-ranked content source and is initially placed in position #3, the primary position, different from the ranking ordered by user interest applied to the selection of content item locations. Alternatively, the ordering of content items in the display interface 210 may reflect a ranking ordered by user interest.
[0042] In one embodiment, the I / O interface 222 supports an interface for detecting an output from a stereo headset 240 or other output device that detects a user's posture change and broadcasts or transmits the user's posture change as user posture information 120. In one embodiment, the stereo headset 240 incorporates a gyroscope 246 that detects changes in the movement of the user's head 244 by measuring rotational movement and reports the change in movement as user posture information 120. In other embodiments, other sensors worn or not worn by the user may monitor changes in the movement of the user's head 244 or other gesture movements by the user and report the change in movement as user posture information 120. In one embodiment, the stereo headset 240 incorporates a gyroscope 246 for detecting changes in the posture of the user's head to dynamically adjust the actual position selected for positioning audio within the 3D acoustic space 230 relative to the position of the user's head, and at the same time, the simultaneous audio playback controller 112 receives the detected change in movement as user posture information 120.
[0043] In one example, the gyroscope 246 detects the rotational movement of the user's head 244 directed to the left, as indicated by the direction of rotational movement at reference numeral 242. In one example, the gyroscope 246 includes one or more sensors that measure angular velocity in units of degrees per second or revolutions per second by monitoring rotation about one or more axes, including but not limited to, for example, three axes. In one example, each sensor within the gyroscope 246 has a resonant mass that shifts as a change in angular velocity, and the movement of the resonant mass is converted into an electrical signal indicative of a change in degrees per second or revolutions per second.
[0044] In this example, in response to I / O interface 222 detecting output from gyroscope 246 as user posture information 120, simultaneous audio playback controller 112 evaluates user posture information 120 relative to the position of the audio selection within 3D acoustic space 230 to determine how to adjust the position of the audio selection within 3D acoustic space 230 based on the user input. In one example, audio selection 236 at "Position #3" is perceived by the user in a primary, central position within 3D acoustic space 230, and simultaneous audio playback controller 112 analyzes a movement of the user's head to the left, indicating user selection of audio selection 234. In response to analyzing the user selection of audio selection 234, simultaneous audio playback controller 112 adjusts the audio attributes of each audio selection to effectively reposition audio selection 234 to a central position within 3D acoustic space and move audio selection 236 to an alternative position within 3D acoustic space 230. Advantageously, the simultaneous audio playback controller 112 adjusts the position of each content source when converted into output as an audio selection based on user posture information 120 indicating a preference for a particular audio selection, placing the user-selected audio selection in a prime position in front of the user, and supporting the user to apply selective listening to efficiently listen to a particular audio selection while all other audio selections continue to play in the background so that the user's attention is directed to listening to the other audio selections.
[0045] FIG. 3 shows a block diagram of an example of components of a simultaneous audio playback controller.
[0046] In this example, the simultaneous audio playback controller 112 includes a content monitor 350. The content monitor 350 monitors content accessed by the user from one or more online sources and stores a record of the content access in a content history log 352. When monitoring the content accessed by the user, the content monitor 350 may also analyze which content accesses indicate a preference by the user, such as a user selection to listen to a particular audio selection from among a plurality of audio selections within the simultaneous stream 116, via one or more types of input by the user, and may update or filter the content history log 352 using the indication of the user's preference via the user's instruction of the content selection.
[0047] In this example, the simultaneous audio playback controller 112 includes a topic selector 360. The topic selector 360 analyzes the content history log 352 to determine the types of topics of interest to the user and the frequency of each type. The topic selector 360 generates a topic preference 362 that includes the topics of interest and the frequency and provides an indication of the priority of the topics of interest. In one example, the topic selector 360 receives a large amount of content from the content history log 352, trains a cognitive model that learns the user's topic preference 362, applies a cognitive analysis function that is periodically updated, and then analyzes new content to quickly determine whether the new content selection matches the topic preference 362 for the user.
[0048] In one embodiment, when a user selects to access the listening service 110 through one or more interfaces, the topic interface controller 370 may initially request, access, and filter available text content from the content aggregator 114 for output by the topic interface controller 370. For example, the topic interface controller 370 may request current content from the content aggregator 114 based on the topic preferences 362 to determine interface content selections 374 for output in the simultaneous streams 116. The topic interface controller 370 may further filter the content received from the content aggregator 114 via the cognitive model of the topic selector 360 to select only the content source selections received from the content aggregator 114 as the interface content selections 374 and may also order the content sources according to their probability of interest to the user based on frequency and other priority information available in the topic preferences 362.
[0049] In another embodiment, output interface configuration 372 instructs topic interface controller 370 to divide the home page of content aggregator 114 into content sources based on one or more graphical characteristics that identify text passage boundaries, such as by dividing a news website into multiple content sources based on the layout of the news website using headline graphical elements or header codes that indicate the start of text passage boundaries. Topic interface controller 370 then analyzes each content source by applying topic preferences 362 to sort and order each content source according to the particular user's interests in the topics present within each content source.
[0050] In this example, the simultaneous audio playback controller 112 includes a text-to-speech converter 310 for converting text to speech. In this example, the topic interface controller 370 selects a plurality of content sources within the interface content selection 374, and the text-to-speech converter 310 converts each text content source into speech specified by one or more selectable features for adjusting the audio attributes of the speech conversion, differentiating different content sources for simultaneous streaming, and simulating an environment that supports a user to selectively listen to one of the speakers while listening to multiple different speakers in parallel and other speakers continue to play in the background. In one embodiment, the plurality of selectable features for adjusting the audio attributes of the speech conversion include the voice position and one or more features including, but not limited to, voice type, voice volume, and voice speed.
[0051] In this example, the voice type selector 312 selects a voice type for application by the text-to-speech converter 310 and specifies the characteristics of the speech conversion of the content source from the interface content selection 374. The voice type may include one or more dubbed voice types in one or more languages.
[0052] In this example, the voice volume selector 314 selects a voice volume for application by the text-to-speech converter 310 and specifies the characteristics of the speech conversion of the content source from the interface content selection 374. The voice volume may include one or more types of volume characteristics including, but not limited to, pitch, duration, timbre, and loudness.
[0053] In this example, the speech speed selector 316 selects the speed of speech for application by the text-to-speech converter 310 and specifies the characteristics of the speech conversion of the content source from the interface content selection 374. The speech speed may include a deceleration or an acceleration of the speed of words per minute in the speech converted from the text by one or more increments.
[0054] In this example, the sound source position selector 318 selects the spatial position for application by the text-to-speech converter 310 and specifies the characteristics of the speech conversion of the content source from the interface content selection 374. The spatial position may represent the cognitive arrangement of sounds within the 3D acoustic space 130, including the arrangement of sources in a plurality of planes at a specific recognition distance from the user.
[0055] In one example, the audio stream mixer 330 initially applies the attribute preference settings 320 to mix the voice attributes of the content source from the interface content selection 374 and outputs an audio stream 332 that simultaneously plays the speech from each of the converted content selections, such as the simultaneous stream 116. In this example, the attribute preference settings 320 specify one or more preferences that designate the voice type selector 312, the volume selector 314, the speech speed selector 316, and the sound source position selector 318.
[0056] In this example, the attribute preference settings 320 may specify preferences based on the topic subject, topic theme, topic priority, and other characteristics that facilitate the user's discrimination between speech converted from different content sources when streamed simultaneously. For example, the user may use different voice types by topic, different volumes by the weighting of interest in the topic and the position or font size of the content on the website of the content aggregator, different speech speeds by the theme of the topic, and different sound source positions by the weighting of interest in the most interesting content placed in the main position to specify the attribute preference settings 320.
[0057] In this example, the user posture interpreter 340 detects user posture information 120 indicating the user's movement, and determines the selection made by the user based on the user's movement relative to the position of the sound source position setting of the audio selection output in the audio stream 332. The audio stream mixer 330 adjusts one or more attributes applied by the voice type selector 312, the volume selector 314, the voice speed selector 316, and the sound source position selector 318, and also reflects the selection of the audio selection by the user based on the user's movement relative to the position of the sound source position setting of the audio selection output in the audio stream 332. Apply the posture-based adjustment setting 342. In one example, the posture-based adjustment setting 342 causes the audio stream mixer 330 to shift the position applied by the sound source position selector 318 to each of the content conversions to the voice in response to the user moving to the right or left of the center towards the audio selection, and to reposition the audio selection signaled by the user to the center within the 3D acoustic space 230. Additionally, the posture-based adjustment setting 342 includes settings for adjusting the voice type, volume, and voice speed of the audio selection in the audio mixed into the audio stream 332 to make the selected audio selection more prominent in a way that supports the user's ability to selectively further focus on the audio selection placed at the center position of the 3D acoustic space 230. Further, the posture-based adjustment setting 342 increases the volume of the audio selection placed at the center over time while all audio selections continue to play in parallel within the audio stream 332, and decreases the volume of the other placed audio selections over time, etc. It may also include settings for adjusting the voice type, volume, and voice speed of the audio mixed into the audio stream 332 to increase the focus on the audio selection placed at the center over time.
[0058] In one example, the attribute setting controller 322 manages a user interface for the user to select one or more preferences in each of the attribute preference setting 320, the pose-based adjustment setting 342, and the output interface setting 372. In one example, in addition to the user setting preferences, the attribute setting controller 322 may prompt the user to indicate whether the application of the attribute preference setting 320, the pose-based adjustment setting 342, and the output interface setting 372 results in the output expected by the user, or alternatively, may monitor user input indicating whether the application of the attribute preference setting 320, the pose-based adjustment setting 342, and the output interface setting 372 results in the output expected by the user. In one embodiment, the attribute setting controller 322 may evaluate the user response and apply machine learning to train a model that manages one or more of the attribute preference setting 320, the pose-based adjustment setting 342, and the output interface setting 372. Advantageously, by supporting an interface for specifying setting preferences for the user and by continuing to train preferences based on the user response, the attribute setting controller 322 supports the specific user preferences for selective listening in an efficient manner, in a manner in which the space is filled by multiple voices, and adjusts the position of the output audio selection and the characteristics of the audio attributes in the simultaneous stream to prepare for the parallel output of multiple content sources within the 3D acoustic space 130.
[0059] FIG. 4 shows a block diagram of an example of a computing device that implements a simultaneous audio playback controller guided by user pose information detected by a visual monitor so that attention can be directed to an audio selection placed at a primary position among a plurality of content sources converted from text to speech and played in parallel.
[0060] In one example, listening service 110 runs on computing device 400 and generates an audio stream that simultaneously streams multiple content selections converted from text to audio selections, such as audio selection 432, audio selection 434, audio selection 436, and audio selection 438. In this example, I / O interface 422 outputs simultaneous stream 416 generated by listening service 110 to stereo speaker system 410. Stereo speaker system 410 includes multiple speaker units controlled to output audio tracks in simultaneous stream 416 in 3D acoustic space 430 that are audibly detectable by a user present in the space.
[0061] In one embodiment, as a user listens to the simultaneous streams 416 through a stereo speaker system 410, a vision monitor 412 detects movement of the user's head 444. In this example, the vision monitor 412 is positioned near the user's head 444 to detect movement of the user's head 444 or other pose information, such as the user rotating their head 444 to the left as depicted by reference numeral 442, or other pose information indicating a selection. For example, the vision monitor 412 represents one or more types of cameras or motion detectors for detecting specific types of movement within a recording area. In this example, the vision monitor 412 is positioned near the components of the stereo speaker system 410 and outputs the user pose information 414 to the I / O interface 422 of the computing device 400 for application by the listening service 110. Alternatively, the visual monitor 412 may be located independently of the stereo speaker system 410 component and may be a component of the computing device 400 for outputting user posture information 414 for application by the listening service 110.
[0062] FIG. 5 shows a block diagram of an example of an adjustment to the position of an audio selection in a simultaneous stream of audio selections converted from a plurality of text content sources in response to detecting user pose information indicating a selection of a particular audio selection output within a 3D acoustic space.
[0063] In one example, the 3D acoustic space 530 represents an acoustic space in which the listening service 110 outputs a simultaneous stream 116 that includes "audio selection #1", "audio selection #2", "audio selection #3", and "audio selection #4". In this example, as indicated by reference numeral 532, "audio selection #3" is initially placed at "position #1" within the 3D acoustic space 530. As indicated by reference numeral 534, "audio selection #2" is initially placed at "position #2" within the 3D acoustic space 530. As indicated by reference numeral 536, "audio selection #1" is initially placed at "position #3" within the 3D acoustic space 530. As indicated by reference numeral 538, "audio selection #4" is initially placed at "position #4" within the 3D acoustic space 530. In addition to each of the audio selections being dispersed at different positions within the 3D acoustic space 530, each of the audio selections may be output with an audio attribute that aurally differentiates one or more characteristics of the respective audio selection.
[0064] In this example, as shown at reference numeral 510, the listening service 110 detects user posture information 120 that reflects the movement of the user's head rotating to the right. Advantageously, in response to the movement of the user's head rotated to the right towards "Audio Selection #4" at reference numeral 538, the listening service 110 selects to relocate the position of "Audio Selection #4" to the primary position of "Position #1" as shown at reference numeral 540, and supports the user to apply selective listening and pay more attention to "Audio Selection #4" by positioning "Audio Selection #4" at the primary position in front of the user. Additionally, the listening service 110 selects to relocate "Audio Selection #1" to the available position of "Position #4" as shown at reference numeral 542. In addition to relocating "Audio Selection #1" and "Audio Selection #4", advantageously, the listening service 110 adjusts the audio attributes of one or more characteristics of each audio selection to aurally distinguish "Audio Selection #4" at the primary position from other audio selections, increase the characteristics that increase the aural focus on "Audio Selection #4", and decrease the characteristics that increase the aural focus on other audio selections, thereby supporting the user to apply selective listening and pay more attention to "Audio Selection #4".
[0065] FIG. 6 shows an example of a block diagram of an interface content selection identified from an interface of a content aggregator by a listening service for managing simultaneous streams of text source content that is converted to audio and played in parallel.
[0066] In one example, the content aggregator interface 602 includes multiple types of text content. The topic interface controller 370 applies an output interface setting 372 that includes at least one graphical characteristic for identifying boundaries between text messages, evaluates the graphical characteristics of the text content within the content aggregator interface 602, and identifies one or more content sources in the interface content selection 630 for text-to-speech conversion and parallel playback in a simultaneous stream.
[0067] In this example, the content aggregator interface 602 includes four separate messages of text content, exemplified as "Text Content #1" 612, "Text Content #2" 616, "Text Content #3A" 622, and "Text Content #3B" 626. In this example, the topic interface controller 370 evaluates the graphical characteristics identified by the output interface setting 372 within the content aggregator interface 602. The topic interface controller 370 has a code style that reflects the boundary where the heading 610 starts "Text Content #1" 612, and the line 614 is a graphical element that reflects the boundary where "Text Content #2" 616 starts. The heading 620 has a code style that reflects the boundary where "Text Content #3A" 622 starts, and the ad 624 is a graphical element that graphically separates "Text Content #3A" 622 and "Text Content #3B" 626, but it is identified that it is not the boundary between different text messages.
[0068] Advantageously, by evaluating the graphical characteristics within content aggregator 602 and identifying the boundaries between text messages, topic interface controller 370 enables splitting text messages from different content aggregators into separate content sources. In this example, topic interface controller 370 splits content aggregator interface 602 into three content sources within interface content selection 630 based on the boundaries identified between text messages. In this example, topic interface controller 370 splits a first content selection 632 from heading 610 and "text content #1" 612, a second content selection 634 from heading 620 and "text content #2" 616, and a third content selection 636 from "text content #3A" and "text content #3B".
[0069] FIG. 7 shows a block diagram of an example of a computer system in which an embodiment of the present invention may be implemented. The present invention may be executed in various systems and system combinations composed of functional components such as the functional components described with reference to computer system 700, and may be communicatively connected to a network such as network 702.
[0070] Computer system 700 includes a bus 722 or other communication device for communicating information within computer system 700, and at least one hardware processing device such as processor 712 coupled to bus 722 for processing information. Bus 722 preferably includes low latency and high latency paths connected by bridges and adapters and controlled within computer system 700 by multiple bus controllers. When implemented as a server or node, computer system 700 may include multiple processors designed to improve network service power.
[0071] During normal operation, the processor 712 may be at least one general-purpose processor that processes data under the control of software 750, which may include at least one of application software, an operating system, middleware, and other code and computer-executable programs accessible from a dynamic storage device such as random access memory (RAM) 714, a static storage device such as read-only memory (ROM) 716, a data storage device such as mass storage device 718, or other data storage media. Software 750 may include, but is not limited to, code, applications, protocols, interfaces, and processes for controlling one or more systems within a network including, but not limited to, adapters, switches, servers, cluster systems, and grid environments.
[0072] The computer system 700 may communicate with a remote computer such as server 740 or a remote client. In one example, server 740 may be connected to computer system 700 via any type of network such as network 702, via a communication interface such as network interface 732, or via a network link that may be connected to network 702, for example.
[0073] In this example, a plurality of systems within a network environment may be communicatively connected via network 702, and the network is a medium used to provide communication links between various devices and computer systems that are communicatively connected. Network 702 may include, for example, permanent connections such as wired or fiber optic cables, and temporary connections made via telephone connections and wireless transmission connections, and may include, for example, routers, switches, gateways, and other hardware for enabling communication channels between systems connected via network 702. Network 702 may represent one or more of a packet-switched network, a telephone-based network, a broadcast television network, a local area and wired area network, a public network, and a restricted network.
[0074] Network 702, and the systems communicatively connected to computer 700 via network 702, may implement one or more layers of one or more types of network protocol stacks that may include one or more of the physical layer, link layer, network layer, transport layer, presentation layer, and application layer. For example, network 702 may implement one or more of a Transmission Control Protocol / Internet Protocol (TCP / IP) protocol stack or an Open Systems Interconnect (OSI) protocol stack. Additionally, for example, network 702 may represent a worldwide collection of networks and gateways that use the TCP / IP suite of protocols that communicate with each other. Network 702 may implement a Secure HTTP protocol layer or other security protocols for securing communication between systems.
[0075] In this example, network interface 732 connects computer system 700 to network 702 via a link and includes adapter 734 for communicatively connecting computer system 700 to a server 740 or other computer systems via network 702. Although not shown, network interface 732 may include additional software, such as device drivers, additional hardware, and other controllers to enable communications. When implemented as a server, computer system 700 may include multiple communication interfaces accessible, for example, through multiple Peripheral Component Interconnect (PCI) bus bridges connected to input / output controllers. In this manner, computer system 700 may enable connection to multiple clients through multiple separate ports, each of which may also support multiple connections to multiple clients.
[0076] In one embodiment, the operations performed by processor 712 may control the flowcharts of Figures 8-11 and other operations described herein. The operations performed by processor 712 may be called for by software 750 or other code, or the steps of an embodiment of the present invention may be performed by specific hardware components that contain hardwired logic to perform the steps, or any combination of programmed computer components and customized hardware components. In one embodiment, one or more components of computer system 700, or other components that may be integrated into one or more components of computer system 700, may include hardwired logic to perform the operations of the flowcharts of Figures 8-11.
[0077] In addition, computer system 700 may include a plurality of peripheral components that facilitate input and output. These peripheral components are connected to expansion slots such as input / output (I / O) interface 726 coupled to one of a plurality of controllers, adapters, and buses 722 at multiple levels. For example, input device 724 may include a microphone, video capture device, image reading system, keyboard, mouse, or other input peripheral device that is communicatively enabled on bus 722 via, for example, I / O interface 726 that controls the input. Additionally, output device 720, communicatively enabled on bus 722 via, for example, I / O interface 726 for controlling output, may include, for example, one or more graphical display devices, audio speakers, and a tactilely detectable output interface, but may include other output interfaces. In alternative embodiments of the present invention, additional or alternative input and output peripheral components may be added.
[0078] With respect to FIG. 7, the present invention may be a system, method, or computer program product or a combination thereof. The computer program product may include a computer-readable storage medium (media) having thereon computer-readable program instructions for causing a processor to execute aspects of embodiments of the present invention.
[0079] A computer-readable storage medium may be a tangible device that holds and stores instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A more specific non-exhaustive list of examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy (registered trademark) disk, punch card, or a mechanically encoded device such as a raised structure in a groove having recorded instructions, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not itself a propagated signal such as a radio wave, a freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through a fiber optic cable), or an electrical signal transmitted through a wire that is interpreted as a transient signal per se.
[0080] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computers / processing devices or to an external computer or external storage device via a network, such as, for example, the Internet, a local area network, a wide area network, or a wireless network, or a combination of these. The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination of these. A network adapter card or network interface in each computer / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium within each computing / processing device.
[0081] Computer-readable program instructions for carrying out operations of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine language instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented languages such as Smalltalk®, C++, or the like, and conventional procedural languages such as the C programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server, as a stand-alone software package. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, the electrical circuitry may execute the computer-readable program instructions by utilizing state information of the computer-readable program instructions to individualize the electrical circuitry, which may include, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), to perform aspects of the invention.
[0082] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0083] These computer-readable program instructions can be provided to a general-purpose computer, special-purpose computer processor, or other programmable data processing apparatus to produce machine-generated means for implementing the functions / acts identified in the block or blocks of the flowchart diagrams and / or block diagrams. These computer-readable program instructions can also be stored on a computer-readable storage medium capable of instructing a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, such that a computer-readable storage medium having instructions stored thereon includes an article of manufacture containing instructions that implement aspects of the functions / acts identified in the block or blocks of the flowchart diagrams and / or block diagrams.
[0084] The computer readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device and cause the computer, other programmable data processing apparatus, or other device to perform a series of operational steps to generate a computer-implemented process such that the instructions executing on the computer, other programmable data processing apparatus, or other device implement aspects of the functions / actions identified in a block or blocks of the flowchart diagrams and / or block diagrams.
[0085] The flowchart diagrams and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart diagrams or block diagrams may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing a particular logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order depicted in the figures. For example, two blocks shown in succession may in fact be executed concurrently or substantially concurrently, or the blocks may be executed in the reverse order, depending on the functionality involved. It should be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a special-purpose hardware-based system that performs particular functions or actions or that implements a combination of special-purpose hardware and computer instructions.
[0086] Those of ordinary skill in the art will appreciate that the hardware depicted in Figure 7 may vary. Additionally, those skilled in the art will appreciate that the depicted example is not meant to imply architectural limitations with respect to the present invention.
[0087] FIG. 8 illustrates a high-level logical flowchart of a process and computer program for tracking user topic preferences for application by a listening service to determine and prioritize content sources for text-to-speech conversion and parallel playback in simultaneous streams.
[0088] In one example, the process and computer program start at block 800 and then proceed to block 802. Block 802 indicates monitoring for user instructions for content selection. Next, block 804 indicates a determination of whether a user instruction for content selection has been detected. At block 804, if a user instruction for content selection is detected, the process proceeds to block 806. Block 806 indicates recording the content selection user instruction and the content in a content history log. Next, block 808 indicates analyzing the content history log for the type and frequency of the user's content selections. Thereafter, block 810 indicates generating a topic preference for the user based on the topics in the types of content frequently selected by the user. Next, block 812 indicates training a topic preference model using the topic preference and ends the process.
[0089] Figure 9 shows a high-level logical flowchart of a process and computer program for managing a listening service that implements a simultaneous audio playback controller guided by user posture information.
[0090] In one example, the process and computer program start at block 900 and then proceed to block 902. Block 902 represents a determination of whether the user selects to listen to text-to-speech-based content. At block 902, if the user selects to listen to text-to-speech-based content, the process proceeds to block 904. Block 904 represents identifying a collection of potential selections of text messages available for listening from a content aggregator. Next, block 906 represents splitting the text messages from the content aggregator into separate content sources, identified by boundaries based on graphical characteristics identified in the output interface settings applied to the output interface of the content aggregator, such that each content source reflects a separate text message from the collection. Thereafter, block 908 represents identifying a selection of content sources whose content meets the user's topic preferences. Next, block 910 represents sorting the selection of content sources according to the priority levels identified for each of the topic preferences. Thereafter, block 912 represents applying the attribute preference settings to select the auditory positions within the selected 3D acoustic space according to the sorted priorities and specifying the audio attributes of the speech conversion of the selection of content sources. Next, block 914 represents acoustically outputting each of the text messages from the selection of content sources as audio selections associated with different positions within the 3D acoustic space according to the sorted priorities and simultaneously streaming the text-to-speech conversion of the selection of content sources to a single stereo output stream according to the applied audio attributes, and the process proceeds to block 916.
[0091] Block 916 indicates monitoring the input of user posture information. Next, block 918 indicates determining whether the input of user posture information has been received. In block 918, if the input of user posture information is received, the process proceeds to block 920. Block 920 indicates triggering the evaluation of the input of user posture information, and the process is passed to block 922. When returning to block 918, if the input of user posture information is not received, the process proceeds to block 922.
[0092] Block 922 indicates determining whether there is a user input to stop listening to text-to-speech-based content. In block 922, if a user input to stop listening to text-to-speech-based content is received, the process ends. Otherwise, if a user input to stop listening to text-to-speech-based content is not received in block 922, the process returns to block 916.
[0093] FIG. 10 shows an example of a high-level logic flowchart of a process and program for managing a simultaneous audio playback controller guided by user posture information to arrange an audio selection converted from text to audio playback of a specific content source from among a plurality of content sources provided by a content aggregator at a primary position from among a plurality of content sources converted from text to audio and played back in parallel.
[0094] In one example, the process and computer program begin at block 1000 and then proceed to block 1002. Block 1002 illustrates determining whether evaluation of user posture information input is triggered. If evaluation of user posture information input is triggered at block 1002, the process proceeds to block 1004. Block 1004 illustrates analyzing the direction of movement of the input user posture information relative to the positioning of the audio selection within the 3D acoustic space to identify a specific audio selection associated with the direction of movement within the 3D acoustic space. Next, block 1006 illustrates adjusting the reordering of the content source selection to direct placement of the specific audio selection at a key location within the 3D acoustic space. Thereafter, block 1008 illustrates applying posture-based adjustment settings to the adjusted content source selection to adjust audio attributes of the audio conversion of the content source selection to audibly adjust the position of the audio selection within the 3D acoustic space and adjust audio characteristics of the audio selection to enhance the user's ability to selectively focus on the specific audio selection. Next, block 1010 depicts acoustically outputting each of the text passages from the selection of content sources as an audio selection, and adjusting the simultaneous streams of text-to-speech conversions of the selection of content sources according to the adjusted audio attributes into a single stereo output stream while highlighting the particular audio selection at a key location within the 3D acoustic space, and the process ends.
[0095] FIG. 11 illustrates an example high level logical flowchart of a process and program for setting user preferences for attribute preference settings, user postural movements, and posture-based adjustment settings for application by a simultaneous audio playback controller guided by user posture information.
[0096] In one example, the process and computer program start at block 1100 and then proceed to block 1102. Block 1102 indicates a determination of whether the user selects to specify preferences based on text-to-speech. At block 1102, if the user selects to specify preferences based on text-to-speech, the process proceeds to block 1104. Block 1104 indicates prompting the user to select an ordering preference for audio priority positions within a 3D acoustic space, including selecting a primary position among a plurality of positions. Then, block 1106 indicates identifying a selection of a topic from among the user topic preferences. Next, block 1108 indicates prompting the user to select one or more voice types, voice volumes, voice positions, and voice speed attributes based on the topic, theme, and presentation characteristics of the content. Then, block 1110 indicates storing the user's selection as the user's attribute preference setting, and the process proceeds to block 1112.
[0097] Block 1112 indicates prompting the user to make movements of one or more postures to sample the user's selection of movement for selecting different positions within the 3D audio space. Next, block 1114 indicates training a cognitive model for interpreting the movement of the user's posture based on the sampling of the movement of the user's posture. Next, block 1116 indicates prompting the user to select one or more voice types, voice volumes, voice positions, and voice speed attributes in relation to the movement of the user's posture for different positions within the 3D audio space. Then, block 1118 indicates storing the user's selection as a posture-based adjustment setting for the user, and the process ends.
[0098] In one or more embodiments, a method, computer system, and computer program product are directed to identifying a plurality of content sources, each including text content, where the plurality of content sources are ordered by interest to a user. The method, computer system, and computer program product are directed to converting each of the plurality of content sources from text content into speech including distinct audio selections of a plurality of audio selections. The method, computer system, and computer program product are directed to specifying audio attributes that apply one or more audio attributes to each audio selection to select a respective position of each audio selection from among a plurality of positions in a multi-dimensional acoustic space and to aurally distinguish one or more characteristics of each audio selection from other audio selections of the plurality of audio selections, where the respective position of each audio selection reflects the rank of the respective audio selection ordered by interest. The method, computer system, and computer program product are directed to outputting a simultaneous stream of the plurality of audio selections to an audio output device for stereoscopic playback of the plurality of audio selections to the user at a plurality of positions in a multi-dimensional acoustic space that reflects the plurality of content sources ordered by interest. One advantage of specifying audio attributes that apply one or more audio attributes to each audio selection to select a respective position of each audio selection from among a plurality of positions in a multi-dimensional acoustic space and to aurally distinguish one or more characteristics of each audio selection from other audio selections of the plurality of audio selections is that the simultaneous stream of audio selections provides an audio output that supports a user in scanning through the content of the plurality of audio selections presented in parallel at different audible positions detectable to the user and identifying a topic of interest to the user from among the audio selections.One advantage of identifying a plurality of content sources ordered by user interest, specifying audio attributes, and selecting each position of the audio selection that reflects the rank of each audio selection ordered by interest is that the simultaneous stream of audio selections provides an audio output that places the content source that the user is most likely to be interested in at one or more primary positions, where this one or more primary positions improves the predictability for the user that the audible selection at the one or more primary positions will have the highest interest for the user when scanning the auditory positions.
[0099] Additionally, in one or more embodiments, a method, computer system, and computer program product are directed to identifying a plurality of content sources, each including text content, where the plurality of content sources are ordered by interest to the user. The method, computer system, and computer program product may be directed to monitoring content accessed by the user as an indication of the user's selection of content. The method, computer system, and computer program product may further be directed to recording the indication of selection and the content in a content history log in response to detecting an indication of the user's selection of content. The method, computer system, and computer program product may further be directed to analyzing the content history log for the frequency of at least one type of content. The method, computer system, and computer program product may further be directed to generating at least one topic preference for the user based on at least one type of content and its respective frequency. The method, computer system, and computer program product may be directed to training a cognitive model using at least one topic preference and its respective frequency to identify other content including at least one topic. The method, computer system, and computer program product may be directed to identifying a plurality of content sources via the cognitive model from a selection of a plurality of text passages including at least one topic in response to the user accessing a website including a plurality of text messages. The method, computer system, and computer program product may further be directed to ordering the plurality of content sources according to respective priority levels identified by the cognitive model based on the respective frequency of at least one topic present in each of the plurality of content sources.One advantage of training a cognitive model using topic preferences and frequencies based on the monitored user activity of accessing and responding to content, and then selectively identifying and ordering content sources by the trained cognitive model, is that the selection of content sources that are played in parallel in a simultaneous stream is ordered in an auditory position according to the likelihood of interest to a particular user, improving the efficiency for a particular user to scan through multiple audio selections streamed in parallel and find at least one audio selection of interest to the user.
[0100] In addition, in one or more embodiments, a method, computer system, and computer program product for identifying a plurality of content sources, each including text content, where the plurality of content sources are ordered by interest to a user, may be directed to identifying an interface that includes a plurality of text messages available from a content aggregator. The method, computer system, and computer program product may further be directed to dividing, by a computer system, a plurality of text messages from a content aggregator into separate content sources of a plurality of content sources based on an output interface setting that specifies at least one graphical characteristic for identifying a boundary between text messages in the interface. One advantage of identifying a plurality of text messages available from a content aggregator and dividing the plurality of text messages into content sources based on an output interface setting that specifies at least one graphical characteristic for identifying a boundary between text messages in the interface is that the content sources are identified independently of any particular content aggregator interface, and simultaneous streams are generated for content sources from multiple types of content aggregators and independent of whether the content aggregator outputs text messages in the interface for conversion to speech.
[0101] In addition, in one or more embodiments, for each audio selection, one or more audio attributes are applied to specify an audio attribute that selects a respective position of each audio selection from among a plurality of positions in a multi-channel audio space, and a method, computer system, and computer program product for aurally distinguishing one or more characteristics of each audio selection from other audio selections of the plurality of audio selections, wherein the respective position of each audio selection reflects the rank of each audio selection ordered by interest, the method, computer system, and computer program product may be directed to applying one or more audio attributes to each audio selection to specify an audio attribute for a particular position among a plurality of positions at a primary position in front of the user in a multi-channel audio space for a particular audio selection converted from a particular content source having the highest rank ordered by interest. One advantage of applying one or more audio attributes to specify an audio attribute for a particular audio selection having the highest rank ordered by interest at a primary position in front of the user in a multi-channel audio space is to improve the efficiency when the user determines an order to scan and focus on each of the audio selections that are played in parallel while the audio selection that is most likely to be of most interest to the user is predictably positioned at a position in front of the user along with a plurality of audio selections at different positions.
[0102] In addition, in one or more embodiments, the method, computer system, and computer program product may be configured to monitor an input indicative of a user posture change. The method, computer system, and computer program product may also be configured to analyze, by a computer system, a direction of movement of a user posture change relative to positioning of a plurality of audio selections in a multi-dimensional audio space including a 3D audio space in response to detecting an input indicative of a user posture change. The method, computer system, and computer program product may also be configured to identify, by a computer system, a particular audio selection from among the plurality of audio selections selected by a user based on a direction of movement of a user posture change relative to positioning of the plurality of audio selections in a 3D audio space. The method, computer system, and computer program product may also be configured to specify, by a computer system, audio attributes that apply one or more updated audio attributes to each audio selection to select a respective position of each audio selection from among a plurality of positions in a 3D audio space and to aurally distinguish one or more characteristics of each audio selection from other audio selections of the plurality of audio selections, wherein the particular audio selection is moved to a primary position in front of the user in the 3D audio space and the volume of the particular audio selection is increased relative to other audio selections of the plurality of audio selections. One advantage of monitoring an input indicative of a user posture change and detecting a direction of movement in a user posture change relative to positioning of a plurality of audio selections in a multi-dimensional audio space is to provide an input for a user to focus on a particular audio selection with a minimal movement that enables the user to select and focus on a particular audio selection from among a plurality of audio selections in a simultaneous stream without hands-free and without requiring precise movement.
[0103] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that the terms "comprise" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, or components, or any combination thereof, and do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups or combinations thereof.
[0104] In the following claims, equivalent structures, materials, acts, and all means or step-plus-function element equivalents are intended to include any structure, material, or act performed in combination with other claimed elements as explicitly claimed. The description of one or more embodiments of the present invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or to limit the invention to the precise form disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the invention. The embodiments were chosen and described to best explain the principles and practical application of the invention and to enable those skilled in the art to understand the invention in various embodiments, with various modifications suited to the particular use contemplated.
[0105] Although the present invention has been particularly shown and described with reference to one or more embodiments, it will be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the invention.
Claims
1. A method comprising: identifying, by a computer system, a plurality of content sources each including text content, wherein the plurality of content sources are ordered according to a user's interest; converting, by the computer system, each of the plurality of content sources from the text content into an audio including distinct audio selections of a plurality of audio selections; specifying, by the computer system, audio attributes for each of the plurality of audio selections to place each of the plurality of audio selections at a respective position among a plurality of positions in a multi-dimensional acoustic space and to aurally distinguish one or more characteristics of each of the plurality of audio selections from other audio selections of the plurality of audio selections, wherein the respective positions of the plurality of audio selections reflect ranks of the respective audio selections ordered according to the interest; outputting, by the computer system, a simultaneous stream of the plurality of audio selections to an audio output device for stereo playback of the plurality of audio selections at the plurality of positions in the multi-dimensional acoustic space that reflect the plurality of content sources ordered according to the interest; and a method comprising the steps of:
2. The step of identifying, by the computer system, the plurality of content sources each including text content, wherein the plurality of content sources are ordered according to the user's interest, comprises: monitoring, by the computer system, content accessed by a user as an indication of a selection of the content by the user; recording, by the computer system, the indication of the selection and the content in a content history log in response to detecting the indication of the selection of the content by the user; analyzing, by the computer system, the content history log for frequencies of at least one type of content; generating, by the computer system, at least one topic preference for the user based on the at least one type of content and respective frequencies thereof; training, by the computer system, a cognitive model using the at least one topic preference and the respective frequencies to identify other content including the at least one topic; identifying, by the computer system, via the cognitive model, the plurality of content sources from a selection of the plurality of text passages including the at least one topic in response to the user accessing a website including a plurality of text passages; ordering, by the computer system, the plurality of content sources according to respective priority levels identified by the cognitive model based on the respective frequencies of the at least one topic present in the respective content sources of the plurality of content sources; The method according to claim 1, comprising:
3. The step of identifying, by the computer system, the plurality of content sources each including text content, wherein the plurality of content sources are ordered by interest for a particular user, comprises: identifying, by the computer system, an interface including a plurality of text passages available from a content aggregator; dividing, by the computer system, the plurality of text passages from the content aggregator into separate content sources of the plurality of content sources based on an output interface setting that specifies at least one graphical characteristic to identify boundaries between text passages in the interface; The method according to claim 1, comprising:
4. The plurality of selectable features for adjusting the audio attribute includes one or more of one or more voice type attributes, one or more voice volume attributes, and one or more voice speed attributes. The method according to claim 1.
5. The step of specifying the audio attribute further comprises: The step of specifying the audio attributes by the computer system for each audio selection, by topic, by the graphical representation characteristics of each content source, and based on the attribute preference settings specified by the theme The method according to claim 4, comprising: **Claim 6** The step of specifying the audio attributes further comprises: The step of specifying the audio attributes for a specific position among the plurality of positions at a main position in front of the user in the multi-channel audio space for a specific audio selection converted from a specific content source having the highest rank ordered by the interest by the computer system The method according to claim 1, comprising: **Claim 7** The method comprises: The step of monitoring, by the computer system, an input indicating a user posture change; In response to detecting the input indicating the user posture change, the step of analyzing, by the computer system, the direction of movement of the user posture change relative to the positioning of the plurality of audio selections in the multi-channel audio space including the 3D audio space; The step of identifying, by the computer system, a specific audio selection from among the plurality of audio selections selected by the user based on the direction of movement of the user posture change relative to the positioning of the plurality of audio selections in the 3D audio space; The step of specifying updated audio attributes by the computer system to relocate each of the plurality of audio selections to its respective position from among the plurality of positions in the 3D audio space and to aurally distinguish the one or more characteristics of each of the audio selections from other audio selections of the plurality of audio selections, wherein the specific audio selection is moved to a main position in front of the user in the 3D audio space and the volume of the specific audio selection is increased relative to other audio selections of the plurality of audio selections The method according to claim 1, further comprising: **Claim 8** [[ID= The method of claim 7, comprising the step of receiving, by the computer system, an output from a gyroscope that measures an angular velocity to measure a rotational movement of a head of a user who listens to the audio output of the simultaneous stream. **Claim 9** A computer system comprising one or more processors, one or more computer-readable memories, one or more computer-readable storage devices, and program instructions stored in at least one of the one or more computer-readable storage devices and executable by at least one of the one or more processors via at least one of the one or more memories, wherein the stored program instructions are program instructions that identify a plurality of content sources each including text content, and the plurality of content sources are ordered by interest to the user. are program instructions that convert each of the plurality of content sources from the text content into audio including distinct audio selections of a plurality of audio selections. are program instructions that place each audio selection of the plurality of audio selections at a respective position among a plurality of positions in a multi-channel acoustic space and specify audio attributes to aurally distinguish one or more characteristics of each audio selection from other audio selections of the plurality of audio selections, wherein the respective positions of the respective audio selections reflect the rank of the respective audio selections ordered by the interest. are program instructions that output a simultaneous stream of the plurality of audio selections to an audio output device for stereo playback of the plurality of audio selections to the user within the plurality of positions in the multi-channel acoustic space that reflect the plurality of content sources ordered by the interest. A computer system comprising. **Claim 10** The stored program instructions are program instructions that identify the plurality of content sources each including text content, and the plurality of content sources are ordered by interest to the user. Program instructions for monitoring content accessed by a user as an indication of the user's selection of the content; Program instructions for recording the selection instruction and the content in a content history log in response to detecting an indication of the user's selection of the content; Program instructions for analyzing the content history log for the frequency of at least one type of content; Program instructions for generating at least one topic preference for the user based on the at least one type of content and its respective frequency; Program instructions for training a cognitive model using the at least one topic preference and its respective frequency to identify other content including the at least one topic; Program instructions for identifying a plurality of content sources via the cognitive model from a selection of the plurality of text passages including the at least one topic in response to the user accessing a website including the plurality of text passages; Program instructions for ordering the plurality of content sources according to respective priority levels identified by the cognitive model based on the respective frequency of the at least one topic present in each of the plurality of content sources; The computer system according to claim 9, comprising:
11. Program instructions for identifying the plurality of content sources each including text content, the plurality of content sources being ordered by interest for a particular user, the stored program instructions further comprising: Program instructions for identifying an interface including a plurality of text passages available from a content aggregator; Program instructions for splitting the plurality of text passages from the content aggregator into separate content sources of the plurality of content sources based on an output interface setting that specifies at least one graphical characteristic for identifying a boundary between text passages in the interface; The computer system according to claim 9, comprising:
12. The computer system according to claim 9, wherein the plurality of selectable features for adjusting the audio attribute includes one or more of one or more voice type attributes, one or more voice volume attributes, and one or more voice speed attributes.
13. The stored program instructions for specifying the audio attribute further include program instructions for specifying the audio attribute for each audio selection based on an attribute preference setting specified by a topic, by graphical representation characteristics of each content source, and by a theme. The computer system according to claim 12.
14. The stored program instructions for specifying the audio attribute further include program instructions for specifying the audio attribute for a specific position among the plurality of positions at a main position in front of the user in the multi-channel audio space for a specific audio selection converted from a specific content source having the highest rank ordered by the interest. The computer system according to claim 9.
15. The stored program instructions further include program instructions for monitoring an input indicating a user posture change, program instructions for analyzing a direction of movement of the user posture change relative to the positioning of the plurality of audio selections in the multi-channel audio space including the 3D audio space in response to detecting the input indicating the user posture change, and program instructions for identifying a specific audio selection from among the plurality of audio selections selected by the user based on the direction of movement of the user posture change relative to the positioning of the plurality of audio selections in the 3D audio space. Program instructions that relocate each of the plurality of audio selections to respective positions from among the plurality of positions in the 3D acoustic space and specify updated audio attributes that are updated to aurally distinguish one or more characteristics of each of the audio selections from other audio selections of the plurality of audio selections, wherein the particular audio selection is moved to a primary position in front of the user in the 3D acoustic space and the volume of the particular audio selection is increased relative to other audio selections of the plurality of audio selections, and the program instructions The computer system according to claim 9, comprising: **Claim 16** The stored program instructions, which monitor an input indicating a change in the user's posture, further Program instructions that receive an output from a gyroscope that measures angular velocity, which measures a rotational movement of the head of a user listening to the audio output of the simultaneous streams The computer system according to claim 15, comprising: **Claim 17** A computer program that causes a computer system to execute each step in the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Document read-aloud device
JP2000099307A
Method and device for reproducing synthesized voice
JP2003099078A
Voice output controlling device
JP2006115364A
Voice portlet providing system and voice portlet providing method
JP2012195661A
Voice output device and voice output method
JP2015232849A