Voice synthesis device using multimodal mixing
Patent Information
- Application Number
- JP2022526206
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-11-07
- Filing Date
- 2020-10-21
- Publication Date
- 2025-06-02
- Estimated Expiration
- 2040-10-21
AI Technical Summary
Existing speech synthesizers lack effective multimodal blending capabilities for enhanced user interaction and pronunciation clarity.
A speech synthesizer with improved multimodal mixing that detects drag velocity on a touch-sensitive display screen to select appropriate audio files for pronouncing words, using multiple drag speed ranges to determine whether to play individual phonemes or a complete word, thereby adapting to user input speed.
Enhances user experience by providing dynamic adaptive speech synthesis with user-desired pronunciation speeds, reducing computational resources and effort, and improving clarity at both slow and fast speeds.
Smart Images

Figure 00000024_0000 
Figure 00000025_0000 
Figure 00000026_0000
Abstract
Description
Technical Field
[0001] The subject matter disclosed herein generally relates to the technical field of special-purpose machines that facilitate speech synthesis. The subject matter relates to software-configured computerized variations of such special-purpose machines and improvements thereto, as well as techniques in which such special-purpose machines are improved compared to other special-purpose machines that facilitate speech synthesis. Specifically, the present disclosure addresses systems and methods for providing a speech synthesizer.
Background Art
[0002] A machine may be configured to interact with one or more users of the machine (e.g., a computer or other device) by presenting exercises that teach one or more reading skills to one or more users or otherwise guide one or more users through the practice of one or more reading skills. For example, the machine may present alphabetic characters (e.g., the character "A" or the character "B") within a graphical user interface (GUI) and synthesize speech by playing the voice or video recording of an actor pronouncing the presented alphabetic character, and then may also prompt the user (e.g., a child learning to read) to pronounce the presented alphabetic character.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, there is room for improvement in speech synthesizers using multimodal blending. [Means for solving the problem]
[0005] An improved speech synthesis device using multimodal mixing is provided.
[0006] Some embodiments are illustrated, and not limited to, in the accompanying drawings. [Brief explanation of the drawing]
[0007] [Figure 1] Surface view of a machine (e.g., apparatus) having a touch-sensitive display screen on which a GUI suitable for speech synthesis is presented, according to several exemplary embodiments. [Figure 2] Surface view of a machine (e.g., device) having a touch-sensitive display screen on which a GUI suitable for speech synthesis is presented, according to several exemplary embodiments. [Figure 3] Surface view of a machine (e.g., device) having a touch-sensitive display screen on which a GUI suitable for speech synthesis is presented, according to several exemplary embodiments. [Figure 4] Surface view of a machine (e.g., device) having a touch-sensitive display screen on which a GUI suitable for speech synthesis is presented, according to several exemplary embodiments. [Figure 5] Surface view of a machine (e.g., device) having a touch-sensitive display screen on which a GUI suitable for speech synthesis is presented, according to several exemplary embodiments. [Figure 6] A block diagram showing the components of a machine according to several exemplary embodiments. [Figure 7] A flowchart illustrating the operation of a machine when performing a speech synthesis method, according to several exemplary embodiments. [Figure 8] A flowchart illustrating the operation of a machine when performing a speech synthesis method, according to several exemplary embodiments. [Figure 9]A flowchart illustrating the operation of a machine when performing a speech synthesis method, according to several exemplary embodiments. [Figure 10] A block diagram showing components of a machine capable of reading instructions from a machine-readable medium and performing one or more of the methodologies discussed herein, according to several exemplary embodiments. [Modes for carrying out the invention]
[0008] The exemplary methods (e.g., algorithms) facilitate speech synthesis, and the exemplary systems (e.g., special-purpose machines configured with special-purpose software) are configured to facilitate speech synthesis. The examples are merely typologies of variations that have been made possible. Unless expressly stated otherwise, the structures (e.g., structural components such as modules) are arbitrary and may be combined or subdivided, and the operations (e.g., in procedures, algorithms, or other functions) may be in different order, combined or subdivided. The following description includes numerous specific details to provide a thorough understanding of various exemplary embodiments for illustrative purposes. However, it will be apparent to those skilled in the art that the subject matter can be carried out without these specific details.
[0009] A machine (e.g., a mobile device or other computing machine) may be specifically configured (e.g., by appropriate hardware modules, software modules, or a combination of both) to operate as a speech synthesizer, such as a speech synthesizer with multimodal mixing, or to function in other ways. According to examples of systems and methods described herein, the machine presents a GUI on a touch-sensitive display screen (e.g., controlled by or communicating with a mobile device). The GUI depicts a word to be pronounced (e.g., "nap", "cat", or "tap") (e.g., as part of a phonics educational game or other application). The depicted word comprises a sequentially first alphabetic letter (e.g., "n") and a sequentially second alphabetic letter (e.g., "a"). The machine then detects the drag speed of a touch input on the touch-sensitive display screen and determines that the detected drag speed of the touch input falls within a first drag speed range of a plurality of drag speed ranges. Based on (for example, in response to) the detected drag speed falling within a first drag speed range, the machine selects whether to pronounce the word by sequentially playing at least the first and second audio files. The first audio file represents the first phoneme that pronounces the first sequential letter of the word. The second audio file represents the second phoneme that pronounces the second sequential letter of the word.
[0010] Each of the multiple drag speed ranges (e.g., multiple drag speed ranges) may be associated with a corresponding group of audio files (e.g., a stored bank). The multiple ranges subdivide the potential drag speed of the touch input into two or more classifications (e.g., categories or classes) so that they are defined (partitioned) by one or more threshold drag speeds, and each threshold drag speed demarks one or both of two adjacent ranges. For example, if the touch input is detected to have a slow drag speed, the machine identifies a first group of audio files (e.g., slow speed) and retrieves one or more audio files from that first group for playback. As another example, if the touch input is detected to have a fast drag speed, the machine identifies a second group of audio files (e.g., non-slow speed) and retrieves audio files from that second group for playback.
[0011] In some exemplary embodiments, three classifications are implemented for slow drag speed, medium drag speed, and fast drag speed, each corresponding to three groups of audio files. For example, the first group of audio files for slow drag speed may consist of individual audio files of individual recorded phonemes spoken at a normal speed. The second group of audio files for medium drag speed may consist of audio files of entire recorded words spoken at a slow speed (e.g., with overpronunciation in the articulation of each constituent phoneme). The third group of audio files for fast drag speed may consist of audio files of the same entire recorded word spoken at a normal speed (e.g., without overpronunciation) or at another speed faster than the slow speed.
[0012] The group of audio files selected by the machine may depend in part or entirely on the drag speed of the touch input. According to the systems and methods discussed herein, one available classification (e.g., category) of the drag speed (e.g., first drag speed range) corresponds to the sequential playback of individual audio files for pronouncing a word, where each sequentially played audio file corresponds to an individual phoneme of the word. As described above, the phonemes recorded in these single-phoneme audio files may be spoken at a normal speed. In some exemplary embodiments, one available classification (e.g., second drag speed range) corresponds to the playback of a single audio file for pronouncing a word, where multiple phonemes of a complete word are recorded in a single audio file. As described above, the multiple phonemes of the word recorded in this single audio file may be uttered at a slow speed (e.g., slower than the normal speed). In certain exemplary embodiments, one available classification (e.g., third drag speed range) corresponds to the playback of an alternative single audio file for pronouncing a word. The multiple phonemes of the complete word in this alternative single audio file may be spoken at a normal speed instead of a slow speed. In various exemplary embodiments, multiple phonemes of a word in this or further alternative single audio file are spoken at a fast rate (e.g., faster than normal).
[0013] In the presented GUI, the first alphabet character has a corresponding area (e.g., a first sub-area), which is configured to detect the drag speed of touch input (e.g., based on a first portion occurring in the corresponding area). The detected drag speed may be applied to the entire word or only to the portion corresponding to the first alphabet character. Similarly, the second alphabet character may have a corresponding area (e.g., a second sub-area), which is configured to detect or update the drag speed of touch input (e.g., based on a second portion occurring in the corresponding area). The detected or updated drag speed may be applicable only to the rest of the complete word or only to the portion corresponding to the second alphabet character.
[0014] In some exemplary embodiments, the GUI includes a slider bar (e.g., positioned below a word or otherwise visually adjacent to the word) that moves along the direction in which the word is read (e.g., the reading direction of the word as depicted in the GUI). The movement of the slider bar may be based on touch input representing the movement of the user's finger. Furthermore, the GUI may include a visual indicator that moves along the reading direction of the word based on components of the touch input (e.g., input components parallel to the reading direction of the word or other projected components) and moves at a speed based on (e.g., proportional to) the drag speed of the touch input. The visual indicator may also move in conjunction with the playback of one or more audio files selected by the machine to pronounce the word.
[0015] Figures 1 to 5 are drawings of a machine 100 (e.g., a device such as a mobile device) having a display screen 101 presenting a GUI_110 suitable for speech synthesis according to some exemplary embodiments. As shown in FIG. 1, the display screen 101 is touch-sensitive and is configured to receive one or more touch inputs from one or more fingers of a user (e.g., a child learning phonics by playing a phonics educational game). As an example, finger 140 is shown as touching the display screen 101 of the machine 100.
[0016] The GUI_110 is presented on the display screen 101. The GUI_110 depicts a word 120 (e.g., "nap" (nap. taking a nap), or alternatively "dog" (dog), "mom" (mom), "dad" (dad), "baby" (baby), "apple" (apple), "school" (school), or "backpack" (backpack) as depicted) to be pronounced by the machine 100 (which functions temporarily or permanently as a speech synthesis device), the user, or both of them. Also, the GUI_110 is shown as including a slider control 130 (e.g., a slider bar or other control area of the GUI_110). The slider control 130 may be visually aligned with the word 120. For example, both the slider control 130 and the word 120 may be along the same straight line (e.g., the direction in which the word 120 is read), or along two parallel straight lines (e.g., both in the direction in which the word 120 is read). As another example, both the slider control 130 and the word 120 can follow the same curve or two curves separated by a certain distance.
[0017] As shown in Figure 1, the slider control unit 130 may include a slide element 131 such as a position indicator bar or other visual indicator (e.g., a cursor or other marker) that shows the progress of pronunciation of a word 120, its constituent letters, its phonemes, or any appropriate combination thereof. As further shown in Figure 1, the word 120 comprises one or more alphabetic characters, and therefore may comprise a sequential first alphabetic character 121 (e.g., "n") and a sequential second alphabetic character 122 (e.g., "a") (e.g., among other text characters). The word 120 may further comprise a sequential third alphabetic character 123 (e.g., "p"). For example, the word 120 may be a consonant-vowel-consonant (CVC) word such as "nap" or "cat," and accordingly comprises a sequential first alphabetic character 121, a sequential second alphabetic character 122, and a sequential third alphabetic character 123, all ordered and aligned in the direction in which the word 120 is to be read.
[0018] Different sub-regions of the slider control unit 130 may correspond to different alphabetic characters of the word 120, or they may be used to detect or update the drag speed of a swipe touch input within the slider control unit 130. Each sub-region of the slider control unit 130 may be visually aligned to the corresponding alphabetic character of the word 120. Thus, referring to Figure 1, the first sub-region of the slider control unit 130 may correspond to the sequential first alphabetic character 121 (e.g., "n") and may be visually aligned to the sequential first alphabetic character 121. The second sub-region of the slider control unit 130 may correspond to the sequential second alphabetic character 122 (e.g., "a") and may be visually aligned to the sequential second alphabetic character 122. Similarly, the third sub-region of the slider control unit 130 may correspond to the sequential third alphabetic character 123 (e.g., "p") and may be visually aligned to the sequential third alphabetic character 123.
[0019] Furthermore, the GUI_110 may include a visual indicator 150 of the progress of Reading the word 120, Pronouncing the word 120, or both. The visual indicator 150 may be or include one or more visual elements indicating the degree to which the word 120 is read or pronounced. As shown in FIGS. 1-5, the visual indicator 150 is a vertical line. However, in various exemplary embodiments, the visual indicator 150 can include a change in color, a change in brightness, a change in fill pattern, a change in size, a change in position (e.g., a vertical displacement perpendicular to the direction in which the word 120 is read), a visual element (e.g., an arrowhead), or any suitable combination thereof.
[0020] As shown in FIG. 1, the finger 140 is performing a touch input (e.g., a swipe gesture or other touch-and-drag input) on the display screen 101. To initiate the touch input, the finger 140 touches the display screen 101 at a location within the slider control portion 130 (e.g., a first location that may be present within the first sub-region of the slider control portion 130), and the display screen 101 detects that the finger 140 is touching the display screen 101 at that location. Thus, a touch input is started (e.g., a touch-down) within the slider control portion 130. In response to detecting that the finger 140 is touching at the illustrated location within the GUI_110, the GUI_110 presents a slide element 131 at the same location.
[0021] In response to a portion (e.g., a first portion) of the touch input occurring within the first sub-region of the slider control unit 130, the machine 100 detects the drag speed of the touch input. The first sub-region of the slider control unit 130 may correspond to the sequential first alphabet character 121 of the word 120. The machine 100 then classifies the detected drag speed and, based on this, determines what blend mode to use to pronounce the word 120. For example, the machine 100 may choose to pronounce the word 120 by sequentially playing individual audio files, each containing a recording of the phonemes corresponding to the sequential first alphabet character 121, the sequential second alphabet character 122, and the sequential third alphabet character 123 of the word 120, or by using some alternative blend mode (e.g., by playing a single audio file containing a recording of the word 120 spoken as a whole).
[0022] As shown in Figure 2, finger 140 continues to perform touch input on display screen 101. At the time shown, finger 140 is touching display screen 101 at a certain location (e.g., a second location) within the slider control unit 130, and display screen 101 detects that finger 140 is touching the display screen 101 at that location. Therefore, the touch input continues to move within the slider control unit 130. In response to the detection that finger 140 is touching the location shown in GUI_110, GUI_110 presents slide element 131 at the same location. As described above, slide element 131, visual indicator 150, or both, can indicate the degree of progress achieved in the pronunciation of word 120 (e.g., progress to the pronunciation of phonemes corresponding to sequential first alphabet letters 121, as shown in Figure 2).
[0023] According to some exemplary embodiments, in response to a portion of the touch input (e.g., a second portion) occurring within a second sub-region of the slider control unit 130, the machine 100 detects or updates the drag speed of the touch input. The second sub-region of the slider control unit 130 may correspond to the second sequential alphabet character 122 of the word 120. The machine 100 then classifies the detected or updated drag speed and, based on that, may determine what mixed mode to use to pronounce the rest of the word 120 (e.g., from the second sequential alphabet character 122 onward, or otherwise, without including the first sequential alphabet character 121 corresponding to the first sub-region of the slider control unit 130). For example, the machine 100 may choose to pronounce the rest of word 120 by sequentially playing individual audio files, each containing a recording of the phonemes corresponding to the sequential second alphabet character 122 and the sequential third alphabet character 123 of word 120, or to pronounce the rest of word 120 using some alternative mixed mode (for example, by playing at least a portion of a single audio file containing a recording of word 120 being spoken as a whole).
[0024] As shown in Figure 3, finger 140 continues to perform touch input to the display screen 101. At the time shown, finger 140 is touching the display screen 101 at a certain position (e.g., the third position) within the slider control unit 130, and the display screen 101 detects that finger 140 is touching the display screen 101 at that position. Therefore, the touch input continues to move within the slider control unit 130. In response to the detection that finger 140 is touching the shown location within GUI_110, GUI_110 presents a slide element 131 at the same location. As described above, the slide element 131, the visual indicator 150, or both, can indicate the degree of progress achieved in the pronunciation of word 120 (e.g., progress to the pronunciation of the phonemes corresponding to the second alphabet letter 122 sequentially, as shown in Figure 3).
[0025] According to a particular exemplary embodiment, in response to a portion of the touch input (e.g., a third portion) occurring within a third sub-region of the slider control unit 130, the machine 100 detects or updates the drag speed of the touch input. The third sub-region of the slider control unit 130 may correspond to the sequential third alphabet character 123 of the word 120. The machine 100 then classifies the detected or updated drag speed and, based on this, may determine what mixed mode to use to pronounce the rest of the word 120 (e.g., from the sequential third alphabet character 123 onward, or otherwise without using the sequential first alphabet character 121 and the sequential second alphabet character 122 corresponding to the first and second sub-regions of the slider control unit 130). For example, the machine 100 may choose to pronounce the rest of the word 120 by sequentially playing one or more individual audio files. The one or more audio files store recordings of one or more phonemes (e.g., among the further alphabet characters) corresponding to the sequential third alphabet character 123. Alternatively, machine 100 may decide to pronounce the remaining portion of word 120 using some alternative mixed mode (for example, by playing at least a portion of a single audio file that stores a recording of word 120 being spoken as a whole).
[0026] As shown in Figure 4, finger 140 continues to perform touch input on the display screen 101. At the time shown, finger 140 is touching the display screen 101 at a certain position (e.g., the fourth position) within the slider control unit 130, and the display screen 101 detects that finger 140 is touching the display screen 101 at that position. Therefore, the touch input continues to move within the slider control unit 130. In response to the detection that finger 140 is touching the illustrated location within GUI_110, GUI_110 presents a slide element 131 at the same location. As described above, the slide element 131, the visual indicator 150, or both, can indicate the degree of progress achieved in the pronunciation of word 120 (e.g., progress to the pronunciation of phonemes corresponding to sequential third alphabet letters 123, as illustrated in Figure 4).
[0027] As shown in Figure 5, the touch input to the display screen 101 is terminated simply by lifting the finger 140 from the display screen 101 at a certain position (e.g., the fifth position) within the slider control unit 130, and the display screen 101 detects that the finger 140 has moved to this position on the display screen 101 and has subsequently stopped contacting the display screen 101. Thus, the touch input terminates its movement within the slider control unit 130. In response to detecting that the finger 140 has been lifted from the display screen 101 at the illustrated position within GUI_110, GUI_110 presents the slide element 131 at the same position. As described above, the slide element 131, the visual indicator 150, or both, can indicate the degree of progress achieved in the pronunciation of word 120 (e.g., progress toward completion, as illustrated in Figure 5).
[0028] Figure 6 is a block diagram showing components of a machine 100 (e.g., a device such as a mobile device) according to several exemplary embodiments. The machine 100 is shown to include a GUI generator 610, a touch input detector 620, a drag speed classifier 630, a speech synthesizer 640, and a display screen 101, all of which are configured to communicate with each other (e.g., via a bus, shared memory, or switches). The GUI generator 610 may be a GUI module or similar appropriate software code for generating GUI_110, or may include such modules. The touch input detector 620 may be a touch input module or similar appropriate software code for detecting one or more touch inputs (e.g., touch-and-drag inputs or swipe inputs) occurring on the display screen 101, or may include such modules. The drag speed classifier 630 may be a speed classifier module or similar appropriate software code for detecting, updating, or otherwise determining the drag speed of touch inputs. The speech synthesizer 640 may be a speech module or similar appropriate software code for pronouncing the word 120 (for example, via GUI_110, via the audio playback subsystem of machine 100, or via machine 100 or any part thereof, which may have both).
[0029] As shown in Figure 6, the GUI generator 610, the touch input detector 620, the drag speed classifier 630, the speech synthesizer 640, or any suitable combination thereof may form all or part of an application 600 (e.g., a mobile application) that is stored (e.g., installed) in the machine 100 (e.g., in response to or as a result of data received from one or more server machines via a network). Furthermore, one or more processors 699 (e.g., hardware processors, digital processors, or any suitable combination thereof) may be included (e.g., temporarily or permanently) in the application 600, the GUI generator 610, the touch input detector 620, the drag speed classifier 630, the speech synthesizer 640, or any suitable combination thereof.
[0030] Any one or more components (e.g., modules) described herein may be implemented using hardware alone (e.g., one or more processors 699) or a combination of hardware and software. For example, any component described herein may physically comprise one or more arrangements of processors 699 (e.g., a subset of processors 699, or between processors 699) configured to perform the operations described herein for that component. As another example, any component described herein may comprise software, hardware, or both, which configure one or more arrangements of processors 699 to perform the operations described herein for that component. Thus, different components described herein may comprise or configure different arrangements of processors 699 at different times, or constitute a single arrangement of processors 699 at different times. Each component (e.g., module) described herein is an example of means for performing the operations described herein for that component. Furthermore, any two or more components described herein may be combined into a single component, and the functions described herein for a single component may be subdivided among multiple components. Furthermore, according to various exemplary embodiments, components described herein as being implemented within a single system or machine (e.g., a single device) may be distributed among multiple systems or machines (e.g., multiple devices).
[0031] Machine 100 may be, may comprise, or may be otherwise implemented as a special-purpose computer (e.g., special, or otherwise non-conventional, non-generic) modified to perform one or more of the functions described herein (e.g., composed of or programmed by special-purpose software such as one or more software modules of a special-purpose application, operating system, firmware, middleware, or other software program). For example, a special-purpose computer system capable of carrying out one or more of the methodologies described herein is described later with respect to Figure 10, and such a special-purpose computer may, accordingly, be a means for carrying out one or more of the methodologies described herein. In the art of such special-purpose computers, a special-purpose computer that is specifically modified by the structure discussed herein (e.g., composed of special-purpose software) and performs the functions discussed herein is a technical improvement over other special-purpose computers that lack the structure discussed herein or are otherwise unable to perform the functions discussed herein. Thus, a special-purpose machine configured according to the systems and methods discussed herein provides an improvement over the art of similar special-purpose machines.
[0032] Accordingly, machine 100 may be implemented, in whole or in part, as a special-purpose (e.g., specialized) computer system, as will be described later with respect to Figure 10. According to various exemplary embodiments, machine 100 may be, or may comprise, a desktop computer, a vehicle computer, a home media system (e.g., a home theater system or other home entertainment system), a tablet computer, a navigation device, a portable media device, a smartphone, or a wearable device (e.g., a smart watch, smart glasses, smart clothing, or smart jewelry).
[0033] Figures 7 to 9 are flowcharts showing the operation of the machine 100 when performing the speech synthesis method 700 according to several exemplary embodiments. The operation in method 700 may be performed by the machine 100 using the components (e.g., modules) described above with respect to Figure 6, using one or more processors (e.g., microprocessors or other hardware processors), or any suitable combination thereof. As shown in Figure 7, method 700 comprises operations 710, 720, 730, and 740.
[0034] In operation 710, the GUI generator 610 generates GUI_110 and presents it on the display screen 101, or otherwise causes GUI_110 to be presented on the display screen 101. Execution of operation 710 can display GUI_110 as illustrated in Figure 1.
[0035] In operation 720, the touch input detector 620 detects the drag speed of the touch input (for example, via, using, in conjunction with, or otherwise) based on at least part of the drag speed of the touch input (for example, by detecting the speed at which the touch input is being dragged or otherwise moved on the display screen 101). Detection may be performed by measuring the drag speed of the touch input (for example, pixels per second, inches per second, or other appropriate units of speed). The execution of operation 710 may result in the GUI_110 being displayed as illustrated in Figure 2.
[0036] In operation 730, the drag speed classifier 630 determines the drag speed range to which the drag speed detected in operation 720 belongs. This has the effect of classifying the drag speed into one of several available drag speed ranges (e.g., a first drag speed range). For example, the drag speed classifier 630 may determine that the detected drag speed belongs to a first drag speed range (e.g., slow drag speeds) among two or more ranges (e.g., one or more categories of slow drag speeds and not-too-slow drag speeds).
[0037] In operation 740, based on the range determined in operation 730 (e.g., drag speed classification), the speech synthesizer 640 selects (e.g., chooses or makes other determinations) whether word 120 should be pronounced by sequential playback of audio files for individual phonemes (e.g., playback of at least a first audio file and a second audio file, where the first audio file represents the first phoneme pronouncing the sequential first alphabet character 121 of word 120, and the second audio file represents the second phoneme pronouncing the sequential second alphabet character 122 of word 120), as opposed to pronouncing word 120 through an alternative process (e.g., playback of a single audio file representing the phonemes of the sequential first alphabet character (121-123) of word 120 as a whole).
[0038] As shown in Figure 8, in addition to any one or more of the operations described above, method 700 may include one or more of operations 820, 822, 830, 840, 850, and 860. Operation 820 may be performed as part of operation 720 (e.g., a precursor task, subroutine, or part thereof) in which the touch input detector 620 detects the drag speed of the touch input. In operation 820, the touch input detector 620 detects the drag speed of the touch input based on a first portion of the touch input. For example, the first portion of the touch input may occur within a first sub-region of the slider control unit 130, and the touch input detector 620 may detect the drag speed based on the first portion occurring within the first sub-region. The first sub-region may correspond to the sequential first alphabet character 121 of a word 120, as presented in GUI_110.
[0039] In some exemplary embodiments, the drag speed of the touch input varies in parts, and accordingly, operation 720 may be repeated for additional sub-regions of the slider control unit 130. In such exemplary embodiments, operation 822 may be performed as part of a repeated instance of operation 720. In operation 822, the touch input detector 620 detects or updates the drag speed of the touch input based on a second part of the touch input. For example, the second part of the touch input may occur within a second sub-region of the slider control unit 130, and the touch input detector 620 may detect the drag speed based on the second part occurring within the second sub-region. The second sub-region may correspond to the second sequential alphabet character 122 of word 120, as presented in GUI_110.
[0040] Operation 830 may be performed as part of operation 730, in which case the drag speed classifier 630 determines the drag speed range into which the drag speed falls. In operation 830, the drag speed classifier 630 compares the drag speed with one or more threshold speeds (e.g., one or more threshold drag speeds that divide or otherwise define multiple available ranges of drag speed). For example, a first threshold drag speed may define an upper limit for a first drag speed range corresponding to a first category of drag speed (e.g., slow). Similarly, a second threshold drag speed may define an upper limit for a second drag speed range corresponding to a second category (e.g., medium or fast), and the second drag speed range may be adjacent to the first drag speed range.
[0041] Operation 840 may be performed as part of operation 740, in which case the speech synthesizer 640 selects whether to pronounce word 120 by sequentially playing audio files for each individual phoneme. This selection is made based on (for example, in response to) a range determined to include the detected drag speed of the touch input. One possible outcome is that the speech synthesizer 640 has chosen to actually pronounce word 120 by sequentially playing audio files for each individual phoneme, and this selection of the process for pronouncing word 120 is performed in operation 840.
[0042] In an exemplary embodiment in which operation 740 comprises operation 840, operation 850 may be performed after operation 740. In operation 850, the speech synthesizer 640 sequentially plays individual audio files for individual phonemes, or otherwise sequentially plays them (for example, one by one), in order to pronounce the word 120. For example, the speech synthesizer 640 may sequentially play at least a first audio file and a second audio file, where the first audio file records a first phoneme that pronounces the sequential first alphabet letter 121 of the word 120. The second audio file records a second phoneme that pronounces the sequential second alphabet letter 122 of the word 120.
[0043] In operation 860, the GUI generator 610 moves the visual indicator 150 along the direction in which the word 120 should be read. The visual indicator 150 may move in conjunction with touch input, in conjunction with sequential playback of audio files for individual phonemes, in conjunction with the speech synthesizer 640 pronouncing the word 120, or in any appropriate combination thereof.
[0044] As shown in Figure 9, in addition to any one or more of the operations described above, method 700 may include one or more of operations 940 and 950. In some exemplary embodiments, operation 940 includes operation 942 and operation 950 includes operation 952. In an alternative exemplary embodiment, operation 940 includes operation 944 and operation 950 includes operation 954.
[0045] Operation 940 may be performed as part of operation 740, in which the speech synthesizer 640 selects whether to pronounce word 120 by sequentially playing audio files for each individual phoneme. As described above, this selection is made based on (e.g., in response to) a range determined to include the detected drag speed of the touch input. One of the consequences of enabling this is that the speech synthesizer 640 chooses to pronounce word 120 by playing a single audio file that records all of the phonemes of word 120 (e.g., all phonemes) as a whole, rather than sequentially playing separate audio files for each individual phoneme, and this selection of the alternative process for pronouncing word 120 is performed in operation 940.
[0046] In an exemplary embodiment in which operation 740 comprises operation 940, operation 950 may be performed after operation 740. In operation 950, the speech synthesizer 640 plays such a single audio file, or otherwise causes it to play, in order to pronounce the word 120.
[0047] As described above, in some exemplary embodiments, operation 940 comprises operation 942 and operation 950 comprises operation 952. In operation 942, as part of selecting that word 120 be pronounced by playing a single audio file, the speech synthesizer 640 selects a third audio file to play in order to pronounce word 120. Here, the third audio file represents (e.g., recorded) phonemes corresponding to sequential first alphabet letters 121 to sequential third alphabet letters 123 of word 120 spoken at a slow speed (e.g., a speaking speed slower than normal speaking speed). In the corresponding operation 952, the speech synthesizer 640 pronounces word 120 by playing or causing the third audio file selected in operation 942 to play.
[0048] As described above, in a certain exemplary embodiment, operation 940 comprises operation 944 and operation 950 comprises operation 954. In operation 944, as part of selecting that word 120 be pronounced by playing a single audio file, the speech synthesizer 640 selects a fourth audio file to play in order to pronounce word 120. Herein, the fourth audio file represents (for example, recorded) the phonemes corresponding to sequential first alphabet letters 121 to sequential third alphabet letters 123 of spoken word 120, at a normal speed or at a speed faster than the slow speaking speed of the third audio file. In the corresponding operation 954, the speech synthesizer 640 pronounces word 120 by playing or causing the fourth audio file selected in operation 944 to play.
[0049] According to various exemplary embodiments, one or more of the methodologies described herein can facilitate the provision of a speech synthesizer having multiple modes for mixing phonemes together to pronounce a word. Furthermore, one or more of the methodologies described herein can facilitate the provision of a user-friendly experience in which the drag speed of touch input fully or partially controls which mixing mode is selected by the speech synthesizer. In particular, the drag speed of touch input is the basis for determining whether individual audio files for individual phonemes should be played, in contrast to other processes for pronouncing a word. Thus, one or more of the methodologies described herein can facilitate the pronunciation of words at a speed desired by the user, with enhanced clarity at slower speeds and enhanced smoothness at higher speeds, and can also facilitate the provision of at least one visual indicator (e.g., in its reading direction) of progress toward the completion of word pronunciation, compared to the capabilities of existing systems and methods.
[0050] Considering these effects together, one or more of the methodologies described herein may avoid the need for certain efforts or resources that would otherwise be involved in providing a speech synthesizer. The effort expended by a user in providing a dynamically adaptive speech synthesizer with multimodal mixing may be reduced by the use (e.g., reliance) of a special-purpose machine that implements one or more of the methodologies described herein. Computing resources used by one or more systems or machines (e.g., in a network environment) can similarly be reduced (compared to a system or machine that lacks, for example, the structures discussed herein or is otherwise unable to perform the functions discussed herein). Examples of such computing resources include processor cycles, network traffic, computing capacity, main memory usage, graphics rendering capacity, graphics memory usage, data storage capacity, power consumption, and cooling capacity.
[0051] Figure 10 is a block diagram illustrating components of a machine 1000 that, in several exemplary embodiments, can read an instruction 1024 from a machine-readable medium 1022 (e.g., a non-temporary machine-readable medium, a machine-readable storage medium, a computer-readable storage medium, or any suitable combination thereof) and execute, in whole or in part, one or more of the methodologies discussed herein. Specifically, Figure 10 shows the machine 1000 in an exemplary form of a computer system (e.g., a computer) in which an instruction 1024 (e.g., software, a program, an application, an applet, an app, or other executable code) can be executed, in whole or in part, to cause the machine 1000 to execute one or more of the methodologies discussed herein.
[0052] In alternative embodiments, machine 1000 may operate as a standalone device or be communicatively coupled to other machines (e.g., networked). In a networked deployment, machine 1000 may operate as a server machine or client machine in a server-client network environment, or as a peer machine in a distributed (e.g., peer-to-peer) network environment. Machine 1000 may be a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a mobile phone, a smartphone, a set-top box (STB), a personal digital assistant (PDA), a web appliance, a network router, a network switch, a network bridge, or any other machine capable of sequentially or otherwise executing instructions 1024 that specify the actions to be performed by that machine. Furthermore, although only a single machine is illustrated, the term “machine” shall also be interpreted as comprising any collection of machines that individually or collectively execute instructions 1024 to perform one or more of the methodologies discussed herein.
[0053] The machine 1000 comprises a processor 1002 (e.g., one or more central processing units (CPUs), one or more graphics processing units (GPUs), one or more digital signal processors (DSPs), one or more application-specific integrated circuits (ASICs), one or more high-frequency integrated circuits (RFICs), or any suitable combination thereof), main memory 1004, and static memory 1006, configured to communicate with each other via a bus 1008. The processor 1002 comprises solid-state digital microcircuits (e.g., electronic, optical, or both) that are temporarily or permanently configurable by some or all of the instructions 1024, such that the processor 1002 can be configured to execute one or more of the methodologies described herein, in whole or in part. For example, one or more sets of microcircuits of the processor 1002 may be configured to execute one or more modules (e.g., software modules) described herein. In some exemplary embodiments, the processor 1002 is a multi-core CPU (e.g., a dual-core CPU, a quad-core CPU, an 8-core CPU, or a 128-core CPU) in which each of the multiple cores operates as another processor capable of whole-or partially executing any one or more of the methodologies described herein. While the beneficial effects described herein may be provided by a machine 1000 having at least the processor 1002, these same beneficial effects may be provided by other types of machines that do not include a processor (e.g., a purely mechanical system, a purely hydraulic system, or a hybrid machine-hydraulic system), provided that such a processorless machine is configured to execute one or more of the methodologies described herein.
[0054] The machine 1000 may further include a graphics display (graphics display) 1010 (for example, a plasma display panel (PDP), a light-emitting diode (LED) display, a liquid crystal display (LCD), a projector, a cathode ray tube (CRT), or any other display capable of displaying graphics or video). The machine 1000 may also include an alphanumeric input device 1012 (for example, a keyboard or keypad), a pointer input device 1014 (for example, a mouse, touchpad, touchscreen, trackball, joystick, stylus, motion sensor, eye-tracking device, data glove, or other pointing device), data storage 1016, an audio generating device 1018 (for example, a sound card, amplifier, speakers, headphone jack, or any suitable combination thereof), and a network interface device 1020.
[0055] The data storage 1016 (e.g., data storage device) includes a machine-readable medium 1022 (e.g., a tangible and non-temporary machine-readable storage medium) in which instructions 1024 embodying any one or more methodologies or functions described herein are stored. The instructions 1024 may also exist, all or at least partially, in the main memory 1004, the static memory 1006, the processor 1002 (e.g., the processor's cache memory), or any suitable combination thereof, before or during their execution by the machine 1000. Thus, the main memory 1004, the static memory 1006, and the processor 1002 can be considered machine-readable media (e.g., tangible and non-temporary machine-readable media). The instructions 1024 may be transmitted or received over the network 1090 via the network interface device 1020. For example, the network interface device 1020 may communicate the instructions 1024 using any one or more transfer protocols (e.g., Hypertext Transfer Protocol (HTTP)).
[0056] In some exemplary embodiments, the machine 1000 may be a portable computing device (e.g., a smartphone, tablet computer, or wearable device) and may include one or more additional input components 1030 (e.g., sensors or gauges). Examples of such input components 1030 include image input components (e.g., one or more cameras), audio input components (e.g., one or more microphones), direction input components (e.g., a compass), position input components (e.g., a GPS receiver), orientation input components (e.g., a gyroscope), motion detection components (e.g., one or more accelerometers), altitude detection components (e.g., an altimeter), temperature input components (e.g., a thermometer), and gas detection components (e.g., a gas sensor). Input data collected by any one or more of these input components 1030 may be accessible and available for use by any of the modules described herein (with appropriate privacy notices and protections, such as opt-in or opt-out consent, implemented in accordance with user preferences, applicable regulations, or any appropriate combination thereof).
[0057] As used herein, the term “memory” refers to a machine-readable medium capable of storing data temporarily or permanently, and may include, but is not limited to, random-access memory (RAM), read-only memory (ROM), buffer memory, flash memory, and cache memory. While machine-readable medium 1022 is shown as a single medium in exemplary embodiments, the term “machine-readable medium” should be understood as comprising a single or multiple mediums (e.g., a centralized or distributed database, or associated caches and servers) capable of storing instructions. Furthermore, the term “machine-readable medium” is understood to include any medium, or combination of multiple media, capable of carrying (e.g., storing or communicating) instructions 1024 for execution by machine 1000, so that when executed by one or more processors of machine 1000 (e.g., processor 1002), machine 1000 executes one or more of the methodologies described herein, in whole or in part. Thus, “machine-readable medium” refers not only to a single storage device but also to a cloud-based storage system or storage network comprising multiple storage devices. The term “machine-readable medium” shall be interpreted as comprising, but not limited to, one or more tangible and non-temporary data repositories (e.g., data volumes) in exemplary forms such as solid memory chips, optical disks, magnetic disks, or any suitable combination thereof.
[0058] As used herein, “non-transient” machine-readable media specifically exclude those that propagate signals by themselves. According to various exemplary embodiments, instructions 1024 for execution by machine 1000 can be communicated via a carrier medium (e.g., a machine-readable carrier medium). Examples of such carrier media include a non-transient carrier medium (e.g., a non-transient machine-readable storage medium such as solid memory that is physically movable from one place to another) and a transient carrier medium (e.g., a carrier wave or other propagating signal that transmits the instructions 1024).
[0059] Certain exemplary embodiments are described herein as comprising modules. Modules can comprise software modules (e.g., code stored in a machine-readable medium or transmission medium or otherwise embodied), hardware modules, or any suitable combination thereof. A “hardware module” is a tangible (e.g., non-transient) physical component (e.g., a set of one or more processors) capable of performing specific operations, and may be configured or arranged in a particular physical manner. In various exemplary embodiments, one or more computer systems or one or more hardware modules thereof may be configured by software (e.g., an application or part thereof) as hardware modules that operate to perform the operations described herein for those modules.
[0060] In some exemplary embodiments, the hardware module may be implemented mechanically, electronically, hydraulically, or in any suitable combination thereof. For example, the hardware module may comprise a dedicated circuit or logic permanently configured to perform a specific operation. The hardware module may be, or comprise, a special-purpose processor such as an FPGA (FieldProgrammable GateArray) or ASIC. Alternatively, the hardware module may comprise programmable logic or circuit temporarily configured by software to perform a specific operation. As an example, the hardware module may comprise software contained within a CPU or other programmable processor. It will be understood that the decision to implement the hardware module mechanically, hydraulically, in a dedicated and permanently configured circuit, or in a temporarily configured circuit (e.g., configured by software), may be made based on cost and time considerations.
[0061] Therefore, the term “hardware module” should be understood to encompass tangible entities that are physically constructed and can be permanently configured (e.g., hardwired) or temporarily configured (e.g., programmed) to operate in or perform specific operations as described herein. Furthermore, as used herein, the term “hardware implementation module” refers to a hardware module. Considering exemplary embodiments in which a hardware module is temporarily configured (e.g., programmed), each hardware module does not need to be configured or instantiated in any single instance at any given time. For example, if a hardware module comprises a CPU configured by software to be a special-purpose processor, the CPU may be configured as different special-purpose processors (e.g., each contained in a different hardware module) at different times. Software (e.g., a software module) may accordingly configure one or more processors to become, for example, a particular hardware module in one instance at a time, or otherwise configured, and to become, or otherwise configured, different hardware modules in another instance at a different time.
[0062] Hardware modules can provide information to other hardware modules and receive information from other hardware modules. Therefore, the described hardware modules can be considered to be communicatively coupled. When multiple hardware modules exist simultaneously, communication can be achieved by signal transmission (e.g., via circuits and buses) between or within two or more hardware modules. In embodiments where multiple hardware modules are configured or instantiated at different times, communication between such hardware modules may be achieved, for example, through the storage and retrieval of information in a memory structure accessible to the multiple hardware modules. For example, one hardware module may perform an operation and store the output of that operation in a communicatively coupled memory (e.g., a memory device). Another hardware module can then access the memory at a later date to retrieve and process the stored output. Hardware modules can also initiate communication with input or output devices and operate on resources (e.g., a collection of information from computing resources).
[0063] Various operations of the exemplary methods described herein may be performed, at least partially, by one or more processors that are configured temporarily (e.g., by software) or permanently to perform the operations in question. Whether temporarily or permanently configured, such processors may constitute a processor implementation module that operates to perform one or more operations or functions described herein. As used herein, “processor implementation module” means a hardware module in which the hardware comprises one or more processors. Thus, since a processor is an example of hardware, and at least some operations of any one or more of the methods discussed herein may be performed by one or more processor implementation modules, hardware implementation modules, or any suitable combination thereof, the operations described herein may be at least partially processor implementations, hardware implementations, or both.
[0064] Furthermore, one or more such processors may perform operations in a “cloud computing” environment or as a service (e.g., within a “Software as a Service” (SaaS) implementation). For example, at least some operations within one or more of the methods discussed herein may be performed by a group of computers (e.g., as machines having processors), and these operations may be accessible via a network (e.g., the Internet) and through one or more appropriate interfaces (e.g., application programming interfaces (APIs)). The performance of a particular operation can be distributed among one or more processors, whether it resides only within a single machine or is deployed across multiple machines. In some exemplary embodiments, one or more processors or hardware modules (e.g., processor implementation modules) may be located in a single geographical location (e.g., a home environment, an office environment, or within a server farm). In other exemplary embodiments, one or more processors or hardware modules may be distributed across multiple geographical locations.
[0065] Throughout this specification, multiple instances may implement components, operations, or structures described as a single instance. While individual operations of one or more methods are illustrated and described as separate operations, one or more of these operations may be performed concurrently, and nothing requires the operations to be performed in the illustrated order. Structures and their functions shown as separate components or functions in the illustrated configurations may be implemented as combined structures or components with combined functions. Similarly, structures and functions presented as single components may be implemented as separate components and functions. These and other variations, modifications, additions, and improvements are within the scope of the subject matter of this specification.
[0066] Some parts of the subject matter discussed herein may be presented in terms of algorithms or symbolic representations of actions on data stored as bits or binary digital signals in memory (e.g., computer memory or other machine memory). Such algorithms or symbolic representations are examples of techniques used by those skilled in the art to communicate the nature of a task. As used herein, “algorithm” is a self-consistent sequence of actions or similar processes that lead to a desired result. In this context, algorithms and actions involve the physical actions of physical quantities. Typically, but not always, such quantities can take the form of electrical, magnetic, or optical signals that can be stored, accessed, transferred, combined, compared, or otherwise acted upon by a machine. It is sometimes convenient, primarily for reasons of common usage, to refer to such signals using words such as “data,” “content,” “bit,” “value,” “element,” “symbol,” “character,” “term,” “digit,” and “numeral.” However, these words are merely convenient labels and should be associated with appropriate physical quantities.
[0067] Unless otherwise specified, discussions herein using terms such as “access,” “process,” “detect,” “calculate,” “compute,” “determine,” “generate,” “present,” and “display” refer to operations or processes made executable by a machine (computer) that operates or transforms data represented as a physical quantity (e.g., electronic, magnetic, or optical) in a machine (e.g., volatile memory, non-volatile memory, or any appropriate combination thereof), register, or other machine component that receives, stores, transmits, or displays information. Furthermore, unless otherwise specified, the terms “a” or “an” are used herein to mean having one or more instances, as is common in patent literature. Finally, as used herein, the conjunction “or” means non-exclusive “or” unless otherwise specified.
[0068] The following listed descriptions illustrate various examples of methods, machine-readable media, and systems (e.g., machines, apparatus, or other devices) discussed herein.
[0069] The first embodiment provides a method comprising the following, the following being an example thereof. The method is, A process of presenting a graphical user interface (GUI) on a touch-sensitive display screen using one or more processors of a machine, wherein the GUI depicts a word to be pronounced, and the word comprises sequential first alphabet characters and sequential second alphabet characters; The process of detecting the drag speed of touch input on the touch-sensitive display screen using one or more processors of the machine, The process involves determining, using one or more processors of the machine, that the detected drag speed of the touch input falls within a first drag speed range among a plurality of drag speed ranges, and A method comprising: a step of selecting whether to pronounce a word by sequentially playing at least a first audio file and a second audio file based on whether the detected drag speed falls within a first drag speed range, wherein the first audio file represents a first phoneme that pronounces the sequential first alphabetic letter of the word, and the second audio file represents a second phoneme that pronounces the sequential second alphabetic letter of the word.
[0070] The second embodiment provides the method described in the first embodiment, and the following is an example thereof. The step of selecting whether the word should be pronounced by sequentially playing at least the first audio file and the second audio file comprises the step of selecting that the word should be pronounced by sequentially playing at least the first audio file and the second audio file, and The method further includes the step of causing the word to be pronounced by sequentially playing at least the first audio file and the second audio file.
[0071] The third embodiment provides the method described in the first embodiment, and the following is an example thereof. The step of selecting whether to pronounce the word by sequentially playing at least the first audio file and the second audio file comprises a step of selecting to pronounce the word by playing a third audio file that represents multiple phonemes of the word, without sequentially playing the first audio file and the second audio file, and The method further includes a step of causing the word to be pronounced by playing the third audio file without sequentially playing the first audio file and the second audio file.
[0072] The fourth embodiment provides a method according to any one of the first to third embodiments, the following being an example. The GUI indicates an area configured to accept the touch input, the area comprising a sub-area configured to detect the drag speed of the touch input based on a portion of the touch input, the portion occurring within the sub-area of the area of the GUI, and The step of determining whether the detected drag speed falls within the first drag speed range is based on the portion of the touch input within the sub-region of the area of the GUI.
[0073] The fifth embodiment provides a method according to any one of the first to fourth embodiments, and the following is an example thereof. The second drag speed range among the plurality of drag speed ranges is adjacent to the first drag speed range, and The step of determining whether the detected drag speed of the touch input falls within the first drag speed range includes comparing the detected drag speed with a threshold drag speed that separates at least one of the first drag speed range or the second drag speed range.
[0074] The sixth embodiment provides a method according to any one of the first to fifth embodiments, and the following is an example thereof. The first drag speed range among the plurality of drag speed ranges corresponds to sequential playback of at least the first audio file and the second audio file, and The second drag speed range among the plurality of drag speed ranges corresponds to the playback of a third audio file representing the plurality of phonemes of the word.
[0075] The seventh embodiment provides the method described in the sixth embodiment, and the following is an example thereof. The first audio file represents the first phoneme recorded at a first pronunciation speed different from the second pronunciation speed at which the word was recorded in the third audio file, and, The second audio file represents the second phoneme recorded at a first pronunciation speed different from the second pronunciation speed at which the word was recorded in the third audio file.
[0076] The eighth embodiment provides the method described in the seventh embodiment, and the following is an example thereof. The third drag speed range among the plurality of drag speed ranges corresponds to the playback of a fourth audio file that represents the plurality of phonemes of the word, recorded at a first pronunciation speed different from the second pronunciation speed at which the word was recorded in the third audio file.
[0077] The ninth embodiment provides a method according to any one of the first to eighth embodiments, and the following is an example thereof. The word depicted in the GUI has a direction in which it should be read. The touch input has input components parallel to the direction in which the word should be read, and The GUI includes a visual indicator that moves in the direction in which the word should be read, based on the input components of the touch input.
[0078] The tenth embodiment provides a method according to any one of the first to ninth embodiments, and the following is an example thereof. The GUI indicates an area configured to accept the touch input, the area comprising a first sub-area configured to detect the drag speed of the touch input based on a first portion of the touch input, the first portion occurring within the first sub-area and corresponding to the sequential first alphabetic character of the word. The region further comprises a second sub-region configured to update the drag speed of the touch input based on the second portion of the touch input, the second portion occurring within the second sub-region and corresponding to the second sequential alphabet character of the word. The step of selecting whether to pronounce the word by sequentially playing at least the first audio file and the second audio file is based on the detected drag speed of the first portion of the touch input. The method further comprises the step of selecting whether to pronounce the rest of the word by sequentially playing at least the second audio file based on the updated drag speed of the second portion of the touch input.
[0079] The eleventh embodiment provides a machine-readable medium (e.g., a non-temporary machine-readable storage medium) having instructions that, when executed by one or more processors of a machine, cause the machine to perform an operation comprising the following: The aforementioned operation is, A step of presenting a graphical user interface (GUI) on a touch-sensitive display screen, wherein the GUI depicts a word to be pronounced, and the word comprises sequential first alphabet characters and sequential second alphabet characters. The process of detecting the drag speed of touch input on the touch-sensitive display screen, A step of determining whether the detected drag speed of the touch input falls within a first drag speed range among a plurality of drag speed ranges, The process includes a step of selecting whether or not to pronounce the word by sequentially playing at least a first audio file and a second audio file based on whether the detected drag speed falls within the first drag speed range. The first audio file represents a first phoneme that pronounces the sequential first alphabet letter of the word. The second audio file represents a second phoneme that pronounces the sequential second alphabet letter of the word.
[0080] The twelfth embodiment provides the machine-readable medium described in the eleventh embodiment, and the following is an example thereof. The step of selecting whether the word should be pronounced by sequentially playing at least the first audio file and the second audio file comprises the step of selecting that the word should be pronounced by sequentially playing at least the first audio file and the second audio file, and The operation further includes the step of pronouncing the word by sequentially playing at least the first audio file and the second audio file.
[0081] The 13th embodiment provides the machine-readable medium described in the 11th embodiment, and the following is an example thereof. The step of selecting whether to pronounce the word by sequentially playing at least the first audio file and the second audio file includes the option of pronouncing the word by playing a third audio file that represents multiple phonemes of the word without sequentially playing the first audio file and the second audio file, and The operation further includes a step of causing the word to be pronounced by playing the third audio file without sequentially playing the first audio file and the second audio file.
[0082] The 14th embodiment provides a machine-readable medium as described in any one of the 11th to 13th embodiments, the following being an example thereof. The first drag speed range among the plurality of drag speed ranges corresponds to sequential playback of at least the first audio file and the second audio file, and The second drag speed range among the aforementioned multiple drag speed ranges corresponds to the playback of a third audio file representing multiple phonemes of the word.
[0083] The 15th embodiment provides a machine-readable medium as described in the 14th embodiment, and the following is an example thereof. The third drag speed range among the plurality of drag speed ranges corresponds to the playback of a fourth audio file that represents the plurality of phonemes of the word, which is recorded at a first pronunciation speed different from the second pronunciation speed at which the word was recorded in the third audio file.
[0084] The sixteenth embodiment provides a system (for example, a computer system) comprising the following: The system comprises one or more processors and memory for storing instructions. When an instruction is executed by at least one of the one or more processors, it causes the system to perform an operation comprising the following: The aforementioned operation is, A step of presenting a graphical user interface (GUI) on a touch-sensitive display screen, wherein the GUI depicts a word to be pronounced, and the word comprises sequential first alphabet characters and sequential second alphabet characters. The process of detecting the drag speed of touch input on the touch-sensitive display screen, A step of determining whether the detected drag speed of the touch input falls within a first drag speed range among a plurality of drag speed ranges, The process includes a step of selecting whether or not to pronounce the word by sequentially playing at least a first audio file and a second audio file based on whether the detected drag speed falls within the first drag speed range. The first audio file represents a first phoneme that pronounces the sequential first alphabet letter of the word. The second audio file represents a second phoneme that pronounces the sequential second alphabet letter of the word.
[0085] The 17th embodiment provides the system described in the 16th embodiment, and the following is an example thereof. The step of selecting whether the word should be pronounced by sequentially playing at least the first audio file and the second audio file comprises the step of selecting whether the word should be pronounced by sequentially playing at least the first audio file and the second audio file, and The operation further includes a step of causing the word to be pronounced by sequentially playing at least the first audio file and the second audio file.
[0086] The 18th embodiment provides the system described in the 16th embodiment, and the following is an example thereof. The step of selecting whether the word should be pronounced by sequentially playing at least the first audio file and the second audio file comprises the step of selecting whether the word should be pronounced by playing a third audio file representing multiple phonemes of the word without sequentially playing the first audio file and the second audio file, and The operation further includes a step of causing the word to be pronounced by playing the third audio file without sequentially playing the first audio file and the second audio file.
[0087] The 19th embodiment provides the system described in any one of the 16th to 18th embodiments, and the following is an example thereof. The first drag speed range among the plurality of drag speed ranges corresponds to sequential playback of at least the first audio file and the second audio file, and The second drag speed range among the aforementioned multiple drag speed ranges corresponds to the playback of a third audio file that represents multiple phonemes of the word.
[0088] The 20th embodiment provides the system described in the 19th embodiment, and the following is an example thereof. The first audio file represents the first phoneme recorded at a first pronunciation speed different from the second pronunciation speed at which the word was recorded in the third audio file, and The second audio file represents the second phoneme recorded at a first pronunciation speed different from the second pronunciation speed at which the word was recorded in the third audio file.
[0089] The 21st embodiment provides a carrier medium for carrying machine-readable instructions for controlling a machine to perform an operation (e.g., a method operation) that is performed in any one of the embodiments described above.
Claims
1. presenting, by one or more processors of the machine, a graphical user interface (GUI) on a touch-sensitive display screen, the GUI depicting a word to be pronounced, the word comprising a sequential first alphabetic character and a sequential second alphabetic character; detecting, by one or more processors of the machine, a drag velocity of a touch input on the touch-sensitive display screen; determining, by one or more processors of the machine, that the detected drag velocity of the touch input falls within a first drag velocity range of a plurality of drag velocity ranges; and selecting, by one or more processors of the machine, whether to pronounce the word by sequentially playing at least a first audio file and a second audio file based on the detected drag speed falling within the first drag speed range, the first audio file representing first phonemes that pronounce the sequential first alphabetical letters of the word, and the second audio file representing second phonemes that pronounce the sequential second alphabetical letters of the word; A method of providing
2. selecting whether the word should be pronounced by sequentially playing at least the first audio file and the second audio file comprises selecting that the word should be pronounced by sequentially playing at least the first audio file and the second audio file; and The method further comprises the step of pronouncing the word by sequentially playing at least the first audio file and the second audio file. The method of claim 1.
3. selecting whether the word should be pronounced by sequentially playing at least the first audio file and the second audio file comprises selecting to pronounce the word by playing a third audio file representing a plurality of phonemes of the word without sequentially playing the first audio file and the second audio file; and The method further includes pronouncing the word by playing the third audio file without sequentially playing the first audio file and the second audio file. The method of claim 1.
4. the GUI presents an area configured to accept the touch input; the region comprises a sub-region configured to detect the drag velocity of the touch input based on a portion of the touch input, the portion occurring within the sub-region of the region of the GUI; and determining that the detected drag velocity falls within the first drag velocity range is based on the portion of the touch input within the sub-region of the region of the GUI; The method of claim 1.
5. a second drag speed range of the plurality of drag speed ranges adjacent to the first drag speed range; and determining that the detected drag velocity of the touch input falls within the first drag velocity range comprises comparing the detected drag velocity to a threshold drag velocity that defines at least one of the first drag velocity range or the second drag velocity range; The method of claim 1.
6. the first drag speed range of the plurality of drag speed ranges corresponds to sequential playback of at least the first audio file and the second audio file; and a second drag speed range of the plurality of drag speed ranges corresponds to the playback of a third audio file representing a plurality of phonemes of the word; The method of claim 1.
7. the first audio file represents the first phoneme recorded at a first pronunciation rate that is different from a second pronunciation rate at which the word was recorded in the third audio file; and the second audio file represents the second phoneme recorded at the first pronunciation rate, which is different from the second pronunciation rate at which the word was recorded in the third audio file. The method of claim 6.
8. a third drag speed range of the plurality of drag speed ranges corresponds to playback of a fourth audio file representing the plurality of phonemes of the word recorded at the first pronunciation speed, which is different from the second pronunciation speed at which the word was recorded in the third audio file. The method of claim 7.
9. the words depicted in the GUI have a direction in which the words should be read; the touch input has an input component parallel to the direction in which the word is to be read; and the GUI includes a visual indicator that moves in a direction in which the word should be read based on the input component of the touch input. The method of claim 1.
10. the GUI presents an area configured to accept the touch input, the area comprising a first sub-area configured to detect the drag velocity of the touch input based on a first portion of the touch input, the first portion occurring within the first sub-area and corresponding to the sequential first alphabetic character of the word; the region further comprising a second sub-region configured to update the drag velocity of the touch input based on a second portion of the touch input, the second portion occurring within the second sub-region and corresponding to the sequential second alphabetic character of the word; selecting whether to pronounce the word by sequentially playing at least the first audio file and the second audio file based on the detected drag velocity of the first portion of the touch input; and The method further includes selecting whether to pronounce the remainder of the word by sequentially playing at least the second audio file based on the updated drag velocity of the second portion of the touch input. The method of claim 1.
11. 1. A machine-readable medium comprising instructions that, when executed by one or more processors of a machine, cause the machine to perform operations, the operations including: presenting a graphical user interface (GUI) on a touch-sensitive display screen, the GUI depicting a word to be pronounced, the word comprising a sequential first alphabetic character and a sequential second alphabetic character; detecting a drag speed of a touch input on the touch-sensitive display screen; determining that the drag velocity of the detected touch input falls within a first drag velocity range of a plurality of drag velocity ranges; selecting, based on the detected drag speed falling within the first drag speed range, whether to pronounce the word by sequentially playing at least a first audio file and a second audio file, the first audio file representing first phonemes that pronounce the sequential first alphabetical letters of the word, and the second audio file representing second phonemes that pronounce the sequential second alphabetical letters of the word; 1. A machine-readable medium comprising:
12. selecting whether the word should be pronounced by sequentially playing at least the first audio file and the second audio file comprises selecting that the word should be pronounced by sequentially playing at least the first audio file and the second audio file; and the operations further include pronouncing the word by sequentially playing at least the first audio file and the second audio file. The machine-readable medium of claim 11.
13. selecting whether to pronounce the word by sequentially playing at least the first audio file and the second audio file comprises selecting to pronounce the word by playing a third audio file representing a plurality of phonemes of the word without sequentially playing the first audio file and the second audio file; and the operations further include pronouncing the word by playing the third audio file without sequentially playing the first audio file and the second audio file. The machine-readable medium of claim 11.
14. the first drag speed range of the plurality of drag speed ranges corresponds to sequential playback of at least the first audio file and the second audio file; and a second drag speed range of the plurality of drag speed ranges corresponds to the playback of a third audio file representing a plurality of phonemes of the word; The machine-readable medium of claim 11.
15. a third drag speed range of the plurality of drag speed ranges corresponds to playback of a fourth audio file representing the plurality of phonemes of the word recorded at a first pronunciation speed that is different from a second pronunciation speed at which the word was recorded in the third audio file; 15. The machine-readable medium of claim 14.
16. 1. A system comprising one or more processors and a memory for storing instructions, The instructions, when executed by at least one of the one or more processors, cause the system to perform an operation, the operation comprising: presenting a graphical user interface (GUI) on a touch-sensitive display screen, the GUI depicting a word to be pronounced, the word comprising a sequential first alphabetic character and a sequential second alphabetic character; detecting a drag speed of a touch input on the touch-sensitive display screen; determining that the drag velocity of the detected touch input falls within a first drag velocity range of a plurality of drag velocity ranges; selecting, based on the detected drag speed falling within the first drag speed range, whether to pronounce the word by sequentially playing at least a first audio file and a second audio file, the first audio file representing first phonemes that pronounce the sequential first alphabetical letters of the word, and the second audio file representing second phonemes that pronounce the sequential second alphabetical letters of the word; A system comprising:
17. selecting whether the word should be pronounced by sequentially playing at least the first audio file and the second audio file comprises selecting that the word should be pronounced by sequentially playing at least the first audio file and the second audio file; and the operations further include pronouncing the word by sequentially playing at least the first audio file and the second audio file.
17. The system of claim 16.
18. selecting whether the word should be pronounced by sequentially playing at least the first audio file and the second audio file comprises selecting that the word should be pronounced by playing a third audio file representing a plurality of phonemes of the word without sequentially playing the first audio file and the second audio file; and the operations further include pronouncing the word by playing the third audio file without sequentially playing the first audio file and the second audio file.
17. The system of claim 16.
19. the first drag speed range of the plurality of drag speed ranges corresponds to sequential playback of at least the first audio file and the second audio file; and a second drag speed range of the plurality of drag speed ranges corresponding to playback of a third audio file representing a plurality of phonemes of the word; 17. The system of claim 16.
20. the first audio file represents the first phoneme recorded at a first pronunciation rate that is different from a second pronunciation rate at which the word was recorded in the third audio file; and the second audio file represents the second phoneme recorded at the first pronunciation rate, which is different from the second pronunciation rate at which the word was recorded in the third audio file.
20. The system of claim 19.