Graphical vocalization adjustments recommended
A computing device enhances vocal quality by collecting data to identify and graphically guide facial adjustments, addressing the limitations of conventional systems in transitioning users from incorrect to correct vocal qualities.
Patent Information
- Application Number
- JP2023559107
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-03-29
- Filing Date
- 2022-02-11
- Publication Date
- 2025-10-09
- Estimated Expiration
- 2042-02-11
AI Technical Summary
Conventional computing systems struggle to help users effectively transition from incorrect to correct vocal qualities, such as pronunciation or singing, as they lack real-time visual guidance for facial adjustments.
A computing device collects audio and spatial data to identify facial elements' positions, determining alternative positions to change vocal qualities, and provides real-time graphical representations for users to adjust their facial movements.
Enables users to improve vocal qualities by providing immediate visual guidance, facilitating the transition from undesired to desired vocal characteristics through dynamic, step-by-step adjustments.
Smart Images

Figure 0007751956000001 
Figure 0007751956000002 
Figure 0007751956000003
Abstract
Description
[Background technology]
[0001] Each language has specific audible qualities (e.g., phonemes) that are required for speech to match a predetermined pronunciation. These audible qualities are often controlled by the shape that a user's mouth makes when speaking or producing sounds. For example, certain audible qualities may require the lips of the mouth to define a particular shape, the tongue to define a particular shape, the tongue to be in a particular configuration in the mouth (e.g., touching the upper teeth or elevated or recessed), the user's jaw to be dropped to increase the space in the mouth, etc. In some situations, a person can learn "incorrect" pronunciations (e.g., pronunciations that do not match predetermined pronunciations defined by a dictionary) and must learn how to change their physical speaking style in order to pronounce words "correctly." Summary of the Invention
[0002] Aspects of the present disclosure relate to methods, systems, and computer program products related to providing a user with a graphical representation of suggested facial adjustments, where the adjustments are determined to alter detected qualities of the user's vocalizations. For example, the method includes receiving audio data of a user's voice as the user vocalizes during a time period. The method also includes receiving spatial data of the user's face during the time period. The method also includes using the spatial data to identify positions of facial elements relative to other facial elements during the time period, where the relative positions of the elements cause a plurality of qualities of the user's voice. The method also includes identifying a subset of positions of the one or more elements that cause a first detected quality of the plurality of qualities during the time period. The method also includes determining alternative positions of the one or more elements that are determined to cause the user's voice to have a second quality of the plurality of qualities rather than the first quality. The method also includes providing the user with a graphical representation of the face showing one or more adjustments from the subset of positions to the alternative positions. Systems and computer program products configured to perform the operations of the method are also provided herein.
[0003] The above summary is not intended to describe each illustrated embodiment or every embodiment of the present disclosure.
[0004] The drawings included in this application are incorporated into and constitute a part of this specification. They illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the disclosure. The drawings are merely illustrative of particular embodiments and are not intended to limit the disclosure. [Brief explanation of the drawings]
[0005] [Figure 1] 1 illustrates a conceptual diagram of an example system in which a controller collects information about a user as the user speaks, enabling the controller to graphically provide suggested facial adjustments to the user, which adjustments are determined to alter the detected quality of the user's voice. [Figure 2A] 2A-2C illustrate examples of adjustments to the shape of a user's mouth that the controller of FIG. 1 may graphically provide. [Figure 2B] 2A-2C illustrate examples of adjustments to the position of a user's cheekbones that the controller of FIG. 1 may graphically provide. [Figure 2C] 2A-2C illustrate examples of adjustments to the position of a user's tongue that the controller of FIG. 1 may graphically provide. [Figure 3] FIG. 2 is a conceptual box diagram illustrating an example of components of the controller of FIG. 1. [Figure 4] 10 illustrates an exemplary flowchart by which the controller of FIG. 1 can graphically provide facial adjustment recommendations for modifying a user's voice quality.
[0006] While the invention is amenable to various modifications and alternative forms, specific features thereof have been shown by way of example in the drawings and will be described in detail. It is to be understood, however, that it is not intended to limit the invention to the particular embodiments described. On the contrary, it is intended to cover all modifications, equivalents, and alternatives falling within the scope of the invention. DETAILED DESCRIPTION OF THE INVENTION
[0007] Aspects of the present disclosure relate to providing recommendations for improving vocalizations, and more particular aspects of the present disclosure relate to collecting spatial and audio data of a user while the user is vocalizing and providing a graphical representation of ways the user can physically change how they vocalize to change the quality of their vocalizations. While the present disclosure is not necessarily limited to such applications, various aspects of the present disclosure can be understood through a discussion of various examples using this context.
[0008] A human voice has certain linguistic characteristics that are desirable in various contexts. For example, there may be certain qualities, such as certain phonemes, that a user must produce in order to correctly pronounce a word in a given language (e.g., the correct pronunciation matches a predetermined pronunciation found in a dictionary). In another example, there may be certain qualities, such as timbre or tone, that a user must produce in order to sing in a particular way (e.g., an operatic way). As used herein, a human vocal quality is an audible characteristic that depends on the shape that the oral cavity makes while the person is producing that quality (e.g., as that shape is produced by the relative orientation of the jaw, mouth, lips, tongue, etc.). Generally, for purposes of discussion, this disclosure will be described in terms of a first quality (e.g., this first quality will be discussed herein as being a technically incorrect dictionary definition of an "airy" singing voice or both) and a second quality (this second quality will be technically correct or generally desirable or both), although it will be understood that aspects of this disclosure relate to detecting and addressing hundreds (or more) of different qualities and / or the importance of those qualities as discussed herein.
[0009] When people learn to speak or sing a language, they often don't achieve perfect pronunciation or form the first time because they don't create the correct sequence of physical movements with their face. For example, a person may have difficulty producing one or more phonemes because they make an incorrect shape with their mouth. For example, they may not define the correct shape where their tongue touches the back of their upper incisors, resulting in an incorrect pronunciation of "l" (e.g., as in the case of a "y"), an inability to pronounce a "rolled r" (e.g., using an alveolar trill, alveolar flap, retroflex trill, or palatal trill), or difficulty pronouncing diagonal letters (e.g., pronouncing "three" as "tree").
[0010] People sometimes try to solve this by using the services of a human specialist, such as a speech-language pathologist or voice coach. The specialist may be highly specialized and trained to teach people how to create specific shapes in their mouths to improve the quality of their speech or singing. Because relearning how to pronounce certain words can be difficult (especially after a person has been pronouncing / speaking them a certain way for a long time), this can be a frustrating and challenging process for some people. Furthermore, because the human specialist cannot physically have someone else define the shape, they often must rely on verbal instructions for defining the new shape, demonstrating the new shape to themselves or a doll / mannequin, or some combination thereof. Thus, in some instances, it may be very difficult for someone who speaks with a first quality to understand how to physically change their speech to speak with a second quality.
[0011] Some conventional computing systems have attempted to streamline this situation by automating parts of the process of detecting users who speak with an "incorrect" quality. For example, conventional computing systems can calculate the amount of difference between the correct quality of speech as a person speaks and the incorrect quality of speech. In some instances, conventional computing systems can further identify one or more aspects of the way the person speaks that are physically incorrect. However, while these conventional computing systems may be helpful in correctly diagnosing a speech problem, a person who is speaking with an incorrect / undesired quality may find these conventional computing systems useless in helping them learn how to speak with a correct / desired quality instead.
[0012] Aspects of the present disclosure may solve or address some or all of these problems. A computing device including a processor executing instructions stored on a memory can provide functionality to address these problems, and this computing device is referred to herein as a controller. The controller can collect audio data and spatial data related to a user's face, oral cavity, or both, while the user is vocalizing. The controller can use this collected data to identify relative positions of elements on the user's face while the user is vocalizing (e.g., speaking or singing), including tracking these positions over time as they correspond to different vocalizations. For example, the controller can generate a vector diagram of these elements corresponding to the user's face. The controller may determine whether there are qualities of the user's voice that do not meet certain thresholds. The controller can determine whether any qualities do not meet certain thresholds by analyzing the audio data, by analyzing the spatial data, or by a combination thereof. If any qualities do not meet these thresholds, the controller can determine alternative positions for the determined elements that would change the vocalization from an undesired first quality to a desirable second quality. The controller can then provide a graphical representation of these alternative positions.
[0013] For example, the controller can provide an augmented image of the user that provides a graphical representation of these alternative positions. The controller may provide one or more augmented images of the user; for example, the controller may show one or more of a front view, a side view, or an internal view of the user's mouth, or a combination thereof, while the user is speaking. The controller can use an augmented reality device to provide a graphical representation in real time on a current image of the user's face. The augmented image of the user can highlight or instruct the user on certain changes to make in the formation / shape / movement of the mouth / tongue / face to change from an undesirable first quality to a desirable second quality. In some examples, the controller can functionally provide this graphical representation in real time (e.g., such that the positions of the user's elements are detected and analyzed and responsive alternative positions are provided within milliseconds of the user making a face that defines the positions of these elements), allowing the user to receive immediate and dynamic visual guidance regarding the formation / shape / movement of the mouth / tongue / face.
[0014] Additionally, while this discussion has been primarily in terms of the controller suggesting facial adjustments to change from a first initial quality to a second final quality, in some examples, the controller may suggest the adjustment to achieve the second quality as the first step in a series of steps toward perfect / preferred pronunciation. For example, the controller may determine that 12 adjustments should be made to the user's utterance to match the perfect / preferred pronunciation. The controller may further identify that suggesting 12 adjustments at once may be too confusing and / or difficult. Thus, the controller may identify a series of adjustments for the user to make over time (e.g., from a first initial quality to a second improved quality, a third further improved quality, a fourth final perfect quality), with each adjustment building on the previous adjustments. The controller may present these steps over time as the user masters each individual step. In some examples, such progression may include specific "vector points" (e.g., where an element moves from an initial position to a predetermined intermediate position), and after the user demonstrates that they can achieve the intermediate vector points, the controller may allow the vector points to be further extended along the calculated vector that ultimately results in a preferred final vocal quality, as discussed herein.
[0015] Beyond this, aspects of the present disclosure may be used to diagnose and / or provide treatment for conditions related to facial movement. For example, aspects of the present disclosure may detect a relative drooping of the smile on one side of a given user's face over time, which may indicate a stroke. The present disclosure may detect such a medical condition and provide an alert to a responsible party in response to the detection. Alternatively, or additionally, or both, aspects of the present disclosure may provide a graphical representation of alternative positions that a user can make that reflect facial movements configured to improve facial mobility after a medical incident (e.g., after a stroke or a paralytic event). Similarly, if the controller determines that a user has a condition that affects facial movement and also determines that the user has vocalizations with a quality that the user would prefer to change, the controller may identify and provide actions the user can still perform (e.g., taking into account the medical condition) that would change the first quality to a preferred second quality (e.g., if the preferred second quality is not entirely a third quality).
[0016] Additionally, while this disclosure primarily discusses how to improve a speaker according to a predetermined "dictionary" definition of perfection, those skilled in the art will appreciate that aspects of the present disclosure may also include having the controller suggest facial adjustments configured to improve a user's singing voice, having the controller assist the user in replicating a particular type of smile that the user prefers (e.g., by instructing the user to lift the corners of their mouth, or by instructing the user to lift their cheekbones and "smile with their eyes"), assisting in the creation of a regional accent (e.g., if an actor is trying to perfect a Boston accent for a movie), etc.
[0017] For example, FIG. 1 illustrates an environment 100 in which a controller 110 collects audio and spatial data of a user 120 as the user 120 vocalizes (e.g., speaks, sings, etc.) and provides a graphical representation of the determined adjustments. The controller 110 may include a processor coupled to a memory (as shown in FIG. 3) that stores instructions that cause the controller 110 to perform the operations discussed herein. While the controller 110 in FIG. 1 is depicted and discussed as a single component for purposes of discussion, in other examples, the controller 110 may include (or otherwise be a part of) multiple computing systems that work together in some manner to perform the functions described herein. Furthermore, while the controller 110 is generally discussed as providing the functionality herein as hosted within a computing device in proximity to the user 120, it should be understood that in other examples, some or all of the functionality of the controller 110 may additionally, alternatively, or both be virtualized so that the functionality may be aggregated from multiple distributed hardware components. In such a distributed case, the functionality provided by the controller 110 may be offered as a service in a cloud computing arrangement (e.g., via a web portal on the user device).
[0018] Controller 110 may provide a graphical representation of the adjustment to display 130. Display 130 may be a computing device configured to graphically present images generated by controller 110. For example, display 130 may include a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, etc. In some examples, display 130 is a standalone device (e.g., a computer screen or a television), while in other examples, display 130 may be integrated into a device that houses controller 110 or sensor 140, etc. (e.g., such that display 130 is part of a mobile phone).
[0019] The controller 110 can collect audio and / or spatial data of the user 120 via the sensor 140. The sensor 140 can include a microphone with sufficient accuracy to capture the user's 120 vocalizations, including identifying the user's ability to enunciate words in a predetermined style or sing with a particular tone, or both. The sensor 140 can also include a camera configured to capture many or all of the elements of the user's 120 face that contribute to the user's 120 voice having one or more characteristics. For example, the elements can include different points along the user's 120 lips, cheeks, chin, or tongue, and the different positions of these elements relative to the user's 120 face can change the quality of the user's 120 voice. Examples of this might include a lowered chin changing the quality of the voice, pursed lips changing the quality of the voice, or raised cheekbones changing the quality of the voice. The controller 110 can use multiple elements to identify positional data associated with each distinct facial feature (e.g., each lip, cheekbone, and chin). For example, the controller 110 may use spatial data of a single lip of the user 120 collected by the camera sensor 140 to determine five, ten, or more elements that are on that single lip and that will change the quality of the speech if any of these elements are deviated during speech.
[0020] In some examples, sensor 140 may include a computing device capable of capturing data regarding the position of user 120's tongue, the shape defined by user 120's oral cavity, or both. For example, sensor 140 may include a computing device configured to identify spatial data through user 120's skin from adjacent user 120's face or cheek, such as an ultrasound device. As another example, sensor 140 may include a device that can fit inside user 120's mouth, such as a mouth guard or dental retainer. Such a mouth guard or retainer may be placed over the top or bottom teeth or both, and configured to detect when the tongue touches the mouth guard / retainer (as it would otherwise). In other examples, sensor 140 may include a tongue sleeve that can be worn over the tongue to detect when the tongue touches the teeth, to detect the shape of the tongue, or both. In certain examples, the sensor 140, which can be inserted into the mouth of the user 120, may further be detected to measure the distance between itself and other objects in the mouth (e.g., the outer periphery of the oral cavity defined by the soft palate, tongue, uvula, etc.), thereby allowing the controller 110 to determine a partial or full three-dimensional map of the user's oral cavity.
[0021] The controller 110 can analyze the spatial and / or audio data collected by the sensors 140 to compile a profile of the user 120. This profile can include a spatial map of the user's 120's face and / or mouth. For example, the controller 110 can generate a vector diagram 132 of the user 120 as it appears on the display 130. The vector diagram 132 can include multiple nodes 134 located at each element of the user's 120's face, with the nodes 134 connected via vectors 136. In some examples, the controller 110 can create a vector diagram 132 with a higher density of nodes 134 in areas that cause different voice qualities (e.g., more nodes 134 near the user's 120's mouth than near the user's 120's forehead). However, the controller 110 may create nodes 134 in specific locations unrelated to the quality of the user's 120's voice (e.g., at the user's 120's ears) to better position the vector diagram 132 on the user 120 as the user 120 naturally moves their head when speaking (e.g., the user's 120's head nods back and forth, shakes, etc.). It will be understood that the specific number and arrangement of nodes 134 and vectors 136 within the vector diagram 132 is provided purely for illustrative purposes, and that more or fewer nodes 134 and vectors 136 provided in different locations are contemplated for purposes consistent with the present disclosure.
[0022] In some examples, the controller 110 may connect nodes 134 via vectors 136 if movement of one node 134 causes movement of each connected node 134. Stated differently, the controller 110 may connect two nodes 134 if movement of the first node 134 essentially causes movement of the second node 134. In other examples, the controller 110 may concentrate nodes 134 in locations that the controller 110 determines the user 120 should focus on so that the controller 110 can better detect finer-grained adjustments the user 120 needs to make (and so that additional computing power is not used and / or "wasted" on calculating and / or depicting nodes 134 and / or vectors 136 that are determined to be relatively unimportant to the user 120).
[0023] The controller 110 analyzes the spatial and audio data of the user 120 to identify whether the user 120 is speaking with a first quality. As discussed herein, a first quality includes a quality that does not conform to technical standards (e.g., a dictionary pronunciation guide) or widely held preferences (e.g., a rich, full timbre of a singing voice). The controller 110 may determine whether the user 120 has such a first quality by comparing the audio and / or spatial data of the user 120's speech with a corpus 150, which includes a large-scale structured repository of data on the speech of a significant number of people. The controller 110 may compare the user 120's speech with data from people in the corpus 150 classified as speaking with a first quality (e.g., if a word is not pronounced according to a dictionary definition, if the singing voice has a timbre or quality that is perceived as nasal or airy, etc.). If controller 110 determines that user 120's audible quality matches (e.g., matches with a threshold amount of confidence) past audio recordings in corpus 150 that are classified in corpus 150 as reflecting such a first quality, controller 110 may determine that user 120's utterance has a first quality.
[0024] Similarly, controller 110 may compare information regarding the relative positions of user 120's facial elements during an utterance with the relative facial positions of people stored in corpus 150. For example, corpus 150 may be structured such that several stored historical relative facial positions during an utterance of a particular word are classified as eliciting a first quality utterance. Controller 110 may compare the relative positions of facial elements as discussed herein with these classified historical relative positions in corpus 150 such that if the relative positions of user 120's elements match historical element positions classified as indicative of the first quality, controller 110 may determine that user 120's utterance is of the first quality.
[0025] In some examples, the controller 110 can determine the importance of a first quality as exhibited by the user's 120 vocalization. For example, some vocalization qualities may be non-binary, such that a measure may not be a measure of whether a phoneme is pronounced correctly or incorrectly, but rather a measure of whether a phoneme is pronounced correctly, slightly incorrectly, or dramatically incorrectly. An example of a voice disorder that is not necessarily binary but is frequently assessed on a spectrum is a lisp in English. Similarly, a singing voice may be identified on a spectrum where a first relatively undesirable quality is "airiness" and a second relatively desirable quality is "richness," with many gradations between airy and rich voices.
[0026] In examples where the first quality is quantifiable on a non-binary importance scale (e.g., a scale of 1 to 10), controller 110 may compare user 120's speech and spatial data against sets of historical records in corpus 150 that are classified (within corpus 150) as being different values on that importance scale, such that each set of historical records that matches the speech or spatial data, or both, indicates an importance of the first quality of user 120's utterance (e.g., if a set of historical records that is 7 on the importance scale matches user 120's speech data, controller 110 determines that the user's speech data has a first quality of importance of 7).
[0027] The controller 110 determines alternative positions for the facial features of the user 120, and these alternative positions are determined to change the voice of the user 120 from a first quality to a second quality, where the second quality discussed herein includes a quality that is consistent with (or a step in the direction of) a technology standard or widely held preference. The controller 110 can determine the alternative positions by comparing spatial data and / or audio data of the user 120 speaking with a corpus 150, where the corpus 150 includes spatial data and / or audio data of the user 120 that is classified as the second quality.
[0028] In some examples, corpus 150 may include historical records of a single historical person speaking with both a first quality and a second quality (e.g., when the historical person learned how to speak with the second quality, potentially as a result of directed assistance from controller 110 as described herein). In such examples, if controller 110 determines that user 120's audio data and / or spatial data matches historical data of a historical person classified as speaking with a first quality, controller 110 may use the spatial data of that same historical person speaking with a second quality to determine alternative positions for facial features of user 120. For example, controller 110 may determine the relative changes in the historical person's facial features as the historical person adjusts from speaking with the first quality to speaking with the second quality, and apply scaled changes to match vector diagram 132 of user 120 to determine the alternative positions.
[0029] In some examples, controller 110 may analyze multiple historical relative changes in corpus 150 having recordings of utterances with a first quality and a second quality so that controller 110 can identify trends in, for example, how different facial structures require different types of changes to change from a first quality to a second quality. Once controller 110 analyzes corpus 150 to determine these trends for various spatial arrangements, controller 110 can apply them to the utterance data of each of users 120 to determine adjustments.
[0030] Additionally, controller 110 can apply trends determined by analyzing corpus 150 to explain the detected importance of a first quality of user 120's utterances. For example, if controller 110 detects that user 120 utters with a first quality, controller 110 can determine the importance of the first quality, if appropriate (e.g., some mispronunciations, etc., may be dualistic and other mispronunciations, etc., may be non-dualistic, as identified by controller 110). In such instances where controller 110 determines the importance, controller 110 can determine an adjustment corresponding to this importance. For example, an utterance having a first quality with a relatively low importance may require a relatively small adjustment to change the utterance to a second quality, while an utterance having a first quality with a relatively high importance may require a relatively large adjustment to change the utterance to the second quality.
[0031] Further, as described above, in some examples, controller 110 may create a multi-stage plan for changing user 120's vocalizations from initial incorrect and / or undesirable qualities to subsequent correct and / or desirable qualities. For example, controller 110 may determine that too many adjustments are needed for user 120 to reliably and accurately perform in one adjustment. In such examples, controller 110 may decompose the full set of adjustments into a series of adjustments that build on each other, providing a graphical representation of a first adjustment until user 120 masters the first adjustment, then providing a second adjustment until mastery is achieved, and so on. By evaluating how user 120 responds to the provided adjustments, controller 110 can learn over time how to decompose the full set of adjustments in a way that builds toward correct and / or desirable vocal qualities (e.g., if some adjustments in some order cause regression, controller 110 will be less likely to provide those adjustments in that order in subsequent sessions with a different user 120).
[0032] In some examples, controller 110 can identify one or more characteristics of user 120 that further define whether the utterance is of a first quality or a second quality, or generally classify the user's vector diagram 132, or both. For example, controller 110 can identify the language in which user 120 speaks, the general facial shape of user 120, the age of user 120, the accent of user 120, etc. In such examples, controller 110 can then compare user 120's utterance to respective populations 152 in corpus 150 that share these characteristics and identify that user 120 is uttering with the first quality, or provide adjustments to user 120 to instead utter with the second quality, or both, to improve results. Controller 110 can identify which characteristics tend to cause user 120 to improve faster over time when controller 110 considers these characteristics, and then organize populations 152 according to these characteristics. For example, controller 110 may determine that people over a certain age tend to have more wrinkles and change how controller 110 identifies features on the faces of these people, causing controller 110 to organize populations according to people of a certain age (or people with a threshold amount of wrinkles, or both).
[0033] In some examples, controller 110 can receive preferred characteristics of a group 152 that user 120 desires to resemble. For example, as noted above, an actor may wish to acquire a regional dialect for an upcoming role, and such actor would select that regional group 152 so that controller 110 can provide the necessary adjustments for the actor to perfect that dialect. Similarly, a regional reporter may wish to drop their regional dialect for a more neutral accent, and thus controller 110 may specify the characteristics of the neutral accent to use.
[0034] In some examples, the controller 110 can input the corpus 150. For example, the controller 110 can autonomously input the corpus 150 with common linked words, phrases, letter pronunciations, etc., so that the controller 110 can search the corpus 150 for predetermined words with predetermined qualities. In other examples, a person trained in natural language processing (NLP) or the like can structure the corpus 150 as described herein and potentially create one or more ensembles 152. In particular examples, a person trained in NLP or the like can construct an initial corpus 150 large enough for the controller 110 to make determinations and / or calculations as discussed herein based on the data in the corpus 150, and then the controller 110 can autonomously grow the corpus 150 and ensembles 152 (including creating entirely new ensembles 152) according to the existing structure and logic of the corpus 150 and / or ensembles 152. The controller 110 can thus grow the corpus 150 and / or ensembles 152 in a supervised or unsupervised manner.
[0035] The controller 110 may interact with the display 130, the sensor 140, or the corpus 150, or a combination thereof, via the network 160. The network 160 may include a computing network through which computing messages may be sent and / or received. For example, the network 160 may include the Internet, a local area network (LAN), a wide area network (WAN), a wireless network such as a wireless LAN (WLAN), and the like. The network 160 may be configured with copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, edge servers, or a combination thereof. A network adapter card or network interface of each computing / processing device (e.g., the controller 110, the display 130, the sensor 140, or the corpus 150, or a combination thereof) may receive messages and / or instructions from and / or through the network 160 and forward the messages and / or instructions to the respective memory or processor of the respective computing / processing device for storage, execution, etc. Although network 160 is depicted as a single entity in FIG. 1 for purposes of explanation, in other examples, network 160 may include multiple private and / or public networks.
[0036] 2A-2C depict example graphical representations of a user's 120's face showing adjustments from initial positions of elements that trigger vocalizations with a first quality to a set of alternative positions determined to trigger vocalizations with a second quality. For example, FIG. 2A depicts an example adjusted vector diagram 170A having adjusted nodes 172 connected by adjusted vectors, where the adjusted nodes 172 indicate different positions for each element of the user's 120's face. As shown, adjusted vector diagram 170A has an adjustment where the user 120 slightly expands their mouth to form more of an "O" shape. Similarly, FIG. 2B illustrates adjusted vector diagram 170B has an adjustment where the user 120 raises their cheekbones. In some examples, the controller 110 can provide audio or text cues along with the graphical representation, such as saying aloud, "Make an O shape with your mouth," or providing text on the display 130 stating, "Lift your cheekbones," or both.
[0037] In some examples, controller 110 may further identify when any movements of user 120 detected by sensor 140 that attempt to replicate adjusted vector diagrams 170A, 170B (collectively, adjusted vector diagrams 170) do not match adjusted vector diagram 170. For example, controller 110 may notify user 120 if the user “overshoots” the proposed adjustment. Controller 110 may provide this feedback in verbal form (e.g., by stating audibly, “Make your lips just a little smaller”), textual form (e.g., by providing text on display 130 stating, “Lower your cheekbones a little”), or graphical form with an updated adjusted vector diagram 170 showing the adjustment back in the “other” direction.
[0038] As shown, the controller 110 may display an adjusted vector diagram 170 on the display in addition to the initial vector diagram 132. As shown, the adjusted vector diagram 170 may be displayed differently than the initial vector diagram 132. For example, the adjusted vector diagram 170 may be a different color, bold, highlighted, oversized, etc. as depicted on the display 130.
[0039] As shown, the adjusted vector diagram 170 may be provided graphically on the display 130. While the display 130 shows only the lower half of the user 120's face in both FIGS. 2A and 2B for illustrative purposes, it will be understood that the controller 110 may cause the display to show more or less of the user 120's face while providing the graphical representation of the adjustment. As shown, the display may further depict a graphical element 174 that explicitly defines the adjustment. For example, the graphical element 174 in FIG. 2A depicts the outer edge of the mouth closure. In some examples, the controller 110 may provide only the adjusted vector diagram 170 and not the graphical element 174. In some examples, the controller 110 may graphically indicate the adjustment from the initial vector diagram 132 to the adjusted vector diagram 170 (e.g., such that a node 172 of the adjusted vector diagram 170 is graphically depicted as movable from a position according to the adjusted vector diagram 170 to a position of the initial vector diagram 132 that moves along the adjustment 174).
[0040] In some examples, controller 110 may provide a graphical representation of the inside of the user's mouth, as depicted in FIG. 2C. Controller 110 may provide this graphical representation on display 130A, which may be substantially similar to display 130. In some examples, controller 110 may simultaneously provide a graphical representation of the inside of the mouth, as depicted in FIG. 2C, and a graphical representation of user 120's face, as depicted in FIG. 2A or 2B, or both. In other examples, controller 110 may allow user 120 to switch between an interior view similar to FIG. 2C and a front view similar to FIG. 2A or 2B, or both.
[0041] As shown, controller 110 can provide this graphical representation in adjustment 180. For example, as shown in FIG. 2C , controller 110 can detect that tongue 182 in oral cavity 184 is not touching the top set of incisors 186 of user 120. Controller 110 may further detect that user 120 is attempting to say the letter "L," which cannot be pronounced according to dictionary pronunciation, when a tip element of tongue 182 defines the positional state depicted in FIG. 2C , and cause user 120 to instead speak with a first quality (e.g., "L" sounds more like a "Y"). In response to this determination, controller 110 determines adjustment 180 in which a top element of tongue 182 contacts top incisors 186. As shown, adjustment 180 of tongue 182 is depicted with a dashed line. In other examples, the controller 110 may instead show the adjustment using a vector diagram similar to Figures 2A and 2B (in other examples, the controller 110 may depict the adjustment to the front surface using dashed lines or the like rather than the adjusted vector diagram 170).
[0042] As mentioned above, the controller 110 may include, or be part of, a computing device that includes a processor configured to execute instructions stored on a memory to perform the techniques described herein. For example, FIG. 3 is a conceptual box diagram of such a computing system 200 of the controller 110. While the controller 110 is depicted as a single entity (e.g., in a single housing) for illustrative purposes, in other examples, the controller 110 may include two or more discrete physical systems (e.g., in two or more discrete housings). The controller 110 may include an interface 210, a processor 220, and a memory 230. The controller 110 may include any number or quantity of interfaces 210, processors 220, or memories 230, or combinations thereof.
[0043] The controller 110 may include components that enable it to communicate with devices external to the controller 110 (e.g., to transmit data and to receive and utilize transmitted data). For example, the controller 110 may include an interface 210 configured to enable the controller 110 and components within the controller 110 (e.g., processor 220, etc.) to communicate with entities external to the controller 110. Specifically, the interface 210 may be configured to enable components of the controller 110 to communicate with the display 130, the sensors 140, the corpus 150, etc. The interface 210 may include one or more network interface cards, such as an Ethernet card or any other type of interface device capable of transmitting and receiving information, or a combination thereof. Any suitable number of interfaces may be used to perform the described functions according to particular needs.
[0044] As discussed herein, controller 110 may be configured to identify incorrect or otherwise objectionable vocalizations of user 120 and graphically provide adjustments to user 120's face to modify the vocalizations. Controller 110 may utilize processor 220 to provide facial adjustments to user 120 in this manner. Processor 220 may include, for example, a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or equivalent discrete or integrated logic circuitry or combinations thereof. Two or more processors 220 may be configured to cooperate to suggest facial adjustments accordingly.
[0045] The processor 220 may suggest facial adjustments to the user 120 according to instructions 232 stored in the memory 230 of the controller 110. The memory 230 may include a computer-readable storage medium or a computer-readable storage device. In some examples, the memory 230 may include one or more of short-term memory or long-term memory. The memory 230 may include, for example, forms of random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), magnetic hard disk, optical disk, floppy disk, flash memory, electrically programmable memory (EPROM), electrically erasable programmable memory (EEPROM), and the like. In some examples, the processor 220 may suggest facial adjustments as described herein according to instructions 232 of one or more applications (e.g., software applications) stored in the memory 230 of the controller 110.
[0046] In addition to instructions 232, in some examples, collected or predetermined data or techniques, etc., used by processor 220 to suggest facial adjustments as described herein may be stored in memory 230. For example, memory 230 may include the above-mentioned information collected from user 120 during an utterance, such as spatial data 234 and audio data 236. As shown, spatial data 234 and audio data 236 may be stored in memory 230 in an associated manner, with each given set of spatial data 234 corresponding to a point in time at which audio data 236 was recorded. Additionally, in some examples, memory 230 includes some or all of corpus data 238, including respective population data 240 for some or all of corpus 150. In some examples, controller 110 may capture local copies of records from corpus 150 that match characteristics of user 120, and these local copies are stored in memory 230 until analysis of user 120 is complete.
[0047] Additionally, the memory 230 may include threshold and preference data 242. The threshold and preference data 242 may include thresholds that define how the controller 110 provides facial adjustments to the user 120. For example, the threshold and preference data 242 may provide a preferred means for engagement that may detail how the user 120 prefers adjustments to be displayed or suggested. The threshold and preference data 242 may also include a threshold deviation from a baseline required to cause the controller 110 to suggest facial adjustments. For example, the user may specify in the threshold and preference data 242 that the controller 110 provide facial suggestions only for certain types of qualities or for qualities of a certain detected importance.
[0048] Memory 230 may further include natural language processing (NLP) techniques 244. Controller 110 can use NLP techniques to determine what user 120 is saying so that controller 110 can determine whether user 120 is saying it correctly and / or preferably. NLP techniques 244 can include, but are not limited to, semantic similarity, syntactic analysis, and ontology matching. For example, in some embodiments, processor 220 can be configured to aggregate natural language data collected during utterances to determine the words user 120 is saying (e.g., to compare them with similar words in corpus 150) and to determine semantic features (e.g., word meanings, repeated words, keywords, etc.) and / or syntactic features (e.g., word structure, location of semantic features in headlines, titles, etc.) of this natural language data to determine the words user 120 is saying (e.g., to compare them with similar words in corpus 150).
[0049] The memory 230 may further include machine learning techniques 246 that the controller 110 can use to improve over time the process of suggesting facial adjustments to the user as discussed herein. The machine learning techniques 246 may include algorithms or models generated by performing supervised, unsupervised, or semi-supervised training on a dataset and then applying the generated algorithms or models to suggest facial adjustments to the user 120. Using these machine learning techniques 246, the controller 110 may improve its ability to suggest facial adjustments to the user 120 over time. For example, the controller 110 may identify over time certain characteristics of the population 152 that provide more relevant historical insight, what types of adjustments cause the user 120 to change from speaking with a first quality to a second quality more quickly and / or more repeatedly, improve the speed at which the quality and importance are identified and calculated, etc.
[0050] Specifically, the controller 110 can learn how to provide facial adjustments under supervised machine learning techniques 246 from one or more human operators who edit the “vectors” that include the “vector points” when the controller 110 initially provides the vector diagrams 132. For example, this may include the human operator canceling and / or rearranging nodes 134 provided by the controller 110, as well as canceling and / or rearranging nodes 172 of the adjusted vector diagrams 170. This may include teaching the controller 110 what to do in response to detecting that the user 120 has moved their mouth / face / tongue in an attempt to replicate one of the adjusted vector diagrams 170. This may include having the controller 110 provide positive feedback (e.g., congratulating the user 120 verbally and / or graphically for positively matching the graphical representation of the facial adjustment), providing negative feedback (e.g., correcting the user 120 verbally and / or graphically and explaining one or more specific ways in which the user 120 did not map to the graphical representation of the facial adjustment), detecting that the user 120 has mastered the current facial adjustment and instead providing the next step in the sequence of facial adjustments, or repeating training the user 120 to reproduce the provided facial adjustment, or combinations thereof.
[0051] Additionally, one or more trained operators can train controller 110 using machine learning techniques 246 described herein to track the face / mouth / tongue of user 120. For example, one or more trained operators can change over time how nodes 134 and / or vectors 136 are positioned on the face of user 120. This can include providing feedback when nodes 134 and / or vectors 136 dissociate from facial recognition locations, in response to which controller 110 can remap the facial recognition of user 120.
[0052] The machine learning techniques 246 may include, but are not limited to, decision tree learning, association rule learning, artificial neural networks, deep learning, inductive logic programming, support vector machines, clustering, Bayesian networks, reinforcement learning, representation learning, similarity / metric learning, sparse dictionary learning, genetic algorithms, rule-based learning, or other machine learning techniques or combinations thereof. Specifically, the machine learning techniques 246 may utilize one or more of the following example techniques:K-Nearest Neighbors (KNN), Learning Vector Quantization (LVQ), Self-Organizing Maps (SOM), Logistic Regression, Ordinary Least Squares Regression (OLSR), Linear Regression, Stepwise Regression, Multivariate Adaptive Regression Splines (MARS), Ridge Regression, Least Absolute Shrinkage and Selection Operator (LASSO), Elastic Net, Least Angle Regression (LARS), Probabilistic Classifier, Naive Bayes Classifier, Binary Classifier, Linear Classifier, Hierarchical Classifier, Canonical Correlation Analysis (CCA), Factor Analysis, Independent Component Analysis (ICA), Linear Discriminant Analysis (LDA), Multidimensional Scaling (MDS), Non-Negative Metric Decomposition (NMF), Partial Least Squares Regression (PLSR), Principal Component Analysis (PCA), Principal Component Regression (PCR), Sammon Mapping, t-Distributed Stochastic Neighbor Embedding (t-SNE), Bootstrap Aggregation, Ensemble Mean, Gradient Boosting Decision Trees (ABRT), Gradient Boosting Machines (GBM), Inductive Bias Algorithm, Q-Learning, State-Action-Rew ard-state-action (SARSA), Temporal Difference (TD) Learning, Apriori Algorithm, Equivalence Class Transformation (ECLAT) Algorithm, Gaussian Process Regression, Gene Expression Programming, Group Method of Data Processing (GMDH), Inductive Logic Programming, Instance-Based Learning, Logical Model Trees, Information Fuzzy Networks (IFN), Hidden Markov Models, Gaussian Naive Bayes, Multinomial Naive Bayes, Averaged Overdependent Estimator (AODE), Classification and Regression Trees (CART), Chi-Square Automatic Interaction Detection (CHAID), Expectation Maximization Algorithm, Feedforward Neural Networks, Logic Learning Machines, Self-Organizing Maps, Single Linkage Clustering, Fuzzy Clustering, Hierarchical Clustering, Boltzmann Machines, Convolutional Neural Networks, Recurrent Neural Networks, Hierarchical Temporal Memory (HTM), or other machine learning algorithms or combinations thereof.
[0053] Using these components, controller 110 may provide a graphical representation of suggested facial adjustments to user 120 in response to the quality of the detected utterance, as discussed herein. For example, controller 110 may suggest facial adjustments according to flowchart 300 depicted in FIG. 4. While flowchart 300 of FIG. 4 is discussed in connection with FIG. 1 for illustrative purposes, it will be understood that other systems may be used to implement flowchart 300 of FIG. 4 in other embodiments. Furthermore, in some examples, controller 110 may perform a different method than flowchart 300 of FIG. 4, or controller 110 may perform a similar method, such as with more or fewer steps in a different order.
[0054] The controller 110 receives (302) audio data of the user 120 vocalizing. This can include the user singing or speaking in one or more languages. The controller 110 additionally receives (304) spatial data of the user 120 vocalizing. The controller 110 can receive both the vocal audio data and the spatial data from a single sensor 140 (e.g., a single device that records both audio and video). In another example, the controller 110 receives the vocal audio data from a first (set of) sensors 140 and the spatial data from a second (set of) sensors 140. If the controller 110 receives the vocal audio data separately from some or all of the spatial data, the controller 110 may synchronize all of the audio and spatial data.
[0055] The controller 110 determines (306) the location of facial elements of the user 120 during the utterance. As discussed herein, the relative locations of these elements cause various qualities in the voice of the user 120. In some examples, determining the location of the elements includes generating a vector diagram 132 of the face of the user 120, as described herein.
[0056] Controller 110 identifies a subset of locations of one or more elements of the user's face that cause the utterance to be of a first quality (308). For example, controller 110 can compare the audio data or the spatial data, or a combination thereof, with data in corpus 150 to identify the utterance to be of a first quality, and in response, controller 110 determines which particular element locations cause the first quality.
[0057] The controller 110 determines (310) alternative positions for these elements that would alter the utterance to have a second quality. The controller 110 may determine these alternative positions by identifying alternative positions in the historical records of the corpus 150 that would cause similar people to speak with the second quality rather than the first quality. Once the controller 110 determines these alternative positions, the controller 110 provides (312) an adjusted graphical representation of the user 120's face. The adjustment may detail how the user 120 moves their face, tongue, or both from an initial position (where the user 120 spoke with the first quality) to an alternative position (where the user 120 can speak with the second quality). For example, the controller 110 may provide an adjusted vector diagram 170 on the display 130 as described herein.
[0058] The description of various embodiments of the present invention is presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. It will be apparent to those skilled in the art that many modifications and variations are possible without departing from the scope and spirit of the described embodiments. The terms used herein have been selected to best explain the principles of the embodiments, practical applications or technical improvements to technology found in the market, or to enable those skilled in the art to understand the embodiments described herein.
[0059] The present invention may be a system, method, or computer program product, or combination thereof, integrated at any possible level of technical detail. The computer program product may include a computer-readable storage medium (or media) having stored thereon computer-readable program instructions for causing a processor to carry out aspects of the present invention.
[0060] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction execution device. By way of example, and not limitation, a computer-readable storage medium may be an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices having instructions recorded on punch cards or ridge structures in grooves, or the like, and suitable combinations thereof. Computer-readable storage, as used herein, should not be construed as a transitory signal per se, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or an electrical signal transmitted over a wire.
[0061] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof). The network may be comprised of copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface of each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage on a computer-readable storage medium within the respective computing / processing device.
[0062] The computer-readable program instructions for carrying out the operations of the present invention may be either source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and procedural programming languages such as the "C" programming language and similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, as a standalone software package, or partially on the user's computer. Alternatively, the computer may be executed partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the computer-readable program instructions in order to carry out aspects of the present invention.
[0063] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0064] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions can also be stored in a computer-readable storage medium connectable to a computer, programmable data processing apparatus, or other device, or combination thereof, that functions in a particular way, such that the computer-readable program instructions stored therein create one of an article of manufacture containing instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0065] Computer-readable program instructions, such as instructions to perform the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams on a computer, other programmable apparatus, or other device, can also be loaded into a computer, other programmable data processing apparatus, or other device to perform a series of operational steps on the computer, other programmable apparatus, or other device to produce a computer-implemented process.
[0066] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of executable implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, which constitute one or more executable instructions for implementing the specified logical function(s). In some alternative embodiments, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a special-purpose hardware-based system that performs the specified functions or operations or executes a combination of special-purpose hardware and computer instructions.
Claims
1. A computer receiving voice data of a user's utterance when the user utters during a certain period of time; a computer receiving spatial data of the user's face during the period; and a computer using the spatial data to identify positions of the facial elements relative to other facial elements during the time period, the relative positions of the elements causing a plurality of qualities in the user's voice; identifying, by the computer, a subset of the positions of one or more of the elements that cause a first incorrect or objectionable quality detected among the plurality of qualities during the time period; determining an alternative position for the one or more elements that is determined to cause the user's voice to have a correct or preferred second quality of the plurality of qualities rather than the first quality; and a computer providing to the user a graphical representation of the face indicating one or more adjustments from the subset of positions to the alternative positions.
2. The computer determines the location of the element by determining an initial vector diagram of the face of the user; determining the alternative position includes determining an adjusted vector diagram of the face of the user; The method of claim 1 , wherein providing the graphical representation of the face includes providing both the initial vector diagram and the adjusted vector diagram.
3. The method of claim 1 , wherein the graphical representation is provided in real time.
4. The method of claim 3, further comprising the computer using an augmented reality device to provide the graphical representation in real time over a current image of the user's face.
5. The method of claim 1, further comprising: a computer collecting data regarding the position of the user's tongue in the user's mouth; the positions of the facial elements include elements on the user's tongue; and the one or more adjustments include an alternative position of at least one of the elements on the tongue.
6. The method of claim 5 , wherein one of an ultrasound sensor, a mouth guard, a retainer, or a tongue sleeve collects the data regarding the position of the tongue.
7. A computer determining the age and language of the user; and a computer identifying a machine learning model specifically trained for the age and the language, the machine learning model comprising: receiving the audio data; receiving the spatial data; locating the element; and Identifying how the subset of locations give rise to the first quality; and The method of claim 1 , further comprising: determining and identifying the alternative locations of the one or more elements.
8. The method of claim 1, further comprising the computer identifying the user's facial shape by analyzing spatial data, and determining the alternative position includes comparing the face against a corpus of faces having similar facial shapes.
9. The method of claim 1, further comprising the computer determining an importance of the first quality, and the adjustment taking into account the determined importance.
10. a processor; a memory in communication with the processor, the memory, when executed by the processor, causing the processor to: receiving voice data corresponding to an utterance of a user during a certain period of time; receiving spatial data of the user's face during said period; using the spatial data to identify positions of the facial elements relative to other facial elements during the time period, the relative positions of the elements causing multiple qualities of the user's voice; identifying a subset of the locations of one or more of the elements causing a first incorrect or objectionable quality detected in the plurality of qualities during the time period; determining an alternative position for the one or more elements that is determined to cause the user's voice to have a correct or preferred second quality of the plurality of qualities rather than the first quality; and providing the user with a graphical representation of the face indicating one or more adjustments from the subset of positions to the alternative positions.
11. Locating the element includes determining an initial vector diagram of the face of the user; determining the alternative position includes determining an adjusted vector diagram of the face of the user; The system of claim 10 , wherein providing the graphical representation of the face includes providing both the initial vector diagram and the adjusted vector diagram.
12. The system of claim 10 , wherein the graphical representation is provided in real time.
13. 13. The system of claim 12, wherein the memory further comprises instructions that, when executed by the processor, cause the processor to use an augmented reality device to provide the graphical representation in real time over a current image of the face of the user.
14. 11. The system of claim 10, wherein the memory further comprises instructions that, when executed by the processor, cause the processor to collect data regarding a position of the user's tongue in the user's mouth, the positions of the facial elements include an element on the user's tongue, and the one or more adjustments include an alternative position of at least one of the elements on the tongue.
15. 15. The system of claim 14, wherein one of an ultrasound sensor, a mouth guard, a retainer, or a tongue sleeve collects the data regarding the position of the tongue.
16. A computer program that causes a computer to execute a method according to any one of claims 1 to 9.
17. A computer-readable storage medium having recorded thereon a computer program that causes a computer to execute the method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Pronunciation diagnostic device, pronunciation diagnostic method, recording medium, and pronunciation diagnostic program
JP2007122004A
Language pronunciation practice support system
JP2008158055A
Facial animation created from acting
JP2009534774A
Method of learning pronunciation of english vowel
JP2010072606A
Method and apparatus for intraoral palpation feedback
JP2011510349A