Method for collecting voice data for each vocalization method using smartphone application
The method addresses the limitations of existing voice data collection by guiding users to manipulate vocal organs and vocalization methods, reducing trial and error, and providing data for AI models by ensuring voice data quality and analyzing parameter changes.
Patent Information
- Application Number
- PCT/KR2024/017153
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-30
- Filing Date
- 2024-11-04
- Publication Date
- 2025-08-21
AI Technical Summary
Existing voice data collection methods, such as those described in Korean Patent Publication No. 10-2023-0090968, do not effectively handle changes in voice due to varying vocalization methods like head voice or chest voice, making it difficult to apply to voice data collection for artificial intelligence models.
A method using a smartphone application that guides users to manipulate their vocal organs and vocalization methods to produce specific pitches or vowels, collecting voice data by ensuring the voice is within preset frequency and decibel ranges, and analyzing changes in voice parameters over multiple uses.
Reduces trial and error in vocalization, protects the voice, and provides data for developing artificial intelligence models by analyzing stability and changes in voice parameters, predicting potential developments based on continuous use.
Smart Images

Figure KR2024017153_21082025_PF_FP_ABST
Abstract
Description
A method for collecting voice data by pronunciation method using a smartphone application
[0001] The present invention relates to a method for collecting voice data by vocalization method, and more particularly, to a method for collecting voice data by vocalization method using a smartphone application, which changes vocalization by manipulating a specific vocal organ or vocalization method according to a vocalization method guide when vocalizing a specific vowel at a specific pitch or in succession of multiple pitches using a smartphone application, and collects voice data according to the result.
[0002] Today, artificial intelligence technology is growing inevitably based on the advancement of electrical / electronic and communication technologies and big data. Recent voice recognition technology related to such artificial intelligence technology targets voice data of speech sounds for the purpose of communication.
[0003] Voice is produced through the delicate interaction of the larynx, breathing, and resonating organs. Depending on the degree of interaction within the larynx, which generates voice, various parameters that characterize the voice on the voice spectrogram change.
[0004] The sound that passes through the vocal cords is composed of complex sounds, and it forms a frequency spectrum with the fundamental sound, which is the sound corresponding to the fundamental vibration with the lowest frequency, and the overtones which are integer multiples of the fundamental frequency. In addition, as the sound passes through the vocal cords and passes through the vocal tract (the path that the breath from the lungs passes through the mouth or nose) and is emitted as sound, the energy of a specific frequency band in the frequency spectrum is amplified or canceled out depending on the shape or form of the vocal tract. Accordingly, the maximum value of each amplified frequency range is referred to as the first resonant frequency (the first formant, F1), the second resonant frequency (the second formant, F2), etc. from the lowest frequency. Vowels can be identified by the first resonant frequency and the second resonant frequency. The first resonant frequency is inversely proportional to the degree of tongue elevation, and increases as the tongue is lowered and the distance between the tongue and the palate increases. The second resonant frequency is proportional to the degree of tongue forward and backward advancement, and increases as the tongue advances forward.
[0005] Meanwhile, Korean Patent Publication No. 10-2023-0090968 discloses a "method for collecting cough, respiratory, and vocal sound data using a smartphone," and the method for collecting cough, respiratory, and vocal sound data using a smartphone comprises: a) a step of inducing a mouth position by using a camera of a smartphone to produce a mouth shape for uttering an arbitrary letter including a vowel that can be produced by opening the mouth; b) a step of automatically starting recording by activating a recording function of the smartphone when the user's mouth shape is positioned at the induced mouth shape displayed on the screen of the smartphone; c) a step of automatically starting recording, outputting a phrase and / or voice prompting the user to cough or breathe once, or a step of prompting the user to follow along by playing an example vocal sound; d) a step of collecting acoustic data of the cough or respiratory sound by recording the cough or respiratory sound when the user coughs or breathes once, or collecting acoustic data of the vocal sound according to the user's example vocal sound; e) a step of automatically stopping the recording function of the smartphone to automatically end the collection of cough, respiratory, and vocal sound data when N cough or respiratory sound data are collected by repeating steps c) and d) N times or vocal sound data satisfying preset specifications are collected; f) a step of transmitting the cough, respiratory, and vocal sound data collected by the smartphone to a server, and analyzing the cough, respiratory, and vocal sound data by the server; and g) a step of the server transmitting the analysis result to the smartphone and displaying the health signals of the user's cough, respiratory, and vocal sound on the screen of the smartphone.
[0006] As described above, in the case of the above patent document, by always keeping the distance and angle between the smartphone and the point of speech constant, the consistency of the collected data is maintained, so that high-quality sound data can be collected, and by displaying health signals such as the user's cough, breathing, and vocal sounds, there is an advantage in that the user can easily check the health status of his or her respiratory system. However, it does not deal with a method for vocal training by causing a change in voice (vocalization) according to the vocal method such as head voice or chest voice, and therefore contains a problem that makes it difficult to apply it to collecting voice data related to that.
[0007] The present invention was created by comprehensively considering the above matters, and the purpose of the present invention is to provide a method for collecting voice data by voice method using a smartphone application for voice guide, which can help reduce trial and error in the user's voice and protect the voice by having the user utter a specific pitch or utter several pitches consecutively as a designated vowel and by manipulating the vocal organ or the voice method while uttering to change the voice, thereby collecting voice data.
[0008] In addition, another object of the present invention is to provide a method for collecting voice data by voice method using a smartphone application, which can provide data on the stability and degree of change of a user's voice and predict the possibility of development by analyzing the direction of change of parameters, changed values, amount of change, standard deviation, etc. according to the number of repeated uses as a user continuously uses a smartphone application for voice guidance, and the like.
[0009] In addition, another object of the present invention is to provide a method for collecting voice data for various vocalization methods for developing an artificial intelligence model.
[0010] In order to achieve the above purpose, a method for collecting voice data by pronunciation method using a smartphone application according to the present invention is provided.
[0011] Each step is performed by executing a smartphone application for voice guidance installed on the user's smartphone.
[0012] a) A step in which a smartphone application for voice guidance (hereinafter referred to as the “smartphone application”) receives voice data of a user speaking for several seconds through a microphone according to the voice method guide or the voice organ operated according to the voice method guide and recognizes the voice;
[0013] b) a step in which the smartphone application determines whether the voice is a voice within a preset frequency range by pitch;
[0014] c) If the voice is a voice within a preset frequency range, the smartphone application determines whether the voice is a voice within the first and second resonance frequency ranges of the presented vowel;
[0015] d) If the voice is a voice within the first and second resonance frequency ranges, the smartphone application determines whether the voice is a voice within a preset decibel (dB) range; and
[0016] e) The characteristic is that if the above voice is within a preset decibel (dB) range, the smartphone application includes a step of labeling the voice data as a specific voice and storing it in memory.
[0017] Here, prior to the above step a), when collecting voice data according to the voice data collection procedure, a step of determining the user's basic information through a questionnaire in advance may be further included in order to analyze the correlation between the collected voice data and the user's basic information.
[0018] At this time, the user's basic information may include at least one of the user's gender, age, nationality, primarily used language, height, weight, presence / absence of voice lesson experience, and presence / absence of a vocal organ-related disease.
[0019] In addition, in the above step a), the user may be made to vocalize as he or she manipulates the vocal organ according to the vocalization method guide, and in recognizing the voice, the user may be made to listen to and vocalize a sound of a suggested pitch, or the user may be made to recognize the pitch of the voice he or she vocalizes as a frequency range set according to the pitch of the voice he or she vocalizes.
[0020] At this time, the frequency range set as the recognition range as the pitch presented above may be set to exceed the frequency of the pitch that is one semitone lower and below the frequency of the pitch that is one semitone higher, or may be set to exceed the frequency of the pitch that is one semitone lower and below the frequency of the pitch that is one semitone higher, with the frequency of the pitch being centered on the frequency of the pitch.
[0021] Additionally, in the determination of step b), if the voice is not within a preset frequency range by pitch, the smartphone application may further include a step of requesting the user to speak again with the correct pitch and guided vocalization method.
[0022] In addition, in the determination of step c), if the voice is not a voice within the first and second resonance frequency ranges of the presented vowel, the smartphone application may further include a step of requesting the user to speak again using the presented vowel and guided speaking method.
[0023] At this time, if the user response to the above request is 'pronounced with the suggested vowel', the smartphone application may store the voice data by reflecting the user response in the labeling of the voice data spoken for several seconds, or adjust the first and second resonance frequency ranges of the vowel according to the pitch and vocalization stage.
[0024] In addition, in the determination of step d), if the voice is not within a preset decibel (dB) range, the smartphone application may further include a step of requesting the user to speak again at a recognizable volume and in a guided speaking method.
[0025] At this time, when the vocalization method instructed in the above-mentioned guided vocalization method must capture a change in decibel (dB), if the change in decibel (dB) is not clear in relation to the guided vocalization method, a step of requesting vocalization again using the guided vocalization method may be further included.
[0026] Additionally, the preset decibel (dB) range in step d) may be 40 to 90 dB.
[0027] In addition, in the step e), when the smartphone application labels the voice data as a specific voice, the voice data can be labeled as a voice produced by manipulating the vocal organ or manipulating the vocal method according to the pitch corresponding to the set frequency range, the suggested vowel, and the vocal method guide.
[0028] In addition, in step e), when the smartphone application labels the voice data as a specific voice, the user's condition on the day can be checked before and after using each step of the voice method guide and reflected in the labeling of the voice data.
[0029] At this time, the user's state reactions, including discomfort during pronunciation, can be confirmed and reflected in the labeling of voice data.
[0030] In addition, in step e), when the smartphone application labels the voice data as a specific voice, the first time the user uses each step of the pronunciation method guide to label and save the voice data is considered as the first time, and when the user uses it repeatedly in the future, the number of times the repetition is performed can be reflected in the labeling.
[0031] Additionally, some steps of the vocalization method guide in steps a) to e) may collect vocal data according to the results of manipulating a specific vocal organ using a leading voice or a leading vowel.
[0032] At this time, voice data according to the result of manipulating a specific vocal organ using the leading voice or leading vowel may be collected as a single voice data, or changes in the first and second resonance frequencies of the vowels of the voice data according to the result of manipulating a specific vocal organ with respect to the first and second resonance frequencies of the leading voice or leading vowel may be captured, and voice data corresponding to the leading voice or leading vowel pronunciation part and voice data according to the result of manipulating a specific vocal organ corresponding thereto may be collected separately.
[0033] In addition, when collecting voice data by manipulating the vocal organ or manipulating the vocal method according to the vocalization method guide in steps a) to e), if a specific pitch is presented, different pitches can be presented for women, men, and male / female children, respectively.
[0034] According to the present invention, in order to observe changes in voice (vocalization) according to a vocalization method, a smartphone application for vocalization guide is used to vocalize a specific pitch or multiple pitches in succession with a designated vowel, and vocalization is changed by manipulating the vocal organs or the vocalization method during vocalization, thereby collecting voice data, thereby reducing trial and error in the user's vocalization and helping to protect the voice.
[0035] In addition, as the user continuously uses the smartphone application for voice guidance, the direction of change in parameters, the changed values, the amount of change, the standard deviation, etc. are analyzed according to the number of repeated uses, thereby providing data on the stability and degree of change in the user's voice and predicting the possibility of development.
[0036] FIG. 1 is a schematic diagram showing the configuration of a system constructed to implement a method for collecting voice data by pronunciation method using a smartphone application according to the present invention.
[0037] Figure 2 is a flowchart showing the execution process of a voice data collection method for each pronunciation method using a smartphone application according to the present invention.
[0038] Figure 3 is a diagram showing the frequency range set as the recognition range for each pitch suggested for the user to vocalize as he or she manipulates the vocal organ according to the vocalization method guide.
[0039] Figure 4 is a diagram showing the first and second resonance frequency ranges of each vowel of voice data according to the results of manipulating the leading voice, the leading vowel, and a specific vocal organ.
[0040] Figure 5 is a table showing examples of different pitches for women, men, and male / female children, step by step, in the vocalization method guide.
[0041] FIG. 6 is a drawing showing a voice method guide displayed on a screen by a smartphone application employed in the voice data collection method of the present invention.
[0042] Figure 7 is a diagram showing the user's daily condition reflected in the voice data collection, the guide for going back to the sound, and the relative position of the pitch compared to the suggested pitch.
[0043] Figure 8 is a diagram showing how, when a user touches the 'It's weird / uncomfortable' button while speaking, the vocal range and user response are checked and reflected in voice data labeling.
[0044] Figure 9 is a drawing showing pulling the lower jaw as part of creating a good posture for vocalization as a basic posture for vocalization.
[0045] Figure 10 is a diagram showing the inhalation process sequentially in the lower abdomen pulling state of thoracic breathing.
[0046] Figure 11 is a drawing showing the state of the tongue and larynx according to the execution of laryngeal lowering.
[0047] Figure 12 is a drawing showing a state in which the back of the mouth and throat are maintained so that it feels like a ping-pong ball is caught in the back of the mouth while the larynx is lowered for head-voicing.
[0048] Figure 13 is a drawing showing the process of practicing changing only the shape of the mouth to pronounce the vowels u, i, oh, e, and a without making any sound while maintaining the state of Figure 12.
[0049] Figure 14 is a diagram showing the process of practicing the backward sound by lowering the larynx and maintaining the feeling of a ping-pong ball hanging at the back of the mouth.
[0050] Figure 15 is a diagram showing the process of practicing opening the inside of the throat to prevent the throat from narrowing in very high notes.
[0051] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings.
[0052] FIG. 1 is a schematic diagram showing the configuration of a voice data collection system constructed to implement a voice data collection method for each pronunciation method using a smartphone application according to the present invention.
[0053] Referring to FIG. 1, a voice data collection system (100) constructed to implement a voice data collection method for each pronunciation method using a smartphone application according to the present invention is configured to include a user terminal (110) and a service provider server (120).
[0054] The user terminal (110) accesses the service provider server (120), app store, play store, etc., and downloads and installs a smartphone application for voice guidance provided by the service provider server (120), app store, play store, etc., and then executes the application each time the user uses the device, thereby enabling the user to practice (train) voice. Such a user terminal (110) may include a smartphone (110a) (110b), a tablet PC (110c), etc.
[0055] The service provider server (120) provides a smartphone application for voice guidance, which is employed in the method for collecting voice data by voice method using a smartphone application according to the present invention, to the user terminal (110), the App Store, the Play Store, etc., and provides a service that analyzes the voice data collected through the user terminal (110) in depth to provide data on the stability and degree of change in the user's voice and predicts the potential for development. In addition, by combining and analyzing the user's basic information and voice data identified through a questionnaire, in the case of a user with a vocal organ-related disease, data is provided for improving the vocal organ-related disease and indicating the degree of improvement. Here, in providing the above data or services, the system operation program may be configured so that data or services corresponding to basic information are provided free of charge or for a fee, and data or services corresponding to in-depth (advanced) information and services providing additional user convenience are provided for a fee, according to regulations set in advance by the service provider. In Fig. 1, reference numeral 120d denotes a program for system operation, an application program related to voice data collection and analysis, and a database (DB) in which various data or information related to vocalization are stored.
[0056] Then, below, we will explain a method of collecting voice data by pronunciation method using a smartphone application based on a voice data collection system having the above configuration.
[0057] Figure 2 is a flowchart showing the execution process of a voice data collection method for each pronunciation method using a smartphone application according to an embodiment of the present invention.
[0058] Referring to FIG. 2, the method for collecting voice data by vocalization method using a smartphone application according to the present invention is performed in such a way that each step is executed by executing a smartphone application for vocalization guide installed on the user's smartphone (110a)(110b). First, the smartphone application for vocalization guide (hereinafter referred to as the "smartphone application") receives voice data produced by the user manipulating the vocal organs according to the vocalization method guide or vocalizing for several seconds (e.g., 2 to 5 seconds) through a microphone, and recognizes the corresponding voice (step S201). Here, examples of manipulating the vocal organs include instructing to lower the thyroid cartilage or instructing to open the mouth widely from side to side. And examples of manipulating the vocalization method include instructing to vocalize increasingly louder or increasingly quieter. Here, the user may also operate the vocal organ according to the vocalization method guide to vocalize and recognize the voice, and the user may listen to and vocalize a sound of the suggested pitch, or recognize the pitch within a frequency range set according to the voice spoken by the user. At this time, the frequency range set as the recognition range as the suggested pitch may be set to exceed the frequency of the pitch that is one semitone lower and below the frequency of the pitch that is one semitone higher around the frequency of the suggested pitch, or may be set to exceed the frequency of the pitch that is one semitone lower and below the frequency of the pitch that is one semitone higher around the frequency of the suggested pitch. Figure 3 is a table showing the frequency range set as the recognition range for each suggested pitch.
[0059] In this way, after the smartphone application receives the voice data spoken by the user and recognizes the voice, the smartphone application determines whether the voice is within a preset frequency range by pitch (step S202). In this determination, if the voice is not within a preset frequency range by pitch, the smartphone application requests the user to speak again with the correct pitch and guided vocalization method (step S203).
[0060] And in the determination of the step S202, if the voice is a voice within a preset frequency range, the smartphone application determines whether the voice is a voice within the first and second resonance frequency ranges of the presented vowel (step S204). In this determination, if the voice is not a voice within the first and second resonance frequency ranges of the presented vowel, the smartphone application requests the user to re-pronounce the presented vowel and the guided pronunciation method (step S205). At this time, if the user response to the request is 'pronounced with the presented vowel', the smartphone application may reflect the user response in the labeling of the voice data spoken for several seconds and store the voice data, or adjust the first and second resonance frequency ranges of the vowel according to the pitch and pronunciation stage.
[0061] Meanwhile, in the determination of step S204, if the voice is a voice within the first or second resonance frequency range, the smartphone application determines whether the voice is a voice within a preset decibel (dB) range (step S206). Here, the preset decibel (dB) range may be 40 to 90 dB.
[0062] In the determination of the above step S206, if the voice is not within the preset decibel (dB) range, the smartphone application requests the user to speak again at a recognizable volume and in a guided speaking method (step S207).
[0063] At this time, when the vocalization method instructed in the above-mentioned guided vocalization method must capture a change in decibel (dB), if the change in decibel (dB) is not clear in relation to the guided vocalization method, a step of requesting vocalization again using the guided vocalization method may be further included.
[0064] In addition, in the determination of the above step S206, if the voice is within a preset decibel (dB) range, the smartphone application labels the voice data as a specific voice and stores it in memory (step S208).
[0065] In a series of processes as described above, before step S201, when collecting voice data according to a voice data collection procedure, a step of determining the user's basic information through a questionnaire in advance may be further included in order to analyze the correlation between the collected voice data and the user's basic information.
[0066] At this time, the user's basic information may include at least one of the user's gender, age, nationality, primarily used language, height, weight, presence / absence of voice lesson experience, and presence / absence of vocal organ-related diseases (e.g., rhinitis, polyps, nodules, other vocal cord and larynx-related surgery experience, etc.).
[0067] In addition, in the above series of processes, the steps S202, S204, and S206 are not necessarily limited to being performed in the order shown in FIG. 2 (i.e., S202 → S204 → S206), and may be performed in various orders, such as S206 → S202 → S204 or S206 → S204 → S202, depending on the case.
[0068] In addition, in the step S208, when the smartphone application labels the voice data as a specific voice, the voice data can be labeled as a voice produced by operating the vocal organ or using a vocal method according to a pitch corresponding to a set frequency range, a suggested vowel, and a vocal method guide.
[0069] Additionally, in step S208, when the smartphone application labels the voice data as a specific voice, the user's daily condition can be checked before and after using each step of the voice method guide, and this can be reflected in the voice data labeling. Furthermore, the user's condition reactions, including discomfort during voice recording, can be checked and reflected in the voice data labeling.
[0070] In addition, in the step S208, when the smartphone application labels the voice data as a specific voice, the first time the user uses each step of the pronunciation method guide to label and save the voice data is considered as the first time, and when the user uses it repeatedly in the future, the number of times the repetition is performed can be reflected in the labeling.
[0071] In addition, in some steps of the vocalization method guide in steps S201 to S208, voice data may be collected according to the result of manipulating a specific vocal organ using a leading voice or a leading vowel. Here, the leading voice or leading vowel refers to a voice or vowel that helps to properly operate a specific vocal organ when collecting voice data according to the result of manipulating the specific vocal organ. The reason why it is called a voice here is that although "eu" is classified as a vowel in Korea, it is not always classified as a vowel in foreign countries, and in the vocalization method guide employed in the present invention, it is not a clear vowel "eu" pronunciation, but a sound made by biting the molars.
[0072] At this time, voice data according to the result of manipulating a specific vocal organ using the leading voice or leading vowel may be collected as a single voice data, or changes in the first and second resonance frequencies of the vowels of the voice data according to the result of manipulating a specific vocal organ with respect to the first and second resonance frequencies of the leading voice or leading vowel may be captured, and voice data corresponding to the leading voice or leading vowel pronunciation part and voice data according to the result of manipulating a specific vocal organ corresponding thereto may be collected separately.
[0073] Here, the first and second resonance frequency ranges of each vowel of the voice data according to the result of manipulating the leading voice, the leading vowel, and the specific vocal organ can be set as in the table (b) with reference to the graph as in (a) of Fig. 4. At this time, since the first and second resonance frequencies can change depending on the pitch and vocalization method, in response to a request by the second condition set for labeling the voice data, that is, if the voice is not a voice in the first and second resonance frequency ranges of the presented vowel, the smartphone application requests the user to re-pronounce the voice with the presented vowel and guided vocalization method. At this time, when the user responds to this request with something like 'pronounced the presented vowel', the voice data is stored by reflecting the user response in the labeling of the spoken voice data as described above, or the first and second resonance frequency ranges of the vowel can be adjusted depending on the pitch and vocalization stage.
[0074] In addition, when collecting voice data by manipulating the vocal organ or manipulating the vocal method according to the vocalization method guide in steps S201 to S208, if a specific pitch is presented, different pitches can be presented for women, men, and male / female children, respectively.
[0075] That is, when collecting voice data by manipulating the vocal organs or manipulating the vocal method according to the vocalization method guide, if a specific pitch is presented, it is presented based on the gender in the pre-stage questionnaire. At this time, the pitch presented in the vocalization method guide is based on female, and for males, a different pitch is presented. If the user is 10 years old or younger, a female pitch is presented. For male children over 10 years old, an additional questionnaire in the pre-stage explanation asks whether voice change has begun, and if voice change has begun, a male pitch is presented, and if voice change has not begun, a female pitch is presented. In addition, for male users in the 10-14 year old range who have not yet begun voice change, a questionnaire about whether voice change has begun before each use or intermittently, or the pitch can be presented by reflecting the results of a sentence reading test.
[0076] Figure 5 is a table showing examples of different pitches presented for women, men, and male / female children, step by step, in the vocalization method guide.
[0077] Below, a further explanation will be provided regarding the voice method guide employed in the voice data collection method for each voice method using a smartphone application according to the present invention.
[0078] FIG. 6 is a drawing showing a voice method guide displayed on a screen by a smartphone application employed in the voice data collection method of the present invention.
[0079] Referring to Fig. 6, when a user touches any vowel button (e.g., 'A' in column 1) as in (a), a pronunciation guide for the corresponding step is presented. If the user understands the explanation of the pronunciation guide for the corresponding step, the user presses the <Understand> button. Then, the user presses the <Pitch Button> as shown in (b) and produces a sound according to the pronunciation guide for several seconds until the <Recording Bar> is fully colored with the <Proposed Vowel> at the set pitch. That is, when the user touches the <Pitch Button> presented in Fig. 6 (b), a sound of the corresponding pitch is heard, and when the user starts to speak and the smartphone application starts to recognize the user's voice, the pitch sound is blocked.
[0080] The smartphone application recognizes the voice produced by the user for several seconds according to the voice method guide by manipulating the vocal organs, and if the voice is within the frequency range set for each pitch as described above, is within the first and second resonance frequency ranges of the presented vowel, and is within the range of 40 dB to 90 dB, the application labels the voice data as 'voice produced according to the vocal organs manipulating the vocal organs according to the voice method guide, with a pitch corresponding to the set frequency range, and with the presented vowel' and stores the label in the memory.
[0081] Figure 7 is a diagram showing the user's daily condition reflected in the voice data collection, the guide for going back to the sound, and the relative position of the pitch compared to the suggested pitch.
[0082] Referring to Figure 7, the smartphone application collects voice data by checking the user's daily condition before and after using the vocalization guide, as shown in (a), and reflecting this in data labeling. Furthermore, when the user is vocalizing, for example, in the case of "Saying Backward," as shown in (b), a reference image is visually displayed along with guidance text, such as, "At that part, try saying the sound as if it were going to the back of your throat."
[0083] When the user sees the pronunciation guide description for that step and touches it to indicate that he or she understands it (i.e., touches the 'I Understood' button), the smartphone application switches to the next voice data collection step.
[0084] When the user touches the <suggested pitch> button on the screen (c), the corresponding pitch sound is heard. Accordingly, when the user begins to vocalize, the pitch sound is blocked until the user's voice is recognized. In addition, when the user pronounces the <suggested vowel>, the pitch is visually presented in relative position compared to the <suggested pitch>, as shown in (c).
[0085] The user speaks according to the pronunciation method guide of the corresponding stage for several seconds until the <recording bar> is fully colored with the <suggested pitch> and <suggested vowel>. At this time, if the <recording bar> satisfies the conditions (i.e., the voice is within the frequency range set for each pitch, the voice is within the first and second resonance frequency ranges of the suggested vowel, and the voice is within the range of 40 dB to 90 dB), the voice is recognized and recorded, and the recording progress is notified. Then, if the <suggested pitch> satisfies the above conditions within the range of the currently suggested pitch (e.g., E4) and the voice data is recorded, the next pitch is suggested. At this time, the pitch may be suggested at semitone intervals, as shown in (c), or may be suggested using a major or minor scale, or a pentatonic scale.
[0086] Additionally, if the <proposed vowel> falls within the pitch range, the proposed vowel will be placed inside the shape. If the pitch of the user's voice is lower than the set pitch range, the proposed vowel will be below the shape (red circle), and if the pitch of the user's voice is higher than the set pitch range, the proposed vowel will be above the shape.
[0087] Figure 8 is a diagram showing how, when a user touches the 'It's weird / uncomfortable' button while speaking, the vocal range and user response are checked and reflected in voice data labeling.
[0088] Referring to FIG. 8, as described above, the user produces a sound according to the vocalization method guide of the corresponding step for several seconds until the <recording bar> is fully colored with the <proposed pitch> and <proposed vowel> as in (a), and the smartphone application recognizes this and, if the condition of the voice being a voice within the frequency range set for each pitch, a voice within the first and second resonance frequency ranges of the proposed vowel, and a voice within the range of 40 dB to 90 dB is satisfied, the corresponding voice data is labeled as 'a voice produced by operating the vocal organ or manipulating the vocalization method according to the vocalization method guide, the pitch corresponding to the set frequency range, the proposed vowel, and the vocalization method' and stored in the memory.
[0089] While the above series of processes are in progress, as in (b), in the case of <It's strange / inconvenient>, a guidance message such as "If your throat is tight or uncomfortable, or the sound is strange, please touch. I can give you feedback." may be displayed on the screen as a [User Guide]. Also, in the case of , a guidance message such as "If the pitch doesn't go up any more, please touch. It may be recorded in your vocal range." may be displayed.
[0090] When the user touches the "It's weird / uncomfortable" button, the smartphone application checks the user's vocal range and confirms user status reactions, such as discomfort during vocalization, which are reflected in the voice data labeling and provides feedback. That is, the application can request the user to repeat the vocalization at the correct pitch and using the guided vocalization method, or with the suggested vowel and using the guided vocalization method, or with a recognizable volume and using the guided vocalization method, or suggest that the user stop.
[0091] Additionally, if the user touches the <Stop> button, voice data collection for that step can be stopped.
[0092] Here, the labeling can be configured as, for example, "pitch_vowel_vocalization method guide step_round_implementation date_gender_age_language used_voice disease or type of voice disease_condition before use_condition after use_user response_user status response", and the order can be changed or content can be added according to criteria such as voice data classification folder settings in memory.
[0093] Figure 9 is a drawing showing pulling the lower jaw as part of creating a good posture for vocalization as a basic posture for vocalization.
[0094] First, the smartphone app guides you through a warm-up phase, guiding you into a posture conducive to thoracic and abdominal breathing. For example, it suggests, "Stand about halfway in front of a wall. Lean back against the wall with your head, back, and buttocks against the wall. Place both shoulders against the wall. Feel your ribcage expand and your stomach tighten? Pull your lower abdomen toward your spine, about two index finger widths below your navel. Hold this position. This is <lower belly pull>."
[0095] Referring to Figure 9, this shows pulling in the lower jaw as part of creating a good posture for vocalization as the basic posture for vocalization after the above warm-up stage. The smartphone application provides guidance such as, "If you are in a daze, unnecessary strength is lost in the lower jaw and a space is created between the upper and lower teeth. At this time, try pulling the lower jaw back slightly." In addition, it provides guidance such as, "<Pulling in the lower jaw>. Check your face with the lower jaw pulled in in the mirror and remember it with the sensation of your body."
[0096] Figure 10 is a diagram showing the inhalation process sequentially in the lower abdomen pulling state of thoracic breathing.
[0097] Referring to Figure 10, the smartphone application guides users by saying, "Inhale deeply through your nose while slowly counting to 1, 2, 3 in the <lower belly pull-in> state. At this time, your torso will expand from the lower abdomen to the back, under the armpits, and to the rib cage near the chest. Think of an oak barrel that is gradually expanding. Continue to <pull in the lower belly>. And at this time, do not raise your shoulders!" The user follows along and practices the inhalation breathing process.
[0098] Figure 11 is a drawing showing the state of the tongue and larynx according to the execution of laryngeal lowering.
[0099] Referring to Figure 11, this illustrates laryngeal depression, and the smartphone application provides guidance on laryngeal depression, such as, "1. Close your lips and touch the convex part in the middle of your throat with your thumb and index finger. Then, push the soft part at the back of the palate with the tip of your tongue. The convex part at the back of your throat will go down, right? Now, while maintaining the convex part at the back of your throat, lift your tongue from the soft part at the back of your palate. As the back of your tongue goes down, you will feel some force acting on the back of your tongue and the inside of your throat. Remember this feeling."
[0100] Along with the above guidance, as shown in Figure 11, guidance phrases such as “The tongue goes down”, “The uvula and throat are visible”, and “When the neck is touched, the convex part goes down” are displayed along with each picture.
[0101] Figure 12 is a drawing showing a state in which the back of the mouth and throat are maintained so that it feels like a ping-pong ball is caught in the back of the mouth while the larynx is lowered for head-voicing.
[0102] Referring to Figure 12, this is about head voice for producing high notes easily, and the smartphone application provides guidance such as "Try touching the convex part of the throat with your thumb and index finger. <Depress the larynx> and imagine that a ping-pong ball is caught in the back of your mouth.", "<Feeling like a ping-pong ball is caught in the back of your mouth> helps to further activate laryngeal depression and open other spaces in the mouth."
[0103] Figure 13 is a drawing showing the process of practicing changing only the shape of the mouth to pronounce the vowels u, i, oh, e, and a without making any sound while maintaining the state of Figure 12.
[0104] Referring to Figure 13, this is a process of practicing changing only the shape of the mouth for vowels. The smartphone application provides guidance such as, “Instead of making the sound in that state, try changing the shape of the mouth for the vowels “oo, i, oh, eh, ah” by opening only the shape of the mouth wide. Be careful not to raise the convex part in the throat.”, “Practice maintaining <lowering the larynx> and <the feeling of a ping-pong ball caught in the back of the mouth> for all vowels.”, “When making a sound, imagine that the ping-pong ball gets bigger and gets pushed back like that as you go to very high notes.” Then, the user practices by following these guidance.
[0105] Figure 14 is a diagram showing the process of practicing the sound going backwards by lowering the larynx and maintaining the feeling of a ping-pong ball hanging at the back of the mouth.
[0106] Referring to Figure 14, this is a process of practicing the backward sound by maintaining the feeling of lowering the larynx and having a ping-pong ball caught in the back of the mouth. The smartphone application provides guidance and questions such as, "Try the following in one breath. Maintaining <lowering the larynx> and <the feeling of a ping-pong ball caught in the back of the mouth>, and <going backward>, start with a faint, small sound of / u / , then gradually change only the shape of your mouth to / i / . At this time, try changing it 'by opening your mouth as wide as possible to the left and right'. Then, the sound will change to the suggested vowel! Try / u-e / and / u-a / too.", "Do you feel the vocal cords lightly or thinly sticking together in the back of your throat when you make the sound?" In addition, while displaying a picture, it provides guidance such as, "Please pause for a moment when you pronounce 'i, eh, ah' normally", and "Try changing it by opening your mouth as wide as possible to the left and right."
[0107] Figure 15 is a diagram showing the process of practicing opening the inside of the throat to prevent the throat from narrowing in very high notes.
[0108] Referring to Fig. 15, this is about <opening the inside of the throat> as part of <opening the head and throat> to easily produce high notes. The smartphone application provides guidance such as, "<lower the larynx> and <maintain the feeling of a ping-pong ball caught in the back of the mouth> while 'breathing in and saying / hu / , / hi / , / ho / , / he / , / ha / '. You will feel a refreshing open feeling in the inside of your throat. It is important to maintain this feeling and the part where this feeling is felt, that is, the inside of your throat being open.", "When uttering high notes or singing, more force is applied to the back of the tongue and the inside of the throat to maintain this feeling.", and displays pictures such as Fig. 15 (a) and (b) and provides guidance such as, "Bite down on your molars, don't use your nose, and try to inhale as much air as possible between your teeth.", "Imagine a ping-pong ball caught in your throat being pushed back and getting bigger.", "Then you will feel a force acting from the inside of your throat to the outside."
[0109] As described above, the method for collecting voice data by vocalization method using a smartphone application according to the present invention uses a smartphone application for vocalization guide to observe changes in voice (vocalization) according to vocalization method, and by having the user vocalize at a specific pitch or in succession of multiple pitches with a designated vowel and manipulate the vocal organ or vocalization method while vocalizing, thereby changing the vocalization and collecting voice data, thereby having the user reduce trial and error in vocalization and help protect the voice.
[0110] In addition, as the user continuously uses the smartphone application for voice guidance, the direction of change in parameters, the changed values, the amount of change, the standard deviation, etc. are analyzed according to the number of repeated uses, thereby providing data on the stability and degree of change in the user's voice and predicting the possibility of development.
[0111] Additionally, it has the advantage of providing convenience and enjoyment in life, such as providing a vocal guide by combining it with technology that draws the singing into sheet music.
[0112] In addition, by combining and analyzing the user's basic information and voice data obtained through the questionnaire, it has the advantage of providing data that can help improve vocal organ-related diseases and inform the degree of improvement for users with vocal organ-related diseases.
Claims
1. Each step is performed by executing a smartphone application for voice guidance installed on the user's smartphone. a) A step in which a smartphone application for voice guidance (hereinafter referred to as the “smartphone application”) receives voice data of a user speaking for several seconds through a microphone according to the voice method guide or the voice method operated by the user, and recognizes the voice; b) a step in which the smartphone application determines whether the voice is a voice within a preset frequency range by pitch; c) If the voice is within a preset frequency range, the smartphone application determines whether the voice is within the first and second resonance frequency ranges of the presented vowel; d) If the voice is a voice within the first and second resonance frequency ranges, the smartphone application determines whether the voice is a voice within a preset decibel (dB) range; and e) A method for collecting voice data by pronunciation method using a smartphone application, including a step of labeling the voice data as a specific voice and storing it in memory if the voice is within a preset decibel (dB) range.
2. In paragraph 1, A method for collecting voice data by vocalization method using a smartphone application, which further includes a step of determining the user's basic information through a questionnaire in advance in order to analyze the correlation between the collected voice data and the user's basic information when collecting voice data according to the voice data collection procedure prior to the above step a).
3. In paragraph 2, A method for collecting voice data by vocalization method using a smartphone application, wherein the basic information of the user includes at least one of the user's gender, age, nationality, primarily used language, height, weight, presence / absence of voice lesson experience, and presence / absence of vocal organ-related disease.
4. In paragraph 1, A method for collecting voice data by voice method using a smartphone application, characterized in that in step a), the user speaks according to the voice method guide, and in recognizing the voice, the user listens to and speaks a sound of a suggested pitch, or recognizes the pitch of a frequency range set according to the voice spoken by the user.
5. In paragraph 4, A method for collecting voice data by vocalization method using a smartphone application, characterized in that the frequency range set as the recognition range as the pitch presented above is set to exceed the frequency of the pitch that is lower by a semitone and lower by a semitone from the frequency of the pitch, or to exceed the frequency of the pitch that is lower by a semitone and lower by a semitone from the frequency of the pitch.
6. In paragraph 1, A method for collecting voice data by vocalization method using a smartphone application, further comprising a step of requesting the user to respeak with the correct vocalization pitch and guided vocalization method, if the voice is not within a preset frequency range by pitch in the determination of step b).
7. In paragraph 1, A method for collecting voice data by voice method using a smartphone application, further comprising a step of requesting the user to respeak using the vowel and guided voice method presented by the smartphone application, if the voice is not a voice within the first and second resonance frequency ranges of the presented vowel in the determination of the above step c).
8. In paragraph 7, A method for collecting voice data by vocalization method using a smartphone application, characterized in that when the user response to the above request is 'pronounced with the suggested vowel', the smartphone application reflects the user response in labeling the vocalization data for several seconds and stores the vocalization data, or adjusts the first and second resonance frequency ranges of the vowel according to the pitch and vocalization stage.
9. In paragraph 1, A method for collecting voice data by voice method using a smartphone application, further comprising a step of requesting the smartphone application to respeak with a recognizable volume and a guided voice method, if the voice is not within a preset decibel (dB) range in the determination of the above step d).
10. In paragraph 9, A method for collecting voice data by vocalization method using a smartphone application, further comprising a step of requesting vocalization again using the guided vocalization method when the vocalization method instructed in the above-mentioned guided vocalization method requires capturing a change in decibel (dB), if the change in decibel (dB) is not clear in relation to the guided vocalization method.
11. In paragraph 1, A method for collecting voice data by vocalization method using a smartphone application, characterized in that the preset decibel (dB) range in the above step d) is 40 to 90 dB.
12. In paragraph 1, A method for collecting voice data by vocalization method using a smartphone application, characterized in that in step e), when the smartphone application labels the voice data as a specific voice, the voice data is labeled as a voice produced by manipulating the vocal organ or manipulating the vocalization method according to the pitch corresponding to the set frequency range, the suggested vowel, and the vocalization method guide.
13. In paragraph 1, A method for collecting voice data by voice method using a smartphone application, characterized in that in step e), when the smartphone application labels the voice data as a specific voice, the user's condition on the day is checked before and after using each step of the voice method guide and reflected in the labeling of the voice data.
14. In paragraph 13, A method for collecting voice data by voice method using a smartphone application, characterized in that the user's state reaction, including discomfort during voice pronunciation, is confirmed and reflected in the labeling of voice data.
15. In paragraph 1, A method for collecting voice data by voice method using a smartphone application, characterized in that in step e), when the smartphone application labels the voice data as a specific voice, the first time the user uses each step of the voice method guide to label and save the voice data is considered as the first time, and if the user uses it repeatedly thereafter, the number of times the repetition is performed is reflected in the labeling.
16. In paragraph 1, A method for collecting voice data by voice method using a smartphone application, characterized in that some steps of the voice method guide in steps a) to e) above collect voice data according to the result of manipulating a specific vocal organ using a leading voice or a leading vowel.
17. In paragraph 16, A method for collecting voice data by vocalization method using a smartphone application, characterized in that voice data according to the result of manipulating a specific vocal organ using the above-mentioned leading voice or leading vowel is collected as a single voice data, or changes in the first and second resonance frequencies of the vowels of the voice data according to the result of manipulating a specific vocal organ with respect to the first and second resonance frequencies of the above-mentioned leading voice or leading vowel are captured, and voice data corresponding to the part of the vocalization of the above-mentioned leading voice or leading vowel and voice data according to the result of manipulating a specific vocal organ are collected separately.
18. In paragraph 1, A method for collecting voice data by vocalization method using a smartphone application, characterized in that when collecting voice data by manipulating the vocal organ or manipulating the vocalization method according to the vocalization method guide in steps a) to e), when a specific pitch is presented, different pitches are presented for women, men, and male / female children.
Citation Information
Patent Citations
Voice training support program, voice training support method, and voice training support apparatus
JP2023155606A
System and Method for Foreign Language Learning based on Loud Speaking
KR101104822B1
Voice self-practice method for voice disorders and user device for voice therapy
KR102484006B1
System and method for correcting pronunciation
KR102591045B1
Signal way selection circuit for LED lamp and system for improving power efficiency having the smae
KR102663109B1