Electronic apparatus that supports to perform pronunciation evaluation based on the sentence for pronunciation evaluation and operating method thereof
Patent Information
- Application Number
- KR1020230000543
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-01-03
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2043-01-03
Smart Images

Figure 112023000651337-PAT00013_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to an electronic device that supports performing a pronunciation evaluation based on a sentence for pronunciation evaluation and a method of operating the same. Background Technology
[0002] Recently, as interest in Hallyu culture grows, the number of foreigners visiting Korea is increasing, and consequently, the number of foreigners seeking to learn Korean is also on the rise.
[0003] In this regard, many foreigners are learning the four major areas of the Korean language—listening, reading, speaking, and writing—at language schools and private academies, but they are complaining of difficulties. In particular, they are struggling with pronunciation in the speaking area.
[0004] Meanwhile, with the recent advancement of speech recognition technology, services are emerging that display the pronunciation of a specific Korean word in Korean characters when a user speaks that word into a microphone. Accordingly, one may consider introducing a sentence pronunciation evaluation service system that, upon receiving a request for a pronunciation evaluation from a user, instructs the user to speak the sentence into the microphone, generates a new sentence from the user's voice using speech recognition technology, and calculates a pronunciation evaluation score by comparing the original sentence with the generated sentence.
[0005] If such a system is introduced, users will be able to overcome difficulties related to pronunciation by receiving a more objective evaluation of their own pronunciation.
[0006] Therefore, research is needed on technology that supports pronunciation evaluation based on sentences. The problem to be solved
[0007] The present invention aims to support users in receiving a more objective evaluation of their pronunciation by presenting an electronic device and a method of operation thereof that support performing a pronunciation evaluation based on a sentence for pronunciation evaluation. means of solving the problem
[0008] An electronic device that supports performing a pronunciation evaluation based on a sentence for pronunciation evaluation according to an embodiment of the present invention comprises: a sentence storage unit in which a plurality of sentences pre-designated for pronunciation evaluation—each of the plurality of sentences being composed of Hangul—are stored; a display unit that, when a pronunciation evaluation command for a first sentence, which is one of the plurality of sentences, is applied from a user of the electronic device, displays the first sentence on a screen and displays a pronunciation instruction message on a screen that includes content instructing the user to pronounce the first sentence aloud; a recognition sentence generation unit that, after the first sentence and the pronunciation instruction message are displayed on the screen, when the user's voice regarding the first sentence is input through a microphone connected to the electronic device, inputs the user's voice as input to a pre-built speech recognition model to perform speech recognition, thereby generating a first recognition sentence for the user's voice; and, when the first recognition sentence is generated, calculates a pronunciation evaluation score for the first sentence based on the match ratio between the first sentence and the first recognition sentence, and the silent segment ratio, which is the ratio of a silent segment where no voice exists within the entire segment constituting the user's voice, and displays the result on a screen. It includes a pronunciation evaluation score processing unit that displays on the screen.
[0009] In addition, a method of operation of an electronic device that supports performing a pronunciation evaluation based on a sentence for pronunciation evaluation according to an embodiment of the present invention comprises: maintaining a sentence storage unit in which a plurality of sentences pre-designated for pronunciation evaluation—each of the plurality of sentences being composed of Hangul—are stored; when a pronunciation evaluation command for a first sentence, which is one of the plurality of sentences, is applied from a user of the electronic device, the first sentence is displayed on a screen, and a pronunciation instruction message containing content instructing to pronounce the first sentence as speech is displayed on the screen; after the first sentence and the pronunciation instruction message are displayed on the screen, when the user's voice regarding the first sentence is input through a microphone connected to the electronic device, the user's voice is input as an input to a pre-built speech recognition model to perform speech recognition, thereby generating a first recognized sentence for the user's voice; and when the first recognized sentence is generated, a pronunciation evaluation score for the first sentence is based on the match ratio between the first sentence and the first recognized sentence, and the silent segment ratio, which is the ratio of a silent segment where no speech exists within the entire segment constituting the user's voice. It includes the step of calculating and displaying on the screen. Effects of the invention
[0010] The present invention provides an electronic device and a method of operation thereof that support performing a pronunciation evaluation based on a sentence for pronunciation evaluation, thereby enabling a user to receive a more objective evaluation of their own pronunciation. Brief explanation of the drawing
[0011] FIG. 1 is a diagram illustrating the structure of an electronic device that supports performing a pronunciation evaluation based on a sentence for pronunciation evaluation according to an embodiment of the present invention. FIG. 2 is a flowchart illustrating the operation method of an electronic device that supports performing a pronunciation evaluation based on a sentence for pronunciation evaluation according to an embodiment of the present invention. Specific details for implementing the invention
[0012] Embodiments according to the present invention will be described in detail below with reference to the accompanying drawings. This description is not intended to limit the present invention to specific embodiments and should be understood to include all modifications, equivalents, and substitutions that fall within the spirit and scope of the present invention. Similar reference numerals have been used for similar components in describing each drawing, and unless otherwise defined, all terms used in this specification, including technical or scientific terms, have the same meaning as generally understood by a person skilled in the art to which the present invention pertains.
[0013] In this document, when a part is described as "including" a component, it means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, in various embodiments of the present invention, each component, functional block, or means may be composed of one or more sub-components, and the electrical, electronic, or mechanical functions performed by each component may be implemented by various known devices or mechanical elements, such as electronic circuits, integrated circuits, and ASICs (Application Specific Integrated Circuits), and may be implemented separately or two or more may be integrated into one.
[0014] Meanwhile, the blocks in the attached block diagram or the steps in the flowchart may be interpreted as computer program instructions that perform designated functions by being loaded into the processor or memory of data-processing equipment, such as general-purpose computers, specialized computers, portable notebook computers, and network computers. Since these computer program instructions may be stored in memory provided in a computer device or in memory readable by a computer, the functions described in the blocks in the block diagram or the steps in the flowchart may be produced as manufactured products containing means of instruction to perform them. Furthermore, each block or each step may represent a module, segment, or part of code containing one or more executable instructions for executing a specific logical function(s). Also, it should be noted that in some alternative embodiments, the functions mentioned in the blocks or steps may be executed in a different order than the prescribed order. For example, two blocks or steps shown in succession may be performed substantially simultaneously or in reverse order, and in some cases, some blocks or steps may be omitted.
[0015] FIG. 1 is a diagram illustrating the structure of an electronic device that supports performing a pronunciation evaluation based on a sentence for pronunciation evaluation according to an embodiment of the present invention.
[0016] Referring to FIG. 1, the electronic device (110) according to the present invention includes a sentence storage unit (111), a display unit (112), a recognized sentence generation unit (113), and a pronunciation evaluation score processing unit (114).
[0017] A sentence storage unit (111) stores a plurality of pre-designated sentences for pronunciation evaluation (each of the plurality of sentences is composed of Hangul).
[0018] For example, if a plurality of pre-specified sentences are assumed to be 'Sentence 1, Sentence 2, Sentence 3, ...', the sentence storage unit (111) may store the plurality of sentences, 'Sentence 1, Sentence 2, Sentence 3, ...'.
[0019] When a pronunciation evaluation command for a first sentence, which is one of the plurality of sentences, is issued by a user of the electronic device (110), the display unit (112) displays the first sentence on the screen and displays a pronunciation instruction message on the screen that includes content instructing to pronounce the first sentence in speech.
[0020] After the first sentence and the pronunciation instruction message are displayed on the screen, the recognition sentence generation unit (113) generates a first recognition sentence for the user's voice by inputting the user's voice to a pre-built voice recognition model and performing voice recognition when the user's voice is input through a microphone connected to the electronic device (110).
[0021] For example, let us assume that a pronunciation evaluation command for ‘Sentence 2’, which is the first sentence among the plurality of sentences ‘Sentence 1, Sentence 2, Sentence 3, ...’, is issued by a user of the electronic device (110).
[0022] Then, the display unit (112) can display the first sentence, ‘Sentence 2’, on the screen and display a pronunciation instruction message on the screen that includes content instructing to pronounce ‘Sentence 2’ aloud.
[0023] Thus, after the first sentence and the pronunciation instruction message are displayed on the screen by the display unit (112), when the user's voice for 'sentence 2' is input through a microphone connected to the electronic device (110), the recognition sentence generation unit (113) can generate a first recognition sentence for the user's voice by inputting the user's voice as input to a pre-built voice recognition model and performing voice recognition.
[0024] When the first recognized sentence is generated, the pronunciation evaluation score processing unit (114) calculates and displays a pronunciation evaluation score for the first sentence based on the matching ratio between the first sentence and the first recognized sentence, and the silent segment ratio, which is the ratio of the silent segment where no sound exists in the entire segment constituting the user's voice.
[0025] At this time, according to one embodiment of the present invention, the pronunciation evaluation score processing unit (114) may include a pronunciation accuracy score extraction unit (115), a pronunciation fluency score extraction unit (116), and a score display unit (117).
[0026] When the first recognition sentence is generated, the pronunciation accuracy score extraction unit (115) calculates the match ratio between the first sentence and the first recognition sentence, and then extracts the first pronunciation accuracy score recorded in the range value to which the match ratio belongs from a preset accuracy table.
[0027] Here, the accuracy table records a plurality of preset first ratio range values and a preset pronunciation accuracy score corresponding to each of the plurality of first ratio range values.
[0028] For example, the accuracy table above may have information recorded as shown in Table 1 below.
[0030] Multiple first ratio range values Pronunciation accuracy score 0 or more and less than 0.5 0 points 0.5 or more and less than 0.6 20 points 0.6 or more and less than 0.65 40 points 0.65 or higher and less than 0.7 60 points 0.7 or more and less than 0.8 80 points 0.8 or more and 1 or less 100 points
[0032] At this time, according to one embodiment of the present invention, the pronunciation accuracy score extraction unit (115) may include a character generation unit (118), an LCS length calculation unit (119), and a match ratio calculation unit (120).
[0033] The character generation unit (118) generates characters corresponding to the first sentence by dividing the text constituting the first sentence into character units, and generates characters corresponding to the first recognition sentence by dividing the text constituting the first recognition sentence into character units.
[0034] The LCS length calculation unit (119) calculates the length of the longest common subsequence (LCS) between the individual characters corresponding to the first sentence and the individual characters corresponding to the first recognition sentence.
[0035] Here, LCS refers to a sequence of matching strings in two strings. For example, if there are 'ACAYKP' and 'CAPCAK', the LCS between the two strings is 'ACAK', and the length of the LCS is '4'.
[0036] When the length of the LCS is calculated, the matching ratio calculation unit (120) calculates the ratio (B / A) of the length of the LCS (B) to the total number (A) of letters corresponding to the first sentence as the matching ratio.
[0037] For example, if the first recognition sentence generated by the recognition sentence generation unit (113) is assumed to be ‘a monkey appeared in the mile’, the character generation unit (118) included in the pronunciation accuracy score extraction unit (115) can generate characters corresponding to ‘sentence 2’ as ‘ㅇㅜㅓㄴㅅㅜㅇㅣㄱㅏㅁㅏㅇㅡㄹㅇㅔなㅏㅌㅏなㅏㅆㄷㅏ’ by dividing the text constituting the first recognition sentence, ‘sentence 2’, which is ‘a monkey appeared in the village’, into character units, and can generate characters corresponding to the first recognition sentence as ‘ㅇㅜㅓㄴㅅㅜㅇㅣㄱㅏㅁㅏㅇㅣㄹㅇㅔなㅏㄷㅏなㅏㄷㅏ’ by dividing the text constituting the first recognition sentence, which is ‘a monkey appeared in the mile’, into character units.
[0038] Then, the LCS length calculation unit (119) can calculate the length of the longest common subsequence between the individual letters corresponding to 'sentence 2', 'ㅇㅜㅓㄴㅅㅜㅇㅣㄱㅏㅁㅏㅇㅡㄹㅇㅔなㅏㅌㅏなㅏㅆㄷㅏ', and the individual letters corresponding to the first recognition sentence, 'ㅇㅜㅓㄴㅅㅜㅇㅣㄱㅏㅁㅏㅇㅣㄹㅇㅔなㅏㄷㅏなㅏㄷㅏ', as '23'.
[0039] Thus, when the length of the LCS is calculated by the LCS length calculation unit (119), the match ratio calculation unit (120) can calculate the match ratio as the ratio (B / A) of the length (B) of the LCS, '23', to the total number (A) of the letters corresponding to 'sentence 2', which is '26 (pieces)'.
[0040] After that, the pronunciation accuracy score extraction unit (115) can extract the pronunciation accuracy score of ‘100 points’, which is recorded in the range value of ‘0.8 or more and 1 or less’ to which the match ratio of ‘0.885’ belongs, from the accuracy table as shown in Table 1 above, as the first pronunciation accuracy score.
[0041] When the first pronunciation accuracy score is extracted, the pronunciation fluency score extraction unit (116) calculates the silent interval ratio, which is the ratio of the silent interval to the entire interval constituting the user's voice, and then extracts the first pronunciation fluency score recorded in the range value to which the silent interval ratio belongs from a preset fluency table.
[0042] Here, the fluency table records a plurality of preset second ratio range values and a preset pronunciation fluency score corresponding to each of the plurality of second ratio range values.
[0043] For example, the above fluency table may have information recorded as shown in Table 2 below.
[0045] Multiple second ratio range values Pronunciation fluency score 0 or more, less than 0.2 100 points 0.2 or more and less than 0.3 80 points 0.3 or more and less than 0.35 60 points 0.35 or more and less than 0.4 40 points 0.4 or more and less than 0.5 20 points 0.5 or more and less than 1 0 points
[0047] At this time, according to one embodiment of the present invention, the pronunciation fluency score extraction unit (116) may include a silent interval detection unit (121) and a silent interval ratio calculation unit (122).
[0048] The silent section detection unit (121) detects sections in which there is no voice for a period of time longer than a preset first time in the entire section constituting the user's voice as the silent sections.
[0049] When the silent section ratio calculation unit (122) detects the silent section, it generates a summation time by summing the time of each of the sections detected as the silent section, and then calculates the ratio (D / C) of the summation time (D) to the time (C) of the entire section constituting the user's voice as the silent section ratio.
[0050] For example, as in the example above, if the first pronunciation accuracy score is extracted as '100 points' by the pronunciation accuracy score extraction unit (115), the silent section detection unit (121) included in the pronunciation fluency score extraction unit (116) can detect sections in which there is no voice for a continuous period of time longer than a preset first time in the entire section constituting the user's voice as the silent sections.
[0051] Here, assuming that the entire range constituting the user’s voice is ‘0 to 5.19 seconds’ and the preset first time is ‘0.1 seconds’, the silent range detection unit (121) can detect the silent ranges in which there is no voice for a continuous period of at least ‘0.1 seconds’, the preset first time, within the entire range constituting the user’s voice, ‘0 to 5.19 seconds’.
[0052] At this time, if the silent section is detected as ‘Section 1, Section 2, Section 3’ by the silent section detection unit (121), the silent section ratio calculation unit (122) can generate a summation time by summing up the time of each of ‘Section 1, Section 2, Section 3’.
[0053] If the above summing time is assumed to be '1.53 seconds', the silent interval ratio calculation unit (122) can calculate the silent interval ratio as '0.294', which is the ratio (D / C) of the above summing time (D) '1.53 seconds' to the time (C) of the entire interval constituting the user's voice, which is '5.19 seconds'.
[0054] After that, the pronunciation fluency score extraction unit (116) can extract the pronunciation fluency score of ‘80 points’ as the first pronunciation fluency score, which is recorded in the range value of ‘0.2 or more and less than 0.3’ to which the silent interval ratio of ‘0.294’ belongs, from the fluency table as shown in Table 2 above.
[0055] When the first pronunciation fluency score is extracted, the score display unit (117) calculates the average of the first pronunciation accuracy score and the first pronunciation fluency score, and then designates the average as the pronunciation evaluation score for the first sentence and displays it on the screen.
[0056] For example, as in the example above, if the first pronunciation accuracy score is extracted as '100 points' by the pronunciation accuracy score extraction unit (115) and the first pronunciation fluency score is extracted as '80 points' by the pronunciation fluency score extraction unit (116), the score display unit (117) can calculate the average of the first pronunciation accuracy score '100 points' and the first pronunciation fluency score '80 points' as '90 points'.
[0057] After that, the score display unit (117) can designate the above average '90 points' as the pronunciation evaluation score for the above first sentence, 'sentence 2', and display it on the screen.
[0058] According to one embodiment of the present invention, the electronic device (110) may further include a configuration for generating history data and storing it in a pre-designated cloud storage server (140) in an encrypted manner.
[0059] In this regard, according to one embodiment of the present invention, the electronic device (110) may further include a password storage unit (123), a pseudo-random number generation function storage unit (124), a history data generation unit (125), a storage event generation unit (126), a pseudo-random number generation unit (127), an operation matrix generation unit (128), a serial number generation unit (129), an encryption unit (130), and a storage processing unit (131).
[0060] The password storage unit (123) stores a password of n digits (n is a natural number greater than or equal to 2) that has been pre-issued to the user.
[0061] For example, if n is '9' and the 9-digit password previously issued to the user is '827570859', the password storage unit (123) may store the password '827570859'.
[0062] The pseudo-random number generation function storage unit (124) stores different t (t is a natural number greater than or equal to 2) pseudo-random number generation functions that are pre-set.
[0063] Here, pseudo-random numbers refer to random numbers generated by a predetermined random number generation algorithm.
[0064] For example, if t is '3' and three different pre-set pseudo-random number generation functions are 'pseudo-random number generation function 1, pseudo-random number generation function 2, pseudo-random number generation function 3', the pseudo-random number generation function storage unit (124) may store the three pseudo-random number generation functions, namely 'pseudo-random number generation function 1, pseudo-random number generation function 2, pseudo-random number generation function 3'.
[0065] After the pronunciation evaluation score for the first sentence is displayed on the screen, the history data generation unit (125) generates a first date and time information for the time when the pronunciation evaluation history saving command for the first sentence is granted by the user, and then generates first history data consisting of voice data for the user's voice, the pronunciation evaluation score for the first sentence, and the first date and time information.
[0066] When the first history data is generated, the storage event generating unit (126) generates a storage event to encrypt the first history data and store it in the cloud storage server (140).
[0067] When the storage event occurs, the pseudo-random number generator (127) designates the number of characters included in the first sentence as a seed value for pseudo-random number generation, and then applies the seed value as an input to each of the t pseudo-random number generation functions stored in the pseudo-random number generation function storage unit (124) to generate t pseudo-random numbers.
[0068] When the t pseudo-random numbers are generated, the operation matrix generation unit (128) generates a t-dimensional column vector having each of the t pseudo-random numbers as a component and a t-dimensional row vector having each of the t pseudo-random numbers as a component, and then generates an operation matrix of size txt by performing a matrix multiplication between the column vector and the row vector.
[0069] When the above operation matrix is generated, the serial number generation unit (129) generates t constituting the above operation matrix 2 A first serial number is generated by concatenating the components in ascending order.
[0070] When the first serial number is generated, the encryption unit (130) generates a modulo value by performing a modulo operation with the password as the divisor on the first serial number, and then generates a first hash value by applying the modulo value as input to a preset hash function, and generates a first encrypted history data by encrypting the first history data with the first hash value.
[0071] Here, modulo operation refers to an operation that performs division by dividing the dividend by the divisor to produce the remainder.
[0072] When the first encryption history data is generated, the storage processing unit (131) completes the storage process for the first history data by storing the first sentence and the first encryption history data in correspondence with each other in the cloud storage server (140).
[0073] Below, the operation of the history data generation unit (125), the storage event generation unit (126), the pseudo-random number generation unit (127), the operation matrix generation unit (128), the serial number generation unit (129), the encryption unit (130), and the storage processing unit (131) will be explained in detail with examples.
[0074] First, let us assume that, as in the example above, the pronunciation evaluation score of ‘90 points’ for the first sentence, ‘sentence 2’, is displayed on the screen by the pronunciation evaluation score processing unit (114).
[0075] At this time, when a command to save the pronunciation evaluation history for the first sentence, 'Sentence 2', is granted by the user, the history data generation unit (125) can generate first date and time information for the time when the command to save the pronunciation evaluation history is granted, and then generate first history data consisting of voice data for the user's voice, '90 points' which is the pronunciation evaluation score for the first sentence, 'Sentence 2', and the first date and time information.
[0076] Here, assuming that the first date and time information generated by the history data generation unit (125) is '16:30 on December 19, 2022', the history data generation unit (125) can generate first history data consisting of voice data for the user's voice, a pronunciation evaluation score of '90 points' for the first sentence 'sentence 2', and the first date and time information '16:30 on December 19, 2022'.
[0077] Thus, when the first history data is generated by the history data generation unit (125), the storage event generation unit (126) can generate a storage event to encrypt the first history data and store it in the cloud storage server (140).
[0078] Thus, when the storage event is generated by the storage event generating unit (127), the pseudo-random number generating unit (127) may specify '13 (characters)', which is the number of characters included in 'a monkey appeared in the village' of the first sentence 'Sentence 2', as a seed value for pseudo-random number generation. (In the present invention, spaces are treated as characters.)
[0079] After that, the pseudo-random number generator (127) can generate three pseudo-random numbers such as '8, 3, 5' by applying the seed value '13' as input to each of the three pseudo-random number generator functions, 'pseudo-random number generator function 1, pseudo-random number generator function 2, pseudo-random number generator function 3', which are stored in the pseudo-random number generator function storage unit (124).
[0080] Thus, when the three pseudo-random numbers are generated by the pseudo-random number generator (127), the operation matrix generator (128) generates a three-dimensional column vector having each of the three pseudo-random numbers, '8, 3, 5', as a component, ' 'and, a 3-dimensional row vector having each of the three aforementioned pseudo-random numbers '8, 3, 5' as a component, ' After generating ', the above column vector ' ' and the above row vector ' By performing matrix multiplication between ' You can create a 3 x 3 operation matrix like '.
[0081] Then, the serial number generation unit (129) is the above operation matrix ' By concatenating the nine components that make up '64, 24, 40, 24, 9, 15, 40, 15, 25' in ascending order, the first serial number can be generated as '91515242425404064'.
[0082] After that, the encryption unit (130) can generate a modulo value of ‘606874488’ by performing a modulo operation with the password ‘827570859’ as the divisor on the first serial number ‘91515242425404064’, and then apply the modulo value ‘606874488’ as input to a preset hash function to generate a first hash value such as ‘26d103c620a31345ag57qat327891703’.
[0083] Thus, when the first hash value is generated by the encryption unit (130), the encryption unit (130) can generate the first encrypted history data by encrypting the first history data with the first hash value, '26d103c620a31345ag57qat327891703'.
[0084] Then, the storage processing unit (131) can complete the storage process for the first history data by storing the first sentence, 'sentence 2', and the first encrypted history data in correspondence with each other in the cloud storage server (140).
[0085] At this time, according to one embodiment of the present invention, the electronic device (110) may further include a configuration for querying the first history data stored in an encrypted state on a cloud storage server (140).
[0086] In this regard, according to one embodiment of the present invention, the electronic device (110) may further include an input instruction message display unit (132), a restoration unit (133), and a history data lookup unit (134).
[0087] The input instruction message display unit (132) receives the first sentence and the first encryption history data from the cloud storage server (140) and, when a query command instructing the user to query the first history data is authorized, displays an input instruction message on the screen that includes the content instructing the user to input a password for decrypting the first encryption history data.
[0088] When the restoration unit (133) inputs the n-digit password as a password for decrypting the first encryption history data by the user, it sets the number of characters included in the first sentence as the seed value, applies the seed value as input to each of the t pseudo-random number generation functions stored in the pseudo-random number generation function storage unit (124) to generate the t pseudo-random numbers, generates the t-dimensional column vector having each of the t pseudo-random numbers as a component and the t-dimensional row vector having each of the t pseudo-random numbers as a component, and generates the operation matrix by performing a matrix multiplication between the column vector and the row vector, and the t constituting the operation matrix 2 After generating the first serial number by concatenating the components in ascending order, the modulo value is generated by performing a modulo operation on the first serial number with the password entered by the user as the divisor, and the first hash value is generated by applying the modulo value as input to the hash function, and the first history data is restored by decrypting the first encryption history data with the first hash value.
[0089] When the first history data is restored, the history data lookup unit (134) plays voice data for the user's voice included in the first history data and simultaneously displays the first date and time information and the pronunciation evaluation score for the first sentence on the screen.
[0090] Below, the operation of the input instruction message display unit (132), the restoration unit (133), and the history data lookup unit (134) will be explained in detail with examples.
[0091] First, let us assume that the first sentence, 'Sentence 2', and the first encryption history data are stored in the cloud storage server (140) through the storage processing unit (131), as in the example described above.
[0092] At this time, when a query command instructing the user to query the first history data is authorized, the input instruction message display unit (132) can receive the first sentence, 'sentence 2', and the first encrypted history data from the cloud storage server (140).
[0093] Thus, when the first sentence and the first encryption history data are received from the cloud storage server (140), the input instruction message display unit (132) can display an input instruction message on the screen that includes instructions to enter a password for decrypting the first encryption history data.
[0094] At this time, the user can see this input instruction message and input the password they have memorized onto the electronic device (110).
[0095] Thus, when the 9-digit password '827570859' is entered by the user as a password for decrypting the first encryption history data, the recovery unit (133) sets '13', which is the number of characters included in 'a monkey appeared in the village' of the first sentence 'Sentence 2', as the seed value, and then applies the seed value '13' as an input to each of the three pseudo-random number generation functions, 'pseudo-random number generation function 1, pseudo-random number generation function 2, pseudo-random number generation function 3', which are stored in the pseudo-random number generation function storage unit (124), thereby generating the three pseudo-random numbers such as '8, 3, 5'.
[0096] Thus, when the three pseudo-random numbers are generated by the restoration unit (133), the restoration unit (133) generates a three-dimensional column vector, '8, 3, 5', each of the three pseudo-random numbers, as a component. 'and, the 3-dimensional row vector having each of the three pseudo-random numbers '8, 3, 5' as a component, ' After generating ', the above column vector ' ' and the above row vector ' By performing matrix multiplication between ' The above operation matrix can be generated as shown above.
[0097] Then, the restoration unit (133) is the above operation matrix ' By concatenating the nine components constituting '64, 24, 40, 24, 9, 15, 40, 15, 25' in ascending order, the above-mentioned first serial number such as '91515242425404064' can be generated.
[0098] After that, the restoration unit (133) can restore the first history data by performing a modulo operation with the password '827570859' entered by the user on the first serial number '91515242425404064' to generate a modulo value '606874488', and then applying the modulo value '606874488' as input to the hash function to generate the first hash value such as '26d103c620a31345ag57qat327891703', and decrypting the first encryption history data into the first hash value '26d103c620a31345ag57qat327891703'.
[0099] Thus, when the first history data is restored by the restoration unit (133), the history data lookup unit (134) can play voice data for the user's voice included in the first history data, and at the same time display on the screen the first date and time information, 'December 19, 2022, 16:30', and the pronunciation evaluation score of '90 points' for the first sentence, 'Sentence 2'.
[0100] FIG. 2 is a flowchart illustrating the operation method of an electronic device that supports performing a pronunciation evaluation based on a sentence for pronunciation evaluation according to an embodiment of the present invention.
[0101] In step (S210), a sentence storage unit is maintained in which a plurality of pre-designated sentences for pronunciation evaluation (each of the plurality of sentences is composed of Hangul) are stored.
[0102] In step (S220), when a pronunciation evaluation command for a first sentence, which is one of the plurality of sentences, is issued from the user of the electronic device, the first sentence is displayed on the screen, and a pronunciation instruction message containing instructions to pronounce the first sentence in speech is displayed on the screen.
[0103] In step (S230), after the first sentence and the pronunciation instruction message are displayed on the screen, when the user's voice for the first sentence is input through a microphone connected to the electronic device, the user's voice is input to a pre-built voice recognition model to perform voice recognition, thereby generating a first recognized sentence for the user's voice.
[0104] In step (S240), when the first recognized sentence is generated, a pronunciation evaluation score for the first sentence is calculated and displayed on the screen based on the matching ratio between the first sentence and the first recognized sentence, and the silent interval ratio, which is the ratio of the silent interval where no speech exists within the entire interval constituting the user's speech.
[0105] At this time, according to one embodiment of the present invention, in step (S240), when the first recognition sentence is generated, the matching ratio between the first sentence and the first recognition sentence is calculated, and then a first pronunciation accuracy score is extracted from a preset accuracy table (the accuracy table contains a plurality of preset first ratio range values and a preset pronunciation accuracy score corresponding to each of the plurality of first ratio range values). Then, when the first pronunciation accuracy score is extracted, the silent segment ratio, which is the ratio of the silent segment to the entire segment constituting the user's voice, is calculated, and then a first pronunciation fluency score is extracted from a preset fluency table (the fluency table contains a plurality of preset second ratio range values and a preset pronunciation fluency score corresponding to each of the plurality of second ratio range values). When the first pronunciation fluency score is extracted, the first pronunciation accuracy score and the first pronunciation The method may include the step of calculating the average of the fluency scores, and then designating the average as the pronunciation evaluation score for the first sentence and displaying it on the screen.
[0106] At this time, according to one embodiment of the present invention, the step of extracting the first pronunciation accuracy score may include: generating individual characters corresponding to the first sentence by dividing the text constituting the first sentence into individual character units, and generating individual characters corresponding to the first recognized sentence by dividing the text constituting the first recognized sentence into individual character units; calculating the length of the longest common subsequence between the individual characters corresponding to the first sentence and the individual characters corresponding to the first recognized sentence; and, when the length of the LCS is calculated, calculating the ratio (B / A) of the length of the LCS to the total number (A) of individual characters corresponding to the first sentence as the match ratio.
[0107] Additionally, according to one embodiment of the present invention, the step of extracting the first pronunciation fluency score may include the step of detecting as silent sections sections in which there is no voice for a period of time longer than a preset first time in the entire section constituting the user's voice, and when the silent sections are detected, generating a summation time by summing the time of each of the sections detected as silent sections, and then calculating the ratio (D / C) of the summation time (D) to the time (C) of the entire section constituting the user's voice as the silent section ratio.
[0108] Additionally, according to one embodiment of the present invention, a method of operating the electronic device comprises the steps of: maintaining a password storage unit in which a password of n digits (where n is a natural number greater than or equal to 2) pre-issued to the user is stored; maintaining a pseudo-random number generation function storage unit in which t different (where t is a natural number greater than or equal to 2) pre-set pseudo-random number generation functions are stored; after a pronunciation evaluation score for the first sentence is displayed on a screen, when a command to save the pronunciation evaluation history for the first sentence is authorized from the user, generating first date and time information for the time when the command to save the pronunciation evaluation history is authorized, and then generating first history data composed of voice data for the user's voice, a pronunciation evaluation score for the first sentence, and the first date and time information; when the first history data is generated, generating a storage event to encrypt the first history data and store it on a pre-designated cloud storage server; when the storage event occurs, designating the number of characters included in the first sentence as a seed value for pseudo-random number generation, and then the seed value the pseudo A step of generating t pseudo-random numbers by applying them as inputs to each of the t pseudo-random number generation functions stored in the random number generation function storage unit; a step of, once the t pseudo-random numbers are generated, generating a t-dimensional column vector having each of the t pseudo-random numbers as a component and a t-dimensional row vector having each of the t pseudo-random numbers as a component, and then generating an operation matrix of size txt by performing a matrix multiplication between the column vector and the row vector; and, once the operation matrix is generated, t constituting the operation matrix 2A step of generating a first serial number by concatenating components in ascending order; a step of, when the first serial number is generated, generating a modulo value by performing a modulo operation on the first serial number with the password as the divisor, and then generating a first hash value by applying the modulo value as input to a preset hash function, and generating a first encrypted history data by encrypting the first history data with the first hash value; a step of, when the first encrypted history data is generated, completing the storage process for the first history data by storing the first sentence and the first encrypted history data in correspondence with each other in the cloud storage server; a step of, after the first sentence and the first encrypted history data are stored in the cloud storage server, receiving a query command from the user instructing to query the first history data, receiving the first sentence and the first encrypted history data from the cloud storage server, and displaying an input instruction message on the screen containing content instructing to input a password for decrypting the first encrypted history data; and, by the user As a password for decrypting the first encryption history data, when the n-digit password is entered, the number of characters included in the first sentence is designated as the seed value, and the seed value is applied as input to each of the t pseudo-random number generation functions stored in the pseudo-random number generation function storage unit to generate the t pseudo-random numbers, and after generating the t-dimensional column vector having each of the t pseudo-random numbers as a component and the t-dimensional row vector having each of the t pseudo-random numbers as a component, the operation matrix is generated by performing a matrix multiplication between the column vector and the row vector, and the t constituting the operation matrix 2The method may further include the steps of: generating a first serial number by concatenating the components in ascending order; generating a modulo value by performing a modulo operation on the first serial number with the password entered by the user as the divisor; generating a first hash value by applying the modulo value as input to the hash function; and restoring the first history data by decrypting the first encryption history data with the first hash value; and, when the first history data is restored, playing voice data for the user's voice included in the first history data while simultaneously displaying the first date and time information and the pronunciation evaluation score for the first sentence on the screen.
[0109] Hereinafter, a method of operation of an electronic device according to an embodiment of the present invention has been described with reference to FIG. 2. Here, since the method of operation of an electronic device according to an embodiment of the present invention may correspond to the configuration of the operation of the electronic device (110) described using FIG. 1, a more detailed description thereof will be omitted.
[0110] A method of operation for an electronic device that supports performing a pronunciation evaluation based on a sentence for pronunciation evaluation according to an embodiment of the present invention can be implemented as a computer program stored in a storage medium for execution through combination with a computer.
[0111] In addition, a method of operation of an electronic device that supports performing a pronunciation evaluation based on a sentence for pronunciation evaluation according to an embodiment of the present invention may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the present invention, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.
[0112] As described above, the present invention has been explained by specific details such as specific components, limited embodiments, and drawings; however, this is provided merely to aid in a more comprehensive understanding of the invention, and the invention is not limited to the above embodiments. A person skilled in the art can make various modifications and variations from this description.
[0113] Accordingly, the scope of the present invention is not limited to the described embodiments, and all things equivalent to or having equivalent variations to the claims set forth below, as well as the claims set forth below, shall be considered to fall within the scope of the concept of the present invention. Explanation of the symbols
[0114] 110: An electronic device that supports performing pronunciation evaluation based on sentences for pronunciation evaluation 111: Sentence storage unit 112: Display unit 113: Recognition sentence generation unit 114: Pronunciation evaluation score processing unit 115: Pronunciation accuracy score extraction unit 116: Pronunciation fluency score extraction unit 117: Score display unit 118: Character generation unit 119: LCS Length Calculation Unit 120: Matching Ratio Calculation Unit 121: Silent interval detection unit 122: Silent interval ratio calculation unit 123: Password storage unit 124: Pseudo-random number generation function storage unit 125: History data generation unit 126: Save event generation unit 127: Pseudo-random number generator 128: Operation matrix generator 129: Serial number generation unit 130: Encryption unit 131: Storage processing unit 132: Input instruction message display unit 133: Restoration Section 134: History Data Retrieval Section 140: Cloud storage server
Claims
Claim 1 An electronic device that supports performing a pronunciation evaluation based on a sentence for pronunciation evaluation, comprising: a sentence storage unit in which a plurality of sentences pre-designated for pronunciation evaluation—each of said plurality of sentences being composed of Hangul—are stored; a password storage unit in which a password of n digits (where n is a natural number greater than or equal to 2) pre-issued to a user of said electronic device is stored; a pseudo-random number generation function storage unit in which t different (where t is a natural number greater than or equal to 2) pre-set pseudo-random number generation functions are stored; a display unit in which, when a pronunciation evaluation command for a first sentence, which is one of said plurality of sentences, is applied from said user, the first sentence is displayed on a screen and a pronunciation instruction message containing instructions to pronounce said first sentence aloud is displayed on a screen; and, after said first sentence and said pronunciation instruction message are displayed on a screen, when the user's voice regarding said first sentence is input through a microphone connected to said electronic device, said voice is input to a pre-built speech recognition model to perform speech recognition, thereby thereby performing a first recognition of said user's voice. A recognition sentence generation unit that generates a sentence; and a pronunciation evaluation score processing unit that, when the first recognition sentence is generated, calculates a pronunciation evaluation score for the first sentence and displays it on a screen based on the matching ratio between the first sentence and the first recognition sentence, and the silent segment ratio, which is the ratio of a silent segment where no speech exists within the entire segment constituting the user's speech.A history data generation unit that, after the pronunciation evaluation score for the first sentence is displayed on the screen, when a command to save the pronunciation evaluation history for the first sentence is authorized by the user, generates first date and time information regarding the time when the command to save the pronunciation evaluation history is authorized, and then generates first history data composed of voice data for the user's voice, the pronunciation evaluation score for the first sentence, and the first date and time information; a storage event generation unit that, when the first history data is generated, generates a storage event to encrypt the first history data and store it on a pre-designated cloud storage server; a pseudo-random number generation unit that, when the storage event occurs, designates the number of characters included in the first sentence as a seed value for pseudo-random number generation, and then applies the seed value as input to each of the t pseudo-random number generation functions stored in the pseudo-random number generation function storage unit to generate t pseudo-random numbers; and when the t pseudo-random numbers are generated, a t-dimensional column vector having each of the t pseudo-random numbers as a component, and the An operation matrix generation unit that generates an operation matrix of size txt by generating a t-dimensional row vector having each of t pseudo-random numbers as a component, and then performing a matrix multiplication between the column vector and the row vector; when the operation matrix is generated, t constituting the operation matrix; 2 A serial number generation unit that generates a first serial number by concatenating components in ascending order; an encryption unit that, when the first serial number is generated, generates a modulo value by performing a modulo operation on the first serial number with the password as the divisor, generates a first hash value by applying the modulo value as input to a preset hash function, and generates a first encrypted history data by encrypting the first history data with the first hash value; a storage processing unit that, when the first encrypted history data is generated, completes the storage process for the first history data by storing the first sentence and the first encrypted history data in correspondence with each other in the cloud storage server; and, after the first sentence and the first encrypted history data are stored in the cloud storage server, when a query command instructing the user to query the first history data is issued, receives the first sentence and the first encrypted history data from the cloud storage server and displays an input instruction message on the screen containing content instructing the user to input a password for decrypting the first encrypted history data. Input instruction message display unit; when the user inputs the n-digit password as a password for decrypting the first encryption history data, the number of characters included in the first sentence is designated as the seed value, and the seed value is applied as input to each of the t pseudo-random number generation functions stored in the pseudo-random number generation function storage unit to generate the t pseudo-random numbers, and after generating the t-dimensional column vector having each of the t pseudo-random numbers as a component and the t-dimensional row vector having each of the t pseudo-random numbers as a component, the operation matrix is generated by performing a matrix multiplication between the column vector and the row vector, and the t constituting the operation matrix 2 An electronic device comprising: a restoration unit that generates a first serial number by concatenating the components in ascending order, generates a modulo value by performing a modulo operation on the first serial number with a password entered by the user as a divisor, generates a first hash value by applying the modulo value as input to the hash function, and restores the first history data by decrypting the first encryption history data with the first hash value; and a history data lookup unit that, when the first history data is restored, plays voice data for the user's voice included in the first history data and simultaneously displays the first date and time information and the pronunciation evaluation score for the first sentence on a screen. Claim 2 In claim 1, the pronunciation evaluation score processing unit, when the first recognition sentence is generated, calculates the match ratio between the first sentence and the first recognition sentence, and then extracts a first pronunciation accuracy score recorded corresponding to the range value to which the match ratio belongs from a preset accuracy table—in which a plurality of preset first ratio range values and a preset pronunciation accuracy score corresponding to each of the plurality of first ratio range values are recorded—a pronunciation accuracy score extraction unit; when the first pronunciation accuracy score is extracted, calculates the silent segment ratio, which is the ratio occupied by the silent segment in the entire segment constituting the user's voice, and then extracts a first pronunciation fluency score recorded corresponding to the range value to which the silent segment ratio belongs from a preset fluency table—in which a plurality of preset second ratio range values and a preset pronunciation fluency score corresponding to each of the plurality of second ratio range values are recorded—a pronunciation fluency score extraction unit; An electronic device comprising a score display unit that, when the first pronunciation fluency score is extracted, calculates the average of the first pronunciation accuracy score and the first pronunciation fluency score, and then designates the average as the pronunciation evaluation score for the first sentence and displays it on a screen. Claim 3 An electronic device comprising: a pronunciation accuracy score extraction unit that generates individual characters corresponding to the first sentence by dividing the text constituting the first sentence into individual character units, and generates individual characters corresponding to the first recognized sentence by dividing the text constituting the first recognized sentence into individual character units; an LCS length calculation unit that calculates the length of the Longest Common Subsequence (LCS) between the individual characters corresponding to the first sentence and the individual characters corresponding to the first recognized sentence; and a match ratio calculation unit that, when the length of the LCS is calculated, calculates the ratio (B / A) of the length of the LCS to the total number (A) of individual characters corresponding to the first sentence as the match ratio. Claim 4 An electronic device comprising, in paragraph 2, a pronunciation fluency score extraction unit that detects as silent sections sections in which no voice is present for a period of time longer than a preset first time in the entire section constituting the user's voice; and a silent section ratio calculation unit that, when a silent section is detected, generates a summation time by summing the time of each of the sections detected as silent sections, and then calculates the ratio (D / C) of the summation time (D) to the time (C) of the entire section constituting the user's voice as the silent section ratio. Claim 5 delete Claim 6 A method of operation of an electronic device that supports performing a pronunciation evaluation based on a sentence for pronunciation evaluation, comprising: a step of maintaining a sentence storage unit in which a plurality of sentences pre-designated for pronunciation evaluation—each of which is composed of Hangul—are stored; a step of maintaining a password storage unit in which a password of n digits (n is a natural number greater than or equal to 2) pre-issued to a user of the electronic device is stored; a step of maintaining a pseudo-random number generation function storage unit in which t different (t is a natural number greater than or equal to 2) pre-set pseudo-random number generation functions are stored; a step of, when a pronunciation evaluation command for a first sentence, which is one of the plurality of sentences, is applied from the user, displaying the first sentence on a screen and displaying a pronunciation instruction message on a screen that includes content instructing the user to pronounce the first sentence aloud; and, after the first sentence and the pronunciation instruction message are displayed on the screen, when the user's voice regarding the first sentence is input through a microphone connected to the electronic device, inputting the user's voice as input to a pre-established speech recognition model to perform speech recognition, thereby thereby the user A step of generating a first recognition sentence for a voice; and, when the first recognition sentence is generated, a step of calculating and displaying a pronunciation evaluation score for the first sentence on a screen based on a match ratio between the first sentence and the first recognition sentence, and a silent segment ratio, which is the ratio of a silent segment where no voice exists within the entire segment constituting the user's voice.After the pronunciation evaluation score for the first sentence is displayed on the screen, if a command to save the pronunciation evaluation history for the first sentence is authorized by the user, a first date and time information regarding the time when the command to save the pronunciation evaluation history is authorized is generated, and then a first history data composed of voice data for the user's voice, the pronunciation evaluation score for the first sentence, and the first date and time information is generated; when the first history data is generated, a step of generating a storage event to encrypt the first history data and store it on a pre-designated cloud storage server; when the storage event is generated, a step of designating the number of characters included in the first sentence as a seed value for pseudo-random number generation, and then applying the seed value as an input to each of the t pseudo-random number generation functions stored in the pseudo-random number generation function storage unit to generate t pseudo-random numbers; when the t pseudo-random numbers are generated, a t-dimensional column vector having each of the t pseudo-random numbers as a component, and each of the t pseudo-random numbers as a component A step of generating a row vector of t dimension having, and then generating an operation matrix of size txt by performing a matrix multiplication between the column vector and the row vector; when the operation matrix is generated, t constituting the operation matrix; 2 A step of generating a first serial number by concatenating components in ascending order; a step of, when the first serial number is generated, generating a modulo value by performing a modulo operation on the first serial number with the password as the divisor, and then generating a first hash value by applying the modulo value as input to a preset hash function, and generating a first encrypted history data by encrypting the first history data with the first hash value; a step of, when the first encrypted history data is generated, completing the storage process for the first history data by storing the first sentence and the first encrypted history data in correspondence with each other in the cloud storage server; a step of, after the first sentence and the first encrypted history data are stored in the cloud storage server, receiving a query command from the user instructing to query the first history data, receiving the first sentence and the first encrypted history data from the cloud storage server, and displaying an input instruction message on the screen containing content instructing to input a password for decrypting the first encrypted history data; by the user, As a password for decrypting the first encryption history data, when the n-digit password is entered, the number of characters included in the first sentence is designated as the seed value, and the seed value is applied as input to each of the t pseudo-random number generation functions stored in the pseudo-random number generation function storage unit to generate the t pseudo-random numbers, and after generating the t-dimensional column vector having each of the t pseudo-random numbers as a component and the t-dimensional row vector having each of the t pseudo-random numbers as a component, the operation matrix is generated by performing a matrix multiplication between the column vector and the row vector, and the t constituting the operation matrix 2 A method of operation of an electronic device comprising: a step of generating a first serial number by concatenating the components in ascending order, then generating a modulo value by performing a modulo operation on the first serial number with a password entered by the user as a divisor, and generating a first hash value by applying the modulo value as input to the hash function, and then restoring the first history data by decrypting the first encryption history data with the first hash value; and a step of, when the first history data is restored, playing voice data for the user's voice included in the first history data and simultaneously displaying the first date and time information and the pronunciation evaluation score for the first sentence on a screen. Claim 7 In claim 6, the step of calculating and displaying the pronunciation evaluation score on the screen comprises: a step of, when the first recognized sentence is generated, calculating the match ratio between the first sentence and the first recognized sentence, and then extracting a first pronunciation accuracy score recorded corresponding to the range value to which the match ratio belongs from a preset accuracy table—in which a plurality of preset first ratio range values and a preset pronunciation accuracy score corresponding to each of the plurality of first ratio range values are recorded—; and, when the first pronunciation accuracy score is extracted, calculating the silent segment ratio, which is the ratio occupied by the silent segment in the entire segment constituting the user's voice, and then extracting a first pronunciation fluency score recorded corresponding to the range value to which the silent segment ratio belongs from a preset fluency table—in which a plurality of preset second ratio range values and a preset pronunciation fluency score corresponding to each of the plurality of second ratio range values are recorded— A method of operation of an electronic device comprising the step of, when the first pronunciation fluency score is extracted, calculating the average of the first pronunciation accuracy score and the first pronunciation fluency score, and then designating the average as the pronunciation evaluation score for the first sentence and displaying it on a screen. Claim 8 In claim 7, the step of extracting the first pronunciation accuracy score comprises: generating individual characters corresponding to the first sentence by dividing the text constituting the first sentence into individual character units, and generating individual characters corresponding to the first recognized sentence by dividing the text constituting the first recognized sentence into individual character units; calculating the length of the Longest Common Subsequence (LCS) between the individual characters corresponding to the first sentence and the individual characters corresponding to the first recognized sentence; and, when the length of the LCS is calculated, calculating the ratio (B / A) of the length of the LCS (B) to the total number (A) of individual characters corresponding to the first sentence as the match ratio. Claim 9 A method of operation of an electronic device according to claim 7, wherein the step of extracting the first pronunciation fluency score comprises: a step of detecting as silent sections sections in which there is no voice for a period of time longer than a preset first time in the entire section constituting the user’s voice; and a step of, when the silent sections are detected, generating a summation time by summing the time of each of the sections detected as silent sections, and then calculating the ratio (D / C) of the summation time (D) to the time (C) of the entire section constituting the user’s voice as the silent section ratio. Claim 10 delete Claim 11 A computer-readable recording medium having a computer program for executing the method of any one of paragraphs 6 through 9 in combination with a computer. Claim 12 A computer program stored on a storage medium for executing the method of any one of paragraphs 6 through 9 through combination with a computer.
Citation Information
Patent Citations
Mobile terminal and operation method thereof
KR1020140144012A
System and method for studying korean pronunciation using voice analysis
KR1020210086182A
A method and apparatus for providing language evaluation interfaces for pronunciation keywords
KR1020210144016A
Program for providing language evaluation interfaces for pronunciation keywords
KR1020210144019A
Format conversion task allocating apparatus which allocates tasks for converting format of document files to multiple format converting servers and the operating method thereof
KR1020220032158A