Language type-specific customized evaluation device and method based on acoustic feature analysis and phoneme analysis
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- EDUTEM CO LTD
- Filing Date
- 2025-09-24
- Publication Date
- 2026-08-03
Smart Images

Figure 112025109170818-PAT00011_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a technology for evaluating language ability tailored to different language types, and more specifically, to an apparatus and method capable of performing a customized evaluation tailored to different language types by analyzing at the phoneme level and extracting acoustic features not only for speech based on correct answer text but also for autonomous speech without a presented correct answer text. Background Technology
[0003] Existing invention evaluation systems are implemented by calculating pronunciation accuracy by comparing simple speech recognition results with correct text. However, this method has the problem that it fails to adequately reflect high-level speech characteristics, such as phoneme-specific features, prosody, and fluency, and the feedback provided to learners is limited to the level of simple text.
[0004] Furthermore, when performing pronunciation evaluation using a method of forced alignment at the phoneme level, it is limited to simply detecting phoneme error types such as insertion, deletion, and substitution, and there is a problem in that it is difficult to effectively reflect actual speech characteristics, such as discrepancies between word boundaries that occur during natural speech.
[0005] Consequently, there is a possibility that the evaluation system may classify a learner's utterance as a pronunciation error, even if it is actually natural and accurate, thereby generating incorrect evaluation results. Therefore, advanced pronunciation evaluation technology may be required that integrates quantitative evaluation, error type classification, and exception handling based on phoneme-level analysis and acoustic feature analysis. Prior art literature
[0007] Korean Registered Patent No. 10-2274764 (July 2, 2021) The problem to be solved
[0008] One embodiment of the present invention aims to provide a device and method capable of performing customized evaluations by language type by analyzing at the phoneme level and extracting acoustic features for not only speech based on correct text but also autonomous speech without presented correct text. means of solving the problem
[0010] Among the embodiments, the device comprises: a speech voice input unit that receives a user’s speech voice corresponding to a correct answer text expressed in a specific language from a user terminal of a language-type customized evaluation device; a phoneme analysis unit that inputs the speech voice into a phoneme analysis model to generate speech phonemes inferred at the phoneme level; an acoustic feature extraction unit that forcibly aligns the speech phonemes based on the correct answer text and then extracts one or more acoustic features for the speech voice; and an evaluation execution unit that generates an evaluation result for the speech voice using a pronunciation evaluation model based on the speech phonemes and the acoustic features.
[0011] The above-mentioned speech voice input unit can receive autonomous speech voice freely spoken by the user when the above-mentioned correct text is not presented.
[0012] The above-described phoneme analysis unit can divide the user's spoken voice or autonomous speech voice into a plurality of spoken phonemes and generate phoneme feature information including the position, length, and utterance probability of each spoken phoneme as spoken phoneme information.
[0013] The above acoustic feature extraction unit can extract at least one of acoustic features including pitch, intensity, formant, harmonic noise ratio (HNR), duration per phoneme, spectral coefficient, and voice interval per phoneme (VOT) from the user's spoken voice or autonomous speech voice.
[0014] The evaluation unit above can generate evaluation data including a suitability score, pitch, utterance length, and pause interval corresponding to each utterance phoneme according to the result of forced alignment based on the correct answer text or the result of phoneme inference based on autonomous speech, and calculate an evaluation score for the utterance speech using the pronunciation evaluation model based on the evaluation data.
[0015] The evaluation unit described above can calculate at least one partial score and a total score regarding accuracy, completeness, fluency, and prosody for the user's spoken voice or autonomously spoken voice, respectively.
[0016] The evaluation performing unit may include an error detection module that detects an error type corresponding to the position of each utterance phoneme according to the correct answer text-based forced alignment result or the autonomous speech-based phoneme inference result, and generates location and type information of the detected error as structured error data; and an exception processing module that sets an exception condition corresponding to the error type and applies it to the evaluation process of the utterance speech.
[0017] The evaluation performing unit above can generate evaluation data including the error data and the exception conditions above and provide it as input data for the pronunciation evaluation model.
[0018] Among the embodiments, the customized evaluation method by language type comprises: receiving a user’s spoken voice corresponding to a correct answer text expressed in a specific language from a user terminal; inputting the spoken voice into a phoneme analysis model to generate spoken phonemes inferred at the phoneme level; forcibly aligning the spoken phonemes based on the correct answer text and then extracting one or more acoustic features for the spoken voice; and generating an evaluation result for the spoken voice using a pronunciation evaluation model based on the spoken phonemes and the acoustic features. Effects of the invention
[0020] The disclosed technology may have the following effects. However, this does not mean that a specific embodiment must include all of the following effects or only the following effects; therefore, the scope of the rights of the disclosed technology should not be understood as being limited by this.
[0021] A customized evaluation device and method for each language type based on acoustic feature analysis and phoneme analysis according to one embodiment of the present invention can quantify and evaluate various indicators such as pronunciation accuracy, completeness, fluency, and prosody based on phoneme analysis and acoustic feature extraction of spoken speech.
[0022] In addition, according to the present invention, phoneme-unit evaluation is possible for autonomous speech without forced alignment, and more precise and flexible evaluation results can be provided by reflecting natural speech patterns, such as linking or contraction that appear in actual speech, as exception conditions.
[0023] The evaluation results according to the present invention are provided along with feedback by error type, which can provide customized pronunciation correction information to learners and can be utilized in various fields such as language education, pronunciation training, and speaking ability diagnosis. Brief explanation of the drawing
[0025] FIG. 1 is a drawing illustrating an evaluation system according to the present invention. FIG. 2 is a diagram illustrating the system configuration of an evaluation device according to the present invention. FIGS. 3a and 3b are drawings illustrating the functional configuration of an evaluation device according to the present invention. Figure 4 is a flowchart illustrating a customized evaluation method by language type according to the present invention. FIG. 5 is a diagram illustrating an embodiment of a phoneme inference process according to the present invention. Figure 6 is a diagram illustrating the phoneme alignment process according to the present invention. FIG. 7 is a drawing illustrating an example of an evaluation process according to the present invention. FIG. 8 is a drawing illustrating an embodiment of a Korean pronunciation evaluation element and item according to the present invention. FIGS. 9a and 9b are drawings illustrating an example of an evaluation result visualization process according to the present invention. Specific details for implementing the invention
[0026] The description of the present invention is merely an example for structural or functional explanation, and therefore the scope of the present invention should not be interpreted as being limited by the examples described in the text. That is, since the examples are subject to various modifications and may take various forms, the scope of the present invention should be understood to include equivalents capable of realizing the technical concept. Furthermore, the objectives or effects presented in the present invention do not imply that a specific example must include all of them or only such effects; therefore, the scope of the present invention should not be understood as being limited by them.
[0027] Meanwhile, the meaning of the terms described in this application should be understood as follows.
[0028] Terms such as "first," "second," etc., are intended to distinguish one component from another, and the scope of rights shall not be limited by these terms. For example, the first component may be named the second component, and similarly, the second component may be named the first component.
[0029] When it is stated that one component is "connected" to another component, it should be understood that it may be directly connected to that other component, or that there may be other components in between. Conversely, when it is stated that one component is "directly connected" to another component, it should be understood that there are no other components in between. Meanwhile, other expressions describing the relationships between components, such as "between" and "exactly between," or "adjacent to" and "directly adjacent to," should be interpreted in the same way.
[0030] A singular expression should be understood to include a plural expression unless the context clearly indicates otherwise, and terms such as "include" or "have" are intended to specify the existence of the implemented features, numbers, steps, actions, components, parts, or combinations thereof, and should be understood not to preclude the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0031] In each step, identifiers (e.g., a, b, c, etc.) are used for convenience of explanation and do not describe the order of the steps; the steps may occur differently from the specified order unless a specific order is clearly indicated in the context. That is, the steps may occur in the same order as specified, may be performed substantially simultaneously, or may be performed in the reverse order.
[0032] Unless otherwise defined, all terms used herein have the same meaning as generally understood by those skilled in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having meanings consistent with the context of the relevant technology and should not be interpreted as having an ideal or overly formal meaning unless explicitly defined in this application.
[0034] FIG. 1 is a drawing illustrating an evaluation system according to the present invention.
[0035] Referring to FIG. 1, the evaluation system (100) may provide an operating environment that enables a customized evaluation method for each language type according to the present invention to be performed. To this end, the evaluation system (100) may include a user terminal (110) and an evaluation device (130). The evaluation system (100) may be designed to perform analysis and evaluation of a user's spoken voice based on pronunciation rules and error types specialized for a specific language. For example, the evaluation system (100) may apply phoneme analysis, error detection, and exception handling methods based on the initial-medial-final structure of Korean.
[0036] Additionally, the evaluation system (100) can be expanded to receive and evaluate not only reading-based speech in which a correct answer text is presented in advance according to the user's speech method, but also autonomous speech voice freely spoken by the user without a correct answer text. The evaluation system (100) can analyze the spoken voice in the phoneme unit for autonomous speech voice as well, and can calculate quantitative indicators such as pronunciation suitability, prosody, and fluency without forced alignment according to context-based evaluation criteria or phonological rules by language type.
[0037] However, the core configuration of the present invention is not necessarily limited to a specific language and can be adjusted according to the phonological systems and speech characteristics of various languages, such as English, Chinese, and Japanese, so it may be applicable to a multilingual speech evaluation system.
[0038] A user terminal (110) may correspond to a terminal device operated by a user. In an embodiment of the present invention, a user may be understood as one or more users, and each of the one or more users may correspond to one or more user terminals (110). For example, in FIG. 1, the first user may correspond to user terminal 1, and the nth user may correspond to user terminal n (where n is a natural number).
[0039] Additionally, the user terminal (110) may be implemented as one device constituting the evaluation system (100) according to the present invention, and the evaluation system (100) may be implemented in various modified forms according to the purpose of customized evaluation for each language type.
[0040] Additionally, the user terminal (110) can be implemented as a smartphone, laptop, or computer that is connected to and operable with the evaluation device (130), and is not necessarily limited thereto, but can also be implemented as various devices including a tablet PC.
[0041] In particular, the user terminal (110) can install and run a dedicated program or application for evaluating various types of language ability. The user terminal (110) may be implemented to include a microphone module capable of recording user pronunciation and a speaker module capable of playing content including native speaker pronunciation, and may be further implemented to include a camera module for collecting user speech video as needed.
[0042] Meanwhile, a user terminal (110) can be connected to an evaluation device (130) via a network, and multiple user terminals (110) can be connected to the evaluation device (130) simultaneously.
[0043] The evaluation device (130) may be implemented as a computer or server that evaluates the user's language ability. Additionally, the evaluation device (130) may be connected to the user terminal (110) via a wired network or a wireless network such as Bluetooth, WiFi, LTE, etc., and may transmit and receive data with the user terminal (110) through the network.
[0044] Additionally, the evaluation device (130) may be implemented to operate in connection with an independent external system (not shown in FIG. 1) to provide related services. For example, the evaluation device (130) may be implemented to provide various services by linking with a foreign language learning system, an artificial intelligence system, and a blockchain system. In one embodiment, the evaluation device (130) may be implemented as a cloud server and may perform the operation of providing big data information or user-customized information related to language ability evaluation through a cloud service.
[0045] In one embodiment, the evaluation device (130) may include a database (not shown in FIG. 1) that stores various information required during the relevant operation process. In this case, the database may store evaluation data such as content for language ability evaluation and the user's spoken voice, and may store information regarding artificial intelligence models for phoneme inference or pronunciation evaluation, but is not necessarily limited thereto, and may store information collected or processed in various forms during the process in which the evaluation device (130) performs a customized evaluation method for each language type according to the present invention. For example, the database may store contextual information required for autonomous speech voice evaluation, comparison group voice features, or generalized evaluation criteria information for each language type.
[0047] FIG. 2 is a diagram illustrating the system configuration of an evaluation device according to the present invention.
[0048] Referring to FIG. 2, the evaluation device (130) may include a processor (210), memory (230), user input / output unit (250), and network input / output unit (270). At this time, the embodiment of the present invention does not have to include all of the above components simultaneously, and depending on each embodiment, some of the above components may be omitted, or some or all of the above components may be selectively included.
[0049] The processor (210) can execute a procedure for performing a customized evaluation method for each language type according to the present invention, manage memory (230) that is read or written during this process, and schedule the synchronization time between volatile memory and non-volatile memory in memory (230). The processor (210) can control the overall operation of the evaluation device (130) and can control the data flow between the memory (230), user input / output unit (250), and network input / output unit (270) by being electrically connected to them. The processor (210) can be implemented as a CPU (Central Processing Unit), etc., of the evaluation device (130), but is not necessarily limited thereto.
[0050] The memory (230) may include an auxiliary storage device implemented as non-volatile memory such as an SSD (Solid State Disk) or HDD (Hard Disk Drive) and used to store all data required by the evaluation device (130), and may include a main memory implemented as volatile memory such as RAM (Random Access Memory). Additionally, the memory (230) may store a set of instructions for the execution of the present invention, and the instructions may be executed by an electrically connected processor (210) so that a customized evaluation method according to the language type according to the present invention may be performed.
[0051] The user input / output unit (250) includes an environment for receiving user input and an environment for outputting specific information to the user, and may include an input device including an adapter such as a touch pad, touch screen, virtual keyboard, or pointing device, and an output device including an adapter such as a monitor or touch screen. In one embodiment, the user input / output unit (250) may correspond to a computing device connected via remote access, and in such case, the evaluation device (130) may correspond to an independent node of the network to which the computing device is connected.
[0052] The network input / output unit (270) provides a communication environment for connecting to other devices through a network and may include an adapter for communication such as a LAN (Local Area Network), MAN (Metropolitan Area Network), WAN (Wide Area Network), and VAN (Value Added Network). Additionally, the network input / output unit (270) may be implemented to provide short-range communication functions such as WiFi and Bluetooth, or wireless communication functions of 4G or higher for wireless transmission of data.
[0054] FIGS. 3a and 3b are drawings illustrating the functional configuration of an evaluation device according to the present invention.
[0055] Referring to FIG. 3a, the evaluation device (130) may be implemented to include functional configurations for performing a customized evaluation method for each language type according to the present invention. Specifically, the evaluation device (130) may include a speech voice input unit (310), a phoneme analysis unit (330), an acoustic feature extraction unit (350), and an evaluation execution unit (370).
[0056] In one embodiment, the evaluation device (130) may be implemented to further include one or more control modules (not shown in FIG. 3) that control the overall operation and manage the control flow and data flow between each of the modules, and the detailed operation of the control module is omitted to the extent that it overlaps with the description below.
[0057] Furthermore, embodiments of the present invention are not required to include all of the above components simultaneously; depending on each embodiment, some of the components may be omitted, or some or all of the components may be selectively included. The operation of each component will be described in detail below.
[0058] The speech voice input unit (310) can perform the operation of receiving user speech corresponding to a reference text expressed in a specific language from the user terminal (110). Here, the reference text corresponds to text expressed in a specific language such as Korean or English, and the user's speech may correspond to actual voice data generated during the process of the user reading or pronouncing the reference text. For example, the reference text may include learning sentences or scripts presented through the screen of the user terminal (110).
[0059] Additionally, the speech voice input unit (310) can receive a user's voice signal through a voice input module, such as a microphone installed in the user terminal (110), and can convert the voice signal into a digital signal and store it. Additionally, the speech voice input unit (310) can selectively perform preprocessing operations such as noise removal, sampling, and frame splitting, and can store the preprocessed voice data in a database.
[0060] In one embodiment, the speech voice input unit (310) can perform the operation of receiving an autonomous speech voice freely spoken by the user when no correct answer text is presented. That is, the autonomous speech voice is distinguished from the speech voice spoken by the user when the correct answer text is presented, and may correspond to voice data autonomously spoken by the user when no correct answer text is provided.
[0061] In this case, the speech voice input unit (310) can set up an autonomous speech environment through a separate input mode or user interface for receiving autonomous speech voice. For example, the user may speak freely in the form of everyday conversation or express their opinion on a specific topic, and the speech voice input unit (310) can collect the user's natural speech voice in real time or continuously for a certain period of time. In one embodiment, the autonomous speech voice may be converted into text at the syllable or word level through an automatic speech recognition (ASR) model, and the converted text may be provided to a subsequent phoneme analysis or evaluation stage. At this time, the ASR model may be implemented based on Wav2Vec2, etc.
[0062] Meanwhile, since the user's self-speaking voice does not have a predetermined text standard like the correct text, in a subsequent step, the phoneme analysis unit (330) can infer speech phonemes directly based on the self-speaking voice and generate speech phoneme information including the time position, length, and speech probability for each phoneme. In addition, in the self-speaking-based voice evaluation, a non-aligned phoneme inference and error type classification model can be applied instead of forced alignment.
[0063] As a result, the speech voice input unit (310) is designed to receive both reading-based speech voice and autonomous speech voice, thereby enabling flexible evaluation based on speech type and effective support for customized language ability evaluation by language type in various user environments.
[0064] The phoneme analysis unit (330) can perform the operation of inputting the spoken voice into a phoneme analysis model to generate spoken phonemes inferred in phoneme units. Here, the phoneme analysis model may correspond to an artificial intelligence-based model that receives the spoken voice as input and outputs phoneme data converted into phoneme units through an inference step. The phoneme analysis model may be constructed by pre-training learning data of a specific language type, or it may be constructed and utilized independently for each language type. The phoneme analysis unit (330) can obtain phoneme information that can be used for pronunciation evaluation from the user's spoken voice based on the phoneme analysis model. For example, the phoneme information output by the phoneme analysis model may include phoneme sequences and phoneme-specific probability values regarding the spoken phonemes.
[0065] In one embodiment, the phoneme analysis unit (330) may perform the operation of generating speech phoneme information including phoneme information classified according to the phonological system of a specific language and identification information of the corresponding phonemes. That is, the phoneme analysis unit (330) may divide the user's spoken voice into a plurality of speech phonemes and generate phoneme feature information including the position, length, and utterance probability of each speech phoneme as speech phoneme information. In particular, the phoneme analysis unit (330) may apply phoneme division criteria that consider the phonological system of each language so as to effectively respond to various language types.
[0066] Here, the phonological system may correspond to the rules by which phonemes are organized in a specific language. For example, in the case of Korean, it may be expressed as an initial, medial, and final consonant structure, and in the case of English, it may be expressed as a consonant, vowel, and diphthong structure. That is, the phoneme analysis unit (330) can identify the language type from the user's spoken voice and can analyze the spoken voice by subdividing it into phoneme units by applying the phonological system of the identified language type.
[0067] Additionally, the phoneme analysis unit (330) may include a function to directly infer speech phonemes for autonomous speech voices freely uttered by the user without the correct answer text being presented. In this case, the phoneme analysis unit (330) may generate a phoneme sequence at the level of phonetic symbols based on the acoustic characteristics and patterns of the autonomous speech voice, and the phoneme sequence may be used as an evaluation standard without the correct answer text during the evaluation process by the evaluation execution unit (370).
[0068] In other words, in an autonomous speech environment, instead of performing forced alignment, phoneme-level speech data and evaluation data can be generated through autonomous speech-based phoneme inference, and evaluation actions such as accuracy, fluency, and error types can be performed using such data.
[0069] Specifically, the phoneme analysis unit (330) can analyze the spoken speech and sequentially generate spoken phonemes identified on the time axis, and can generate spoken phoneme information, time information, and reliability information (or probability value), etc., in which each spoken phoneme is expressed as a phonetic symbol, as spoken phoneme information. For example, the phoneme analysis unit (330) can generate a phoneme sequence in which spoken phonemes identified sequentially on the time axis are expressed as phonetic symbols as phoneme information. To this end, the phoneme analysis unit (330) can utilize a phoneme analysis model, and the phoneme analysis model can generate output data in the form of a sequence in which phonetic symbols for each phoneme and position information according to the phonological system are connected at the phoneme level.
[0070] Additionally, the phoneme analysis unit (330) can generate position information within a syllable or word of each utterance phoneme as metadata added to the phoneme information. The phoneme analysis unit (330) can analyze the spectrum and temporal characteristics of the utterance speech to estimate the start and end times of each utterance phoneme, and can generate phoneme feature information including the duration and likelihood score of each phoneme segment, and the relative position (index) of each utterance phoneme.
[0071] At this time, the positional information for each phoneme may vary depending on the phonological system of the language. For example, in the case of Korean, the position within a word can be classified into initial consonant, medial vowel, and final consonant, respectively. As such, the phoneme analysis unit (330) can generate phoneme unit data inferred from the user's spoken voice or autonomous speech as spoken phoneme information, and through this, can perform precise error diagnosis and pronunciation evaluation for each language type.
[0072] The acoustic feature extraction unit (350) can perform the operation of forcibly aligning the speech phonemes based on the correct text and then extracting one or more acoustic features for the speech voice. Here, the acoustic feature may correspond to characteristic information extracted from the speech voice and may be used as input data for a pronunciation evaluation model. For example, the acoustic feature may include pitch, intensity, formant, harmonic-to-noise ratio (HNR), duration per phoneme, spectral coefficient (e.g., MFCC, etc.), and voice onset time (VOT) per phoneme.
[0073] Pitch represents intonation information by reflecting the fundamental frequency of the utterance, intensity represents the energy magnitude or strength of the spoken voice, and formant represents the resonance frequency according to the oral structure during vowel utterance, which can be used to analyze phonological quality characteristics. Additionally, the Harmonic-Noise Ratio (HNR) represents the ratio of noise to harmonics within a speech signal and is used for voiced / unvoiced discrimination and noise analysis; phoneme duration measures the length of time each phoneme is uttered and is used to evaluate speech rate and fluency; spectral coefficients represent the frequency-time distribution characteristics of the spoken voice; and phoneme utterance interval (VOT) can represent the time from the occurrence of a plosive to utterance. Acoustic features extracted in this way can be used as input data for pronunciation evaluation models during the pronunciation evaluation process.
[0074] More specifically, the acoustic feature extraction unit (350) can generate a reference phoneme sequence for forced alignment based on the correct text. The acoustic feature extraction unit (350) can forcibly align the spoken speech and the correct text on the time axis based on the phoneme unit sequence, and can calculate start and end time stamp information and alignment information for each phoneme according to the forced alignment result. In addition, the acoustic feature extraction unit (350) can set an analysis section per phoneme or per frame based on the forced alignment result, and can selectively extract and store one or more acoustic features necessary for the purpose of pronunciation evaluation.
[0075] In one embodiment, the acoustic feature extraction unit (350) can extract acoustic features even for autonomous speech freely uttered by a user without presenting a correct answer text. In this case, the acoustic feature extraction unit (350) can set an analysis section based on the results of phoneme inference for the entire or section-by-section of the autonomous speech, and perform an operation to extract various acoustic features based on time sections in frame units or phoneme units even without a correct answer text.
[0076] For example, the acoustic feature extraction unit (350) can estimate the start and end times of each phoneme through frame-based segmentation, and can collect acoustic features such as pitch, duration, energy, formant, and MFCC for each segment so that they can be used in the subsequent evaluation process.
[0077] The evaluation execution unit (370) can perform the operation of generating an evaluation result for a spoken voice using a pronunciation evaluation model based on speech phonemes and acoustic features. Here, the pronunciation evaluation model may correspond to an artificial intelligence-based model that generates an evaluation result by analyzing the language ability of a user reading a correct text. The pronunciation evaluation model may be built based on various learning models such as DNN, Transformer, and XGBoost, and may be designed to generate an evaluation result that additionally reflects structured error data and exception conditions as needed.
[0078] In one embodiment, the pronunciation evaluation model may be designed to infer acoustic and phonological characteristics between speech content and phonemes through AI-based analysis, even for autonomous speech where no correct answer text exists, and to generate evaluation results through comparison with learned criteria. In this case, the evaluation process for autonomous speech may correspond to a process of collecting and evaluating speech naturally spoken by a user as is, and the entire process, including speech recognition, phoneme decomposition, feature extraction, and calculation of evaluation scores, may be performed without a correct answer text.
[0079] That is, the evaluation execution unit (370) can generate input data to be used for pronunciation evaluation according to the structure of the pronunciation evaluation model, provide the generated input data to the pronunciation evaluation model, and generate an evaluation result using the output data generated by the model. At this time, the input data may include speech phoneme information, acoustic feature information, error data and exception conditions, etc., and the output data may include evaluation scores, error information and exception processing results, etc.
[0080] In one embodiment, the evaluation performing unit (370) may generate evaluation data including a suitability score, pitch, utterance length, and pause interval corresponding to each utterance phoneme according to the forced alignment result of the utterance phoneme, and calculate an evaluation score for the utterance speech using a pronunciation evaluation model based on the evaluation data. This is explained in more detail in FIG. 7.
[0081] Additionally, the evaluation execution unit (370) can construct evaluation data based on phoneme information and feature data generated by the phoneme analysis unit (330) and the acoustic feature extraction unit (350) without a correct answer text for evaluating the autonomous speech voice, and can analyze accuracy, fluency, and prosody according to criteria calculated by the pronunciation evaluation model based on the autonomous speech learning data. In this case, the evaluation execution unit (370) can perform evaluation operations in an evaluation mode dedicated to autonomous speech, and error detection and exception handling can also be applied in an extended manner according to the autonomous speech scenario.
[0082] In one embodiment, the evaluation performing unit (370) can calculate at least one partial score and a total score regarding accuracy, completeness, fluency, and prosody for the spoken voice. This is explained in more detail in FIG. 7.
[0083] In one embodiment, the evaluation execution unit (370) may include functional configurations for generating evaluation results based on a pronunciation evaluation model. Referring to FIG. 3b, the evaluation execution unit (370) may be implemented by including a data processing module (371), a model building module (372), an error detection module (374), and an exception handling module (375).
[0084] Meanwhile, the evaluation execution unit (370) may optionally include additional functional configurations for independently processing various operations of the evaluation process. For example, the evaluation execution unit (370) may further include a sentence recognition module that classifies sentence type information (e.g., declarative sentences, interrogative sentences, etc.), and a description thereof is omitted.
[0085] The data processing module (371) can perform the operation of collecting or processing speech phoneme information and acoustic features to generate training data for building a model or evaluation data for evaluating pronunciation. The data processing module (371) can optionally add error detection results and exception conditions as needed, and can generate input data in a form optimized for the model.
[0086] For example, the data processing module (371) can generate evaluation data in the form of a vector or array for each phoneme unit sequence. In this case, the data processing module (371) can generate phoneme sequence data by dividing it according to the phonological system of the language (e.g., the structure of Korean initial, medial, and final consonants). In the case of FIGS. 6 and 7, the data processing module (371) can generate phoneme sequences (631, 632, 633) that are divided into three parts for each region of initial, medial, and final consonants according to the Korean phonological system.
[0087] Here, the first phoneme sequence (631) may include speech phonemes corresponding to the initial consonant of each word, the second phoneme sequence (632) may include speech phonemes corresponding to the medial consonant of each word, and the third phoneme sequence (633) may include speech phonemes corresponding to the final consonant of each word. Additionally, the data processing module (371) may generate an evaluation vector corresponding to each phoneme sequence. In this case, the evaluation data may be represented as a set of evaluation vectors in sequence units that have been decomposed into three parts.
[0088] Meanwhile, the data processing module (371) may generate phoneme unit vectors for each word using phoneme unit sequences that have been divided into three parts for each area of the initial consonant, medial vowel, and final consonant, and then integrate them to generate word unit evaluation vectors. For example, the data processing module (371) may generate phoneme unit vectors for each area of the initial consonant, medial vowel, and final consonant based on the first syllable 'ye' of the correct text, and then integrate them to generate a single word unit evaluation vector. Subsequently, the data processing module (371) may generate evaluation data corresponding to the entire utterance text by connecting the word unit evaluation vectors in chronological order.
[0089] Additionally, in the case of an autonomous speech-based evaluation, the data processing module (371) can perform the operation of clustering consecutive speech phonemes on the time axis into meaningful contextual units (e.g., words, phrases, etc.) without a correct answer text, and generating an evaluation vector for each unit. In this case, the data processing module (371) can generate evaluation data that reflects the flow, stress, intonation, etc. of the speech, and thereby enable effective evaluation of fluency, etc., even for unstructured speech.
[0090] The model building module (372) can perform operations to learn, save, and update a pronunciation evaluation model, and can obtain model data by accessing the model stored in the DB during the pronunciation evaluation process. To this end, the model building module (372) can operate in conjunction with the model DB (373). That is, the model DB (373) may correspond to an independent storage that stores various models used in the pronunciation evaluation process. The model building module (372) can perform the model learning and building process based on the learning dataset, and can store the models that have completed learning in the model DB (373). In addition, the model building module (372) can update the models stored in the model DB (373) through additional learning.
[0091] Additionally, the model building module (372) can independently build and operate not only a correct text-based model but also an autonomous speech-based model. In the case of an autonomous speech-based model, it can be designed to automatically generate evaluation criteria from an autonomous speech corpus or to learn average indicators by speech type and convert them into evaluation indicators.
[0092] In one embodiment, a pronunciation evaluation model may be trained and constructed based on a Dual Learning Strategy to effectively evaluate autonomous speech. In this case, the pronunciation evaluation model may be designed to perform a sub-task for phoneme recognition and a main task for pronunciation score prediction in parallel, and the model's performance may be enhanced through a complementary relationship between the two tasks. For example, the pronunciation evaluation model may be designed to learn a relationship such as, "A person who pronounces this phoneme like this receives this score."
[0093] More specifically, the pronunciation evaluation model can be designed to perform the role of an acoustic feature extractor by applying relatively low weights to the phoneme extraction task, and to make the pronunciation score prediction task the model's ultimate goal by applying relatively high weights. Consequently, the pronunciation evaluation model can be designed to enable balanced learning between the two tasks through weight balancing.
[0094] In addition, the pronunciation evaluation model can be designed with a structure in which the information flow between the two tasks is mutually complementary. For example, the pronunciation evaluation model may include a bottom-up information flow in which pronunciation score prediction is performed based on phoneme recognition results, and may be designed to share common acoustic features between the two tasks, and each task may be designed to prevent overfitting to each other.
[0095] As a result, the evaluation execution unit (370) can perform an evaluation of autonomous speech without correct text based on a pronunciation evaluation model, and can perform multi-scale evaluations such as proficiency, fluency and meaning transmission (comprehension), and can ensure the reliability and consistency of the evaluation results through mutual verification between scores, and can provide real-time evaluation and feedback on user speech through a lightweight inference structure based on Wav2Vec2.0.
[0096] The error detection module (374) can detect an error type corresponding to the time axis position of each utterance phoneme based on the result of forced alignment between the utterance phoneme and the correct phoneme, and can perform the operation of generating the location and type information of the detected error as structured error data. For example, the error detection module (374) can detect various error types on a time axis basis, such as omission of the initial consonant, substitution of the medial vowel, and insertion of the final consonant. The error detection module (374) can match the utterance phoneme on the time axis based on the correct phoneme extracted from the correct text and record whether it matches, the location of the mismatch, and spaces or omissions.
[0097] Additionally, the error detection module (374) can detect error types such as insertion, deletion, and substitution based on the start and end times of each phoneme in the forced alignment result, and can also perform detection operations on predefined error types such as word boundary mismatch. The error detection module (374) can generate the detected error information into a predefined data structure. For example, the error detection module (374) can generate error data expressed in a data structure such as (error occurrence time interval, correct phoneme, utterance phoneme, error type).
[0098] Meanwhile, error types may include mispronunciation, omission, insertion, unnecessary pauses, omitted pauses, and monotone. Mispronunciation is an error where the user mispronounces; omission is an error where the user omits a pronunciation even though it is present in the correct text; insertion is an error where it is detected in the user's spoken voice even though it is not in the correct text; unnecessary pause is an error where an inappropriate pause exists between words within the same sentence; omitted pause is an error where a pause is omitted when there are punctuation marks between words; and monotone is an error where words are pronounced in a flat and lifeless tone without intonation or emotional expression. Additionally, word boundary mismatch errors may refer to errors where the word boundaries of the correct text do not match the boundaries of the spoken phonemes.
[0099] Additionally, the error detection module (374) can detect errors based on abnormalities in phoneme patterns or speech patterns without the need for correct text when evaluating based on autonomous speech. For example, the error detection module (374) can detect phoneme stability errors, prosodic errors, and fluency defects from speech patterns such as repetitive speech, unnecessary pauses, or abrupt changes in intonation.
[0100] The exception handling module (375) can perform an operation to set an exception condition corresponding to an error type and apply it to the evaluation process of the spoken voice. The exception handling module (375) can analyze the structured error data generated by the error detection module (374) to set an exception condition corresponding to the error type, and can perform an operation to apply the set exception condition to the evaluation process of the spoken voice. That is, even if a specific phoneme is detected as an error type such as insertion, omission, or substitution, the exception handling module (375) can set an exception condition when it is necessary to process it as a normal speech pattern such as liaison, contraction, and nasalization, thereby preventing negative evaluation (e.g., deduction) from occurring during the score calculation stage of the pronunciation evaluation model.
[0101] For example, in the case where a user utters "Monmeogeosseo" in response to the correct text "I couldn't eat," the exception handling module (375) may set a normal speech exception condition to process it as a natural speech pattern based on linking and contraction, even if it corresponds to an error of mismatch between the word boundary of the correct text and the boundary of the speech phoneme. Subsequently, the exception condition set by the exception handling module (375) may be transmitted to the data processing module (371) along with the error data, and the data processing module (371) may generate evaluation data with the error data and exception condition added and provide it as input to the pronunciation evaluation model. The pronunciation evaluation model may be designed so that it is not reflected in negative evaluations, such as deductions, when it corresponds to the exception condition, and in this case, the exception condition may correspond to a feedback signal for the evaluation data.
[0102] Additionally, the exception handling module (375) can set variations in expressions within the utterance, differences in colloquial word order and intonation, etc., as exception conditions through predefined exception rules or a learning-based exception predictor in the case of an evaluation based on autonomous speech, thereby supporting the evaluation of phoneme variations that naturally occur in actual speech without deductions.
[0104] Figure 4 is a flowchart illustrating a customized evaluation method by language type according to the present invention.
[0105] Referring to FIG. 4, the evaluation device (130) can perform a customized evaluation method for each language type according to the present invention in steps. More specifically, the evaluation device (130) can receive a user’s spoken voice corresponding to a correct text expressed in a specific language or an autonomous spoken voice freely spoken without a correct text from a user terminal (110) through a spoken voice input unit (310) (step S410). At this time, the autonomous spoken voice may correspond to a natural spoken voice input during the process of the user freely expressing their thoughts or sentences without a presented correct text, and may be distinguished from a speech voice based on a correct text.
[0106] Additionally, the evaluation device (130) can input the spoken speech into a phoneme analysis model through the phoneme analysis unit (330) to generate spoken phonemes inferred at the phoneme level (step S430). In the case of a correct text-based system, the generated phonemes can be forcibly aligned by comparing them with the correct phoneme sequence. In the case of an autonomous speech-based system, the spoken phoneme sequence can be divided into semantic units within a sentence and then automatically clustered and aligned according to evaluation criteria.
[0107] Additionally, the evaluation device (130) can extract one or more acoustic features for the spoken voice after forcibly aligning the spoken phonemes based on the correct text through the acoustic feature extraction unit (350) (step S450). The evaluation device (130) can generate an evaluation result for the spoken voice using a pronunciation evaluation model based on the spoken phonemes and acoustic features through the evaluation execution unit (370) (step S470). In this case, the evaluation execution unit (370) can calculate a score based on suitability, length, and pause intervals for each phoneme through the correct text-based evaluation, and can calculate partial scores for each semantic unit sentence or interval, or calculate a comprehensive score based on the importance (weight) between intervals through the autonomous speech-based evaluation.
[0108] In one embodiment, the evaluation device (130) may further include an evaluation visualization unit that visualizes and displays the evaluation results. That is, the evaluation visualization unit may perform the operation of visually representing the partial score, total score, error type, and exception handling results calculated through the pronunciation evaluation model via a user interface (UI) (see FIG. 9a and 9b).
[0109] Specifically, the evaluation visualization unit can display partial scores and overall scores, such as accuracy, completeness, fluency, and prosody, calculated by the evaluation execution unit (370), through various visual elements such as bar graphs, pie charts, and gauges. The evaluation visualization unit can visually provide the location of errors on the speech phoneme sequence through color or highlighting based on structured error data generated by the error detection module (374). For example, sections where initial consonants are omitted may be displayed in red, sections where final consonants are inserted may be displayed in orange, etc.
[0110] Additionally, the evaluation visualization unit can display exception conditions set by the exception handling module (375) as dotted lines or separate colors, allowing the user to intuitively check normal speech patterns that are not penalized. The evaluation visualization unit can store the evaluation visualization results in a database and provide them to the user terminal (110) to support the user in intuitively understanding their pronunciation problems and utilizing them for future correction learning.
[0112] FIG. 5 is a diagram illustrating an embodiment of a phoneme inference process according to the present invention.
[0113] Referring to FIG. 5, the user can check the correct answer text (510) displayed in a specific language (e.g., Korean) and speak along with it. At this time, the user's spoken voice (520) can be input in the form of a voice file (.wav) through a microphone or user terminal (110). The evaluation device (130) can perform preprocessing operations such as sampling, noise removal, and frame splitting on the input spoken voice to convert it into input data for a phoneme analysis model (530). The evaluation device (130) can convert the correct answer text into a text file (.txt) or break it down into phonemes and, if necessary, generate phoneme sequence data according to a language-specific phonological system (e.g., Korean initial, medial, and final consonant structure).
[0114] Additionally, the evaluation device (130) can perform phoneme-unit inference using a phoneme analysis model (530) based on the preprocessed correct text (510) and the spoken voice (520). At this time, the phoneme analysis model (530) can be built in advance based on Transformer or Wav2Vec2, etc. The phoneme analysis model (530) can infer the most likely phoneme candidate for each time frame and can calculate and output a probability value for each phoneme. In particular, the phoneme analysis model (530) can output the phoneme inference result in the form of a spoken phoneme sequence (540), and the evaluation device (130) can generate spoken phoneme information including phonetic symbols corresponding to each phoneme, location information (e.g., initial, medial, final consonants, etc.), and time information.
[0115] For example, the speech phoneme sequence (540) output by the phoneme analysis model (530) in FIG. 5 can be expressed as 'YEH_m BB_i UH_m N_f S_i AH_m G_i YAH_m NG_f UH_m …'. In this case, each speech phoneme can be represented by a phonetic symbol, and the identification information connected after the phonetic symbol can represent positional information within a syllable. That is, 'i' can represent an initial consonant, 'm' a medial vowel, and 'f' a final consonant.
[0117] Figure 6 is a diagram illustrating the phoneme alignment process according to the present invention.
[0118] Referring to FIG. 6, the phoneme alignment process may correspond to a process of forcibly aligning a speech phoneme sequence (540) based on the answer text (610). To this end, the answer text (610) may be broken down into words according to the phonological system of each language and aligned along the time axis. For example, in FIG. 6, the answer text (610) is 'We watch a beautiful sunset together,' and the speech text (620) corresponding to the user's speech voice is 'We watch a beautiful sunset together.' In this case, the answer text (610) and the speech text (620) may be aligned in word order according to the Korean phonological system. If the answer text (610) is expressed in English, the answer text (610) may be aligned in alphabetical order according to the English phonological system.
[0119] In one embodiment, the evaluation device (130) can divide (or 3-dimensionally) the speech phoneme sequence (540) into three regions of initial, medial, and final consonants according to the Korean phonological system and then align the phonemes corresponding to each region, and can perform an operation of forcibly aligning the speech phonemes by matching them to each region based on the word position of the correct answer text (610) aligned on the time axis. That is, the evaluation device (130) can detect errors according to each language type more precisely by differentiating the alignment criteria by considering the phonological system of each language. For example, it may be designed to perform alignment based on the decomposition of initial, medial, and final consonants within a syllable in the case of Korean, and alignment based on alphabet sequences in the case of non-syllable languages such as English.
[0120] In this case, as shown in FIG. 6, the speech phonemes that have been divided into three parts for each area of the initial consonant, medial consonant, and final consonant can be arranged side by side according to word order based on the correct text (610). That is, the first phoneme sequence (631) corresponds to the set of speech phonemes corresponding to the initial consonant of each word in the speech phoneme sequence (540), the second phoneme sequence (632) corresponds to the set of speech phonemes corresponding to the medial consonant of each word in the speech phoneme sequence (540), and the third phoneme sequence (633) corresponds to the set of speech phonemes corresponding to the final consonant of each word in the speech phoneme sequence (540).
[0121] The evaluation device (130) can perform an operation to generate alignment information, start and end times, and error type information, etc., based on the correct answer matching for each speech phoneme using the forced alignment result. In particular, the evaluation device (130) can perform an operation to detect an error type for speech phonemes that failed to match the correct answer based on the correct answer text (610) of the forced alignment result. For example, if an initial consonant is omitted in a speech phoneme, it may correspond to an omission error, and if a final consonant not present in the correct answer text is added to the speech phoneme, it may correspond to an insertion error.
[0123] FIG. 7 is a drawing illustrating an example of an evaluation process according to the present invention.
[0124] Referring to FIG. 7, the evaluation device (130) can perform an operation to calculate a quantitative evaluation score by performing a pronunciation evaluation of the spoken speech based on the forced alignment results of the spoken phonemes, acoustic features, and error and exception information.
[0125] First, the evaluation device (130) can collect phoneme alignment results, acoustic features, error data, and exception conditions, and can generate evaluation data for pronunciation evaluation based thereon. Acoustic features may include pitch, intensity, resonance frequency, harmonic noise ratio, duration per phoneme, spectral coefficient, and speech interval per phoneme.
[0126] For example, in FIG. 7, the evaluation device (130) can generate evaluation data (710) as a result of structuring the GOP score, pitch, duration, pause, error, and exception into an evaluation vector at the phoneme or sequence level. Here, the GOP score is the pronunciation GOP score for each phoneme, the pitch is the average fundamental frequency of the corresponding phoneme interval, the duration is the duration of the corresponding phoneme, the pause is the length of the pause before and after the corresponding phoneme, the error is an error type code, and the exception is whether the exception condition is satisfied.
[0127] In addition, the sequence unit evaluation vector can be represented as a set of phoneme unit evaluation vectors converted into a sequence form, and may further include statistical values of phoneme unit data as needed. For example, the fitness score among the phoneme unit evaluation vectors can be calculated through the following mathematical formula 1.
[0128] [Mathematical Formula 1]
[0129]
[0131] Here, is the target phoneme, is the speech feature sequence corresponding to the corresponding phoneme segment, is the entire set of phonemes, is the probability of the target phoneme for the corresponding interval, is the phoneme with the highest probability in the corresponding interval. In addition, the average value of the phoneme-unit fitness score can be included in the sequence-unit evaluation vector and used for pronunciation evaluation.
[0132] The evaluation device (130) can provide evaluation data (710), which is sequentially structured for each speech phoneme, as input data to the pronunciation evaluation model (730), and the pronunciation evaluation model (730) can calculate and output a quantitative evaluation score (750) as an evaluation result. At this time, the evaluation score (750) may include at least one partial score and a total score regarding accuracy, completeness, fluency, and prosody.
[0134] FIG. 8 is a drawing illustrating an embodiment of a Korean pronunciation evaluation element and item according to the present invention.
[0135] Referring to FIG. 8, the evaluation device (130) can perform the operation of calculating a quantitative evaluation score by applying multiple evaluation elements and evaluation items, such as accuracy and fluency, in the process of analyzing the user's spoken voice and performing a pronunciation evaluation.
[0136] Specifically, the evaluation device (130) can calculate an accuracy score by evaluating the accuracy of segments and whether phonological changes are reflected through a pronunciation evaluation model based on speech phoneme information and acoustic features. Here, the accuracy of segments represents the accuracy of pronunciation of individual consonant and vowel phonemes, and the reflection of phonological changes may represent whether natural phonological changes such as linking, nasalization, and palatalization occur.
[0137] Additionally, the evaluation device (130) can evaluate fluency based on intonation, speed, pauses, and rhythm information of the spoken voice. Here, intonation indicates intonation patterns by sentence type and whether native intonation remains, speed indicates whether the speech speed is excessively fast or slow as a result of measuring the speech speed, pauses indicate pauses at appropriate locations within the sentence, and rhythm indicates the rhythm, stress, and flow of speech of the entire sentence.
[0138] The evaluation device (130) can calculate scores for each of the various evaluation elements and then combine them to calculate partial scores and a comprehensive score, such as accuracy, completeness, fluency, and rhythm, and, if necessary, can also generate a final evaluation score by reflecting error detection results and exception processing results together.
[0140] FIGS. 9a and 9b are drawings illustrating an example of an evaluation result visualization process according to the present invention.
[0141] Referring to FIG. 9a, the evaluation result for a target sentence or word (e.g., 'school') pronounced by the user can be visually displayed by comparing the correct text (910) with the user's actual pronunciation.
[0142] Specifically, the evaluation device (130) receives the user’s spoken voice corresponding to the correct text (910), analyzes the phoneme sequence of the utterance to form a spoken text (920), and can subdivide it into spoken phonemes (930) arranged according to pronunciation rules. In the case of FIG. 9a, the user can pronounce the word 'school', and the evaluation device (130) can analyze the user’s spoken voice and display the inferred spoken phonemes (930) arranged in phoneme units such as 'ha', 'a', 'g', 'kk', 'yo', and can quantitatively calculate and display the pronunciation accuracy for each spoken phoneme (930). That is, a score for the corresponding pronunciation (e.g., 86, 83, 61, 12, 94, etc.) can be displayed at the bottom of each spoken phoneme (930), and through this, the learner can clearly recognize which phonemes are weak in their pronunciation.
[0143] Additionally, the evaluation device (130) may calculate a score for the entire word or sentence and provide an evaluation score (940) in the form of a percentage (%) as shown in FIG. 9a. In this case, the evaluation score (940) may be calculated based on multiple evaluation items such as accuracy, fluency, and completeness, and may be stored in a database or provided through the screen of the user terminal (110).
[0144] In one embodiment, the evaluation device (130) can detect the type of pronunciation error (e.g., insertion, omission, and substitution, etc.) for each phoneme according to the evaluation score through the evaluation execution unit (370), and can display the phoneme by distinguishing colors according to the evaluation score or visualize them in the form of icons. For example, in FIG. 9a, low scores may be displayed in a red color scheme and high scores in a green color scheme.
[0145] In one embodiment, the evaluation device (130) may provide a list of words (950) in which similar errors frequently occur according to pronunciation rules. For example, the list of words (950) illustrated in FIG. 9a may include example words to which the rule “the ‘ㄱ’ following the final consonant ‘ㄱ’ must be pronounced as ‘ㄲ’” applies (i.e., pharmacy, Taegeukgi, seaweed soup, musical instrument, etc.). That is, the list of words (950) may be provided as corrective learning material corresponding to the user’s weak phoneme types as feedback on the user’s spoken voice. Meanwhile, the area marked as A may correspond to a UI configuration that allows selecting word-unit and sentence-unit speech evaluation.
[0146] Referring to FIG. 9b, the evaluation device (130) can produce a comprehensive evaluation result by analyzing acoustic features, whether pronunciation rules are violated, alignment results, etc., for an entire sentence pronounced by the user (e.g., 'I am a student learning Korean.'). To this end, the evaluation device (130) can display the correct text (910) and the spoken text (920) side by side to allow the user's spoken voice and the target sentence to be compared intuitively.
[0147] At this time, the utterance text (920) can be reconstructed based on utterance phonemes inferred through a phoneme analysis model for the user's utterance voice input from the user terminal (110), and may also include boundary information on the time axis by syllable or by word. Additionally, a solid line graph visually representing changes in the energy intensity or waveform of the utterance voice over time may be displayed at the top of the utterance text (920), and the user can intuitively check their utterance rhythm, intonation changes, and intensity patterns through the graph.
[0148] Additionally, the evaluation device (130) can visually display the pronunciation accuracy obtained from the evaluation execution unit (370) for each spoken word. For example, in FIG. 9b, each word is visualized with a color or highlighting effect according to the score range, so that 'Hangugeoreul' is displayed in a red color scheme to indicate relatively low pronunciation accuracy, and 'Jeoneun' and 'Baeuneuneun' are each displayed in a green color scheme to indicate relatively high pronunciation accuracy. Through this, the user can clearly recognize, on a word-by-word basis, which part of their pronunciation was inaccurate.
[0149] Additionally, the evaluation device (130) may display multiple evaluation scores (940) calculated by the evaluation execution unit (370) as bar-shaped or numerical visual elements. In FIG. 9b, they may be displayed by category such as 'accuracy', 'intonation', 'pause', and 'speed', and the numerical value for each category may be calculated as a score between 0 and 100 based on the learning results of the pronunciation evaluation model. For example, if the user's intonation matches well with the type of sentence (e.g., declarative sentence, interrogative sentence, etc.), the score for the 'intonation' category may appear high, and the 'pause' category may be calculated by quantitatively reflecting whether the pause between words was appropriate.
[0150] In one embodiment, the evaluation device (130) can identify and classify sentence type information, such as declarative sentences or interrogative sentences, when the correct answer text (910) is a sentence through the sentence recognition module of the evaluation execution unit (370). In this case, the evaluation device (130) can enable the learner to learn an intonation pattern suitable for the correct answer text (910) by displaying the sentence type information together with the evaluation score (940).
[0151] In one embodiment, the evaluation device (130) can display information on error types included in the user's spoken voice, such as 'missing words', 'inserted words', and 'non-standard intonation', through sentence-level visualization, and this function can be activated by the error detection module and exception processing module of the evaluation execution unit (370).
[0153] Although the present invention has been described above with reference to preferred embodiments, those skilled in the art will understand that various modifications and changes can be made to the invention without departing from the spirit and scope of the invention as described in the following claims. Explanation of the symbols
[0155] 100: Evaluation System 510: Answer Text 520: Speech Voice 530: Phoneme analysis model 540: Speech phoneme sequence 610: Answer Text 620: Utterance Text 631: 1st phoneme sequence 632: 2nd phoneme sequence 633: Third phoneme sequence 710: Evaluation data 730: Pronunciation evaluation model 750: Evaluation score 910: Answer Text 920: Utterance Text 930: Speech phoneme 940: Evaluation score 950: Word List
Claims
Claim 1 A speech voice input unit that receives a user’s speech voice corresponding to a correct answer text expressed in a specific language from a user terminal; a phoneme analysis unit that inputs the speech voice into a phoneme analysis model to generate speech phonemes inferred at the phoneme level; an acoustic feature extraction unit that forcibly aligns the speech phonemes based on the correct answer text and extracts one or more acoustic features for the speech voice; and an evaluation execution unit that generates an evaluation result for the speech voice using a pronunciation evaluation model based on the speech phonemes and the acoustic features; wherein the evaluation execution unit includes an error detection module that detects an error type corresponding to the position of each speech phoneme and generates the position and type information of the detected error as structured error data; and an exception processing module that, among the detected error types, sets an exception condition corresponding to the error type when it corresponds to a normal speech pattern including liaison, contraction, or phonetic variation rules, so that no deduction occurs in the score calculation stage of the pronunciation evaluation model. A language type-specific customized evaluation device characterized by including a data processing module that generates evaluation data including the error data generated by the error detection module and the exception condition set by the exception processing module, and provides this data as input data for the pronunciation evaluation model. Claim 2 A language-type customized evaluation device according to claim 1, wherein the speech voice input unit receives an autonomous speech voice freely spoken by the user when the correct answer text is not presented. Claim 3 A language type-specific customized evaluation device according to claim 1, wherein the phoneme analysis unit divides the user's spoken voice or autonomous speech voice into a plurality of spoken phonemes and generates phoneme feature information including the position, length, and utterance probability of each spoken phoneme as spoken phoneme information. Claim 4 A language type-specific customized evaluation device according to claim 1, wherein the evaluation performing unit generates evaluation data including a suitability score, pitch, utterance length, and pause interval corresponding to each utterance phoneme according to the correct answer text-based forced alignment result or the autonomous speech-based phoneme inference result, and calculates an evaluation score for the utterance speech using the pronunciation evaluation model based on the evaluation data. Claim 5 A language type-specific customized evaluation device, wherein, in paragraph 4, the evaluation performing unit calculates at least one partial score and a comprehensive score regarding accuracy, completeness, fluency, and prosody for the user’s spoken voice or autonomous spoken voice, respectively. Claim 6 delete Claim 7 delete Claim 8 A method for performing a customized evaluation by language type based on a user's speech voice, which is performed in a customized evaluation device by language type, comprising: receiving a user's speech voice corresponding to a correct answer text expressed in a specific language from a user terminal through a speech voice input unit; inputting the speech voice into a phoneme analysis model through a phoneme analysis unit to generate speech phonemes inferred at the phoneme level; forcibly aligning the speech phonemes based on the correct answer text through an acoustic feature extraction unit and then extracting one or more acoustic features for the speech voice; and generating an evaluation result for the speech voice using a pronunciation evaluation model based on the speech phonemes and the acoustic features through an evaluation execution unit; wherein the step of generating the evaluation result comprises: detecting an error type corresponding to the location of each speech phoneme through an error detection module and generating the location and type information of the detected error as structured error data. A method for customized evaluation by language type, characterized by comprising: a step of, through an exception handling module, setting an exception condition corresponding to the error type when the detected error type corresponds to a normal speech pattern including liaison, contraction, or phonetic variation rules, so as not to deduct points in the score calculation stage of the pronunciation evaluation model; and a step of, through a data processing module, generating evaluation data including the error data generated by the error detection module and the exception condition set by the exception handling module, and providing it as input data for the pronunciation evaluation model.