Method and system for language learning pronunciation evaluation based on error-preserving speech recognition and non-linear confidence analysis via exclusion of context-based automatic correction
Patent Information
- Application Number
- KR1020260154000
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2026-03-06
- Filing Date
- 2026-08-18
- Publication Date
- 2026-09-04
Smart Images

Figure PAT00021_ABST
Abstract
Description
Technology Field
[0001] The embodiments of the present disclosure relate to a technology for evaluating pronunciation by analyzing the spoken voice of a language learner. Specifically, the technology relates to a pronunciation evaluation technology that fundamentally blocks the phenomenon of automatic correction for a learner's speech errors that may occur due to a language model during the speech recognition process, thereby preserving the errors as they are, and derives precise phonological-level evaluation results by setting a variable decoding path depending on the presence or absence of a script. Background Technology
[0002] With the recent advancement of artificial intelligence technology, non-face-to-face pronunciation evaluation services are being widely utilized in the field of foreign language learning. These systems typically employ a method based on Speech-to-Text (STT) engines to quantify a learner's spoken voice and calculate similarity by comparing it to a correct answer script.
[0003] A typical speech recognition engine is usually composed of a combination of an acoustic model and a language model. Here, the language model learns from a large-scale corpus to calculate the statistical probability of connections between words, and performs the role of deriving the result as a standard sentence with the highest probability of occurrence based on the surrounding context, even if the input speech signal is somewhat unclear. This structure serves as a key factor in maximizing recognition rates in speech recognition for the purpose of everyday information transmission.
[0004] However, in pronunciation evaluation environments for language learning, the intervention of such language models results in the system smoothly correcting and outputting standard Korean forms of phonological errors actually committed by the learner before the system even recognizes them. Even if a learner pronounces specific phonemes unclearly or speaks in violation of Korean's unique phonological rules, powerful language models restore these into grammatically perfect sentences, creating a discrepancy between the actual pronunciation and the recognized output. This is a technical characteristic that runs counter to the educational objective of precisely diagnosing learners' weaknesses.
[0005] Furthermore, evaluation in a free-talking environment without a pre-arranged script requires a significantly higher level of technical difficulty. This is because, while a clear standard for the correct answer exists when a script is present, in free speech, the output recognized by the system itself must serve as the pseudo-label—the baseline for evaluation. In this context, a reliability calculation process is essential; this process ensures that the speech recognition system converts the heard sound exactly as it is—including learner errors—into text, while simultaneously proving that the converted text does not deviate significantly from the context intended by the learner.
[0006] Existing reliability calculation methods tend to rely solely on either the scores of acoustic models or the probability values of language models, which carries the potential for misjudging reliability itself in cases of severe background noise or extremely inaccurate speaker pronunciation. In particular, if distortions in the physical acoustic environment or prediction uncertainties within the acoustic model are not comprehensively considered, the system evaluates the learner's pronunciation based on incorrectly generated pseudo-correct answers, ultimately undermining the reliability of the evaluation data.
[0007] Therefore, in the field of language education, there is a need for an advanced recognition and evaluation system that precisely verifies the validity of pseudo-correct answers by non-linearly combining various physical and statistical variables extracted from free speech environments, along with technology that preserves speech errors as they are by strategically controlling arbitrary correction mechanisms based on language models. The problem to be solved
[0008] The embodiments of the present disclosure can provide a pronunciation evaluation system and method for language learning that can recognize phonological errors committed by a language learner by reflecting them directly in the text without correction, by strategically excluding a context-based automatic correction function by a language model during the process of recognizing the utterance of a language learner.
[0009] In addition, embodiments of the present disclosure can provide a system and method capable of precisely determining the validity of a pseudo-label serving as an evaluation criterion based on a reliability index calculated by non-linearly combining the prediction sharpness of an acoustic model, the contextual likelihood of a language model, and the degree of physical signal distortion in a scriptless free speech environment.
[0010] In addition, embodiments of the present disclosure can provide a system and method capable of securing a high-precision correct answer baseline even in free speech situations by variably switching the decoding path according to the calculated confidence index and performing hypothesis cross-validation through a large language model (LLM) when the confidence is below a threshold value.
[0011] The embodiments of the present disclosure can provide a system and method capable of presenting an objective and accurate pronunciation correction guide to a learner by classifying the learner's detailed phonological error types by force alignment of the recognized text with the error preserved and the confirmed correct text at the consonant-character level, and processing and providing this as visualized feedback.
[0012] The technical problems to be solved in the embodiments are not limited to those mentioned above, and other unmentioned technical problems may be considered by those skilled in the art from the various embodiments described below. means of solving the problem
[0013] A method for evaluating a language learner's pronunciation according to one embodiment comprises: an operation of generating a preprocessed user voice signal by removing noise based on raw voice data received from a user terminal and a script presence indicator by a voice data receiving unit; an operation of decoding into text in which speech errors are preserved according to a probability distribution of consonant-character units extracted from the acoustic features of the preprocessed user voice signal and the script presence indicator by an error-preserving speech recognition unit, excluding context-based spelling correction operations by a language model (LM), and decoding according to the probability distribution of consonant-character units extracted from the acoustic features of the preprocessed user voice signal, thereby deriving a final ASR error text, a final alignment acoustic tensor, and a final comparison correct text by means of a speech text and acoustic alignment unit, respectively decomposing the final ASR error text and the final comparison correct text into consonant-character units, calculating the Levenshtein distance, and forcing alignment with the final alignment acoustic tensor based on the time axis to obtain a first-order text-acoustic merge map and time-axis aligned acoustic feature information. The operation includes: by the pronunciation error analysis unit, classifying phonological error types of specific patterns by combining the primary text-acoustic merge map and the time-axis aligned acoustic feature information, and calculating the Character Error Rate (CER) to generate a final pronunciation error analysis data set; by the customized feedback generation unit, generating final pronunciation evaluation and feedback data including user weaknesses and correction guides based on the final pronunciation error analysis data set; and by the evaluation result provision unit, processing the final pronunciation evaluation and feedback data into visualized UI components and transmitting them to the user terminal.
[0014] The operation of deriving the final ASR error text, the final acoustic tensor for alignment, and the final correct text for comparison comprises: an operation by an acoustic feature extraction layer to convert the raw waveform of the preprocessed user voice signal into the frequency domain to generate an initial acoustic feature tensor having time-series features; an operation by an encoder neural network layer to compress the spatiotemporal contextual information of the initial acoustic feature tensor to generate an acoustic encoding hidden vector, and to synchronize the initial acoustic feature tensor for transmission to a lower layer and output it as a transmission acoustic feature tensor; and an operation by a CTC decoder layer to calculate an independent initial character probability distribution for each time frame based on the acoustic encoding hidden vector without intervention of probability values from the language model. The operation may include performing a beam search within a vocabulary network of pre-loaded correct texts based on the initial character probability distribution and the acoustic feature tensor for transmission by means of a restricted search decoding layer, thereby determining the final ASR error text, the final alignment acoustic tensor, and the final comparison correct text that maps the highest probability path of the sound exactly as it is heard.
[0015] The operation of acquiring the above-described primary text-acoustic merging map and time-axis aligned acoustic feature information comprises: an operation in which, by means of a character unit decomposition layer, the final ASR error text and the final correct text for comparison are respectively parsed into minimum phonological sequences including initial consonants, medial vowels, and final consonants rather than morpheme units, and converted into a first character sequence and a second character sequence, and a delay synchronization of the final alignment acoustic tensor for time-axis computation; and an operation in which, by means of a Levenstein distance calculation layer, the edit distance between the first character sequence and the second character sequence is calculated using dynamic programming to generate a character unit node edit map containing accurate location information of nodes where insertion, deletion, and substitution occurred. By means of a phoneme-level forced alignment layer, a forced alignment is performed to map text nodes included in the character unit node editing map based on the timestamp frame of the delayed-synchronized acoustic tensor, thereby allowing the first text-acoustic merge map and the time-axis aligned acoustic feature information to be combined and output.
[0016] The operation of calculating the initial character probability distribution includes an operation of outputting a decoding state vector, a preserved acoustic feature tensor, and a third branch control signal along with the initial character probability distribution by the CTC decoder layer; after the operation of calculating the initial character probability distribution, the operation further includes an operation of interpreting the third branch control signal including the script presence indicator by the script presence primary branch determination unit to determine whether the evaluation is a reading evaluation with a pre-input correct answer text or a free speaking evaluation without a script, and switching the data transfer logic circuit to a first path or a second path; and the operation of determining the final ASR error text, the final alignment acoustic tensor, and the final comparison correct answer text includes, when the first branch determination unit determines that it is a reading evaluation and switches to the first path, performing the beam search within the lexical network of the pre-loaded correct answer text using the first path character probability distribution and the first path preserved acoustic tensor assigned to the first path by the restricted search decoding layer, and the first branch When the judgment unit determines that the free speech evaluation is the case and switches to the second path, the second path character probability distribution, the second path decoding state vector, and the second path preservation acoustic tensor are bypassed to the non-linear reliability evaluation second branch judgment unit by the first branch judgment unit, thereby triggering a subsequent branch operation to calculate the reliability of the pseudo-label text generated by the external model.
[0017] The operation of determining the final ASR error text, the acoustic tensor for final alignment, and the correct answer text for final comparison comprises: when the first branching judgment unit determines that the free speech evaluation is selected and the second path is switched, the operation of calculating a final reliability index for the pseudo-correct answer text by combining the acoustic sharpness per time frame derived from the second path character probability distribution, the contextual log-likelihood of the pseudo-correct answer text evaluated through an external model, and the degree of physical acoustic distortion extracted from the second path preservation acoustic tensor as scaling weights by the non-linear reliability evaluation second branching judgment unit; and when the calculated final reliability index is greater than or equal to a preset threshold by the high-reliability-based pseudo-correct answer confirmation layer, the operation of immediately determining the pseudo-correct answer text as the correct answer text serving as the reference point for evaluation and decoding the second path character probability distribution to derive a second type error text. An operation in which, by means of a multiple hypothesis cross-validation decoding layer, when the calculated final confidence index is less than the threshold, a plurality of candidate text hypotheses are generated from the second path character probability distribution, and for each of the generated candidate text hypotheses, the optimal hypothesis in which the conditional probability of a Korean sentence construction token sequence calculated by an external large language model (LLM) is maximized is selected as the correct text, and then the second path character probability distribution is re-decoded based on the selected correct text to derive a third type of error text;And through the pseudo-correct answer merging layer, the second type error text or the third type error text output from the high-reliability-based pseudo-correct answer confirmation layer or the multi-hypothesis cross-validation decoding layer, the confirmed correct answer text corresponding to each, and the preserved acoustic tensor can be merged into a single data format to form the utterance text and acoustic alignment unit.; Effects of the invention
[0018] According to the embodiments, the system can obtain precise phonological analysis data by blocking arbitrary grammar and spelling correction functions by the language model, thereby preserving phoneme sequences containing speech errors in their original form without distorting and recognizing the learner's incomplete pronunciation as standard language.
[0019] According to the embodiments, the system can technically prove the gap between the sentence intended by the learner and the actual utterance and enhance the objectivity of the evaluation by verifying the validity of evaluation criteria through a non-linear reliability analysis combining acoustic sharpness and contextual likelihood even in a free speech environment where no script exists.
[0020] According to the embodiments, the system can minimize misrecognition by the system even in complex speech situations and maximize the reliability of feedback provided to the learner by dynamically switching the decoding path according to the calculated reliability index and performing a cross-validation process using a large language model (LLM).
[0021] According to the embodiments, the system can provide a customized educational solution capable of practical pronunciation correction beyond simple score-based evaluation by visualizing and presenting specific phonological rules or articulation positions where the learner is weak through detailed forced alignment technology at the consonant-character level.
[0022] According to the embodiments, the system can build a highly efficient digital education infrastructure that can replace or assist professional evaluators in a language speaking assessment environment where a large number of people participate simultaneously, through an automated error-preserving recognition engine.
[0023] The effects obtainable from the embodiments are not limited to those mentioned above, and other unmentioned effects can be clearly derived and understood by a person skilled in the art based on the detailed description below. Brief explanation of the drawing
[0024] The accompanying drawings, included as part of the detailed description to aid in understanding the embodiments, provide various embodiments and explain the technical features of the various embodiments together with the detailed description. FIG. 1 is a diagram showing the configuration of an electronic device according to one embodiment. FIG. 2 is a diagram showing the configuration of a program according to one embodiment. FIG. 3 is an overall configuration diagram of a system for error-preserving language learning through context-based automatic correction exclusion according to an embodiment of the present invention. FIG. 4 is an overall flowchart of a method for evaluating language learner pronunciation including speech error preservation and non-linear reliability analysis according to one embodiment of the present invention. FIG. 5 is a structural diagram of a software layer performed in a system according to one embodiment of the present invention. FIG. 6 is a structural diagram of a neural network used in a system according to one embodiment of the present invention. Specific details for implementing the invention
[0025] The following embodiments are technologies developed through the Seoul Metropolitan Government Seoul Economic Promotion Agency (2025 Artificial Intelligence Technology Commercialization Support Project) (CY250204) (Development of AI Technology for Mathematical Voice Description and Educational Accessibility for Universal Lifelong Learning of Visually and Hearing Impaired People).
[0026] The following embodiments are combinations of the components and features of the embodiments in a predetermined form. Each component or feature may be considered optional unless otherwise explicitly stated. Each component or feature may be implemented in a form not combined with other components or features. Additionally, various embodiments may be constructed by combining some components and / or features. The order of operations described in various embodiments may be changed. Some components or features of one embodiment may be included in another embodiment, or may be replaced with corresponding components or features of another embodiment.
[0027] In the description of the drawings, procedures or steps that could obscure the essence of the various embodiments were not described, nor were procedures or steps that can be understood by a person of ordinary knowledge in the relevant technical field described.
[0028] Throughout the specification, when a part is described as "comprising" or "including" a component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, terms such as "...part," "...unit," and "module" as used in the specification refer to a unit that performs at least one function or operation, and this may be implemented in hardware, software, or a combination of hardware and software. Additionally, "one (a or an)," "one," "the," and similar related terms may be used in the context describing various embodiments (particularly in the context of the following claims) in both singular and plural forms, unless otherwise indicated in the specification or clearly contradicted by the context.
[0029] Hereinafter, embodiments according to various examples will be described in detail with reference to the accompanying drawings. The detailed description disclosed below, together with the accompanying drawings, is intended to describe exemplary embodiments of various examples and is not intended to represent the only embodiment.
[0030] In addition, specific terms used in various embodiments are provided to aid in understanding the various embodiments, and the use of such specific terms may be modified in other forms within the scope of not departing from the technical concept of the various embodiments.
[0031] FIG. 1 is a diagram showing the configuration of an electronic device according to one embodiment.
[0032] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or may communicate with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)). The electronic device (101) may be referred to as a client, terminal, or peer.
[0033] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), for example, and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0034] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model.
[0035] An artificial intelligence model can be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers. The artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, or a combination of two or more of the above, but is not limited to the examples described above. Artificial intelligence models may include software structures, either additionally or as a substitute, in addition to hardware structures.
[0036] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0037] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0038] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0039] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0040] The display module (160) can visually provide information to an external (e.g., user) outside the electronic device (101). The display module (160) can, for example, control a display, a holographic device, or a projector and the device.
[0041] It may include a control circuit for. According to one embodiment, the display module (160) may include a touch sensor set to detect a touch, or a pressure sensor set to measure the intensity of the force generated by the touch.
[0042] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0043] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a motion sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an infrared sensor, a biosensor, an acoustic sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0044] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0045] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0046] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0047] The camera module (180) can capture still images and images. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0048] The power management module (188) can manage the power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0049] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0050] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0051] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for realizing eMBB, loss coverage (e.g., 164 dB or less) for realizing mMTC, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for realizing URLLC.
[0052] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0053] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0054] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0055] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0056] The server (108) is connected to an electronic device (101) and can provide services to the connected electronic device (101). Additionally, the server (108) may proceed with the membership registration process, store and manage various information of users who have registered as members accordingly, and provide various purchase and payment functions related to the service. Furthermore, the server (108) may share execution data of service applications running on each of multiple electronic devices (101) in real time so that services can be shared among users. In terms of hardware, this server (108) may have the same configuration as a conventional web server or service server. However, in terms of software, it may include program modules that perform various functions and are implemented through any language such as C, C++, Java, Python, Golang, Kotlin, etc. Additionally, the server (108) generally refers to a computer system that is connected to an unspecified number of clients and / or other servers through an open computer network such as the Internet, receives requests for task execution from clients or other servers, and derives and provides the results of the task, as well as computer software (server program) installed for this purpose. Furthermore, the server (108) should be understood as a broad concept that includes, in addition to the aforementioned server program, a series of application programs running on the server (108) and, in some cases, various databases (DB: Database, hereinafter referred to as "DB") built internally or externally. Accordingly, the server (108) classifies membership registration information and various information and data regarding games, stores them in the DB, and manages them; such DB can be implemented internally or externally of the server (108).Additionally, the server (108) can be implemented using various server programs provided for general server hardware according to operating systems such as Windows, Linux, UNIX, and Macintosh. Representative examples include IIS (Internet Information Server) used in Windows environments and CERN, NCSA, APPACH, and TOMCAT used in UNIX environments to implement web services. Additionally, the server (108) may be linked with an authentication system and a payment system for user authentication of the service or purchase payment related to the service.
[0057] The first network (198) and the second network (199) refer to a connection structure capable of exchanging information between each node, such as terminals and servers, or a network connecting a server (108) and electronic devices (101, 104). The first network (198) and the second network (199) include, but are not limited to, the Internet, LAN (Local Area Network), Wireless LAN (Wireless Local Area Network), WAN (Wide Area Network), PAN (Personal Area Network), 3G, 4G, LTE, 5G, Wi-Fi, etc. The first network (198) and the second network (199) may be closed types such as LAN, WAN, etc., but it is preferable that they be open types such as the Internet. The Internet refers to a global open computer first network (198) and second network (199) structure that provides protocols such as TCP / IP protocol, TCP, and UDP (user datagram protocol), and various services existing in the upper layer, namely HTTP (HyperText Transfer Protocol), Telnet, FTP (File Transfer Protocol), DNS (Domain Name System), SMTP (Simple Mail Transfer Protocol), SNMP (Simple Network Management Protocol), NFS (Network File Service), and NIS (Network Information Service).
[0058] A database may have a general data structure implemented in the storage space (hard disk or memory) of a computer system using a database management program (DBMS). A database may have a data storage form that allows for the free retrieval (extraction), deletion, editing, and addition of data. A database may be implemented to suit the purpose of an embodiment of the present disclosure using a relational database management system (RDBMS) such as Oracle, Informix, Sybase, and DB2, an object-oriented database management system (OODBMS) such as Gemston, Orion, and O2, and an XML native database such as Excelon, Tamino, and Sekaiju, and may have appropriate fields or elements to achieve its functions.
[0059] FIG. 2 is a diagram showing the configuration of a program according to one embodiment.
[0060] FIG. 2 is a block diagram (200) illustrating a program (140) according to various embodiments. According to one embodiment, the program (140) may include an operating system (142), middleware (144), or an application (146) executable on the operating system (142) for controlling one or more resources of an electronic device (101). The operating system (142) may include, for example, Android™, iOS™, Windows™, Symbian™, Tizen™, or Bada™. At least some of the programs (140) may be preloaded into the electronic device (101) at manufacturing time, for example, or downloaded or updated from an external electronic device (e.g., electronic device (102 or 104), or server (108)) when used by a user. All or part of the program (140) may include a neural network.
[0061] The operating system (142) can control the management (e.g., allocation or reclamation) of one or more system resources (e.g., processes, memory, or power) of the electronic device (101). The operating system (142) may additionally or substantially include one or more driver programs for driving other hardware devices of the electronic device (101), e.g., an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197).
[0062] Middleware (144) may provide various functions to an application (146) so that functions or information provided from one or more resources of an electronic device (101) can be used by the application (146). Middleware (144) may include, for example, an application manager (201), a window manager (203), a multimedia manager (205), a resource manager (207), a power manager (209), a database manager (211), a package manager (213), a connectivity manager (215), a notification manager (217), a location manager (219), a graphics manager (221), a security manager (223), a call manager (225), or a voice recognition manager (227).
[0063] The application manager (201) can, for example, manage the life cycle of the application (146). The window manager (203) can, for example, manage one or more GUI resources used on the screen. The multimedia manager (205) can, for example, identify one or more formats required for the playback of media files and perform encoding or decoding of the corresponding media files among the media files using a codec that matches the selected corresponding format. The resource manager (207) can, for example, manage the source code of the application (146) or the memory space of the memory (130). The power manager (209) can, for example, manage the capacity, temperature, or power of the battery (189) and, using the relevant information, determine or provide relevant information required for the operation of the electronic device (101). According to one embodiment, the power manager (209) can interact with the BIOS (basic input / output system) (not shown) of the electronic device (101).
[0064] The database manager (211) can, for example, create, search, or modify a database to be used by the application (146). The package manager (213) can, for example, manage the installation or update of the application distributed in the form of a package file. The connectivity manager (215) can, for example, manage a wireless or direct connection between the electronic device (101) and an external electronic device. The notification manager (217) can, for example, provide a function to notify the user of the occurrence of a specified event (e.g., an incoming call, a message, or an alarm). The location manager (219) can, for example, manage location information of the electronic device (101). The graphics manager (221) can, for example, manage one or more graphic effects or related user interfaces to be provided to the user.
[0065] The security manager (223) may, for example, provide system security or user authentication. The telephony manager (225) may, for example, manage voice call functions or video call functions provided by the electronic device (101). The voice recognition manager (227) may, for example, transmit user voice data to the server (108) and receive from the server (108) a command corresponding to a function to be performed on the electronic device (101) based on at least part of the voice data, or text data converted based on at least part of the voice data. According to one embodiment, the middleware (244) may dynamically delete some existing components or add new components. According to one embodiment, at least part of the middleware (144) may be included as part of the operating system (142) or implemented as separate software different from the operating system (142).
[0066] The application (146) may include, for example, a home (251), a dialer (253), an SMS / MMS (255), an IM (instant message) (257), a browser (259), a camera (261), an alarm (263), a contact (265), a voice recognition (267), an email (269), a calendar (271), a media player (273), an album (275), a watch (277), a health (279) (e.g., measuring biometric information such as exercise volume or blood sugar), or an environmental information (281) (e.g., measuring atmospheric pressure, humidity, or temperature information). According to one embodiment, the application (146) may further include an information exchange application (not shown) capable of supporting information exchange between the electronic device (101) and an external electronic device. The information exchange application may include, for example, a notification relay application configured to transmit information (e.g., a call, a message, or an alarm) designated to an external electronic device, or a device management application configured to manage the external electronic device. The notification relay application may transmit notification information corresponding to a designated event (e.g., receiving mail) generated in another application of the electronic device (101) (e.g., an email application (269)) to the external electronic device. Additionally or alternatively, the notification relay application may receive notification information from the external electronic device and provide it to the user of the electronic device (101).
[0067] A device management application can control the power (e.g., turn-on or turn-off) or function (e.g., brightness, resolution, or focus) of an external electronic device or a part of its components (e.g., a display module or camera module of the external electronic device) that communicates with the electronic device (101). The device management application can additionally or substantially support the installation, deletion, or updating of applications running on the external electronic device.
[0068] Throughout this specification, the terms neural network, neural network, and network function may be used interchangeably. A neural network may consist of a set of interconnected computational units, which may generally be referred to as "nodes." These "nodes" may also be referred to as "neurons." A neural network is composed of at least two nodes. The nodes (or neurons) constituting neural networks may be interconnected by one or more "links."
[0069] In a neural network, two or more nodes connected via links can form a relative relationship between an input node and an output node. The concepts of input and output nodes are relative; any node in an output node relationship with respect to one node may be in an input node relationship with respect to another node, and vice versa. As previously mentioned, the input node versus output node relationship can be generated based on links. One or more output nodes may be connected to a single input node via links, and vice versa.
[0070] In a relationship between an input node and an output node connected through a single link, the value of the output node can be determined based on data input to the input node. Here, the nodes interconnecting the input node and the output node may have weights. The weights may be variable and may be varied by a user or an algorithm to enable the neural network to perform the desired function. Here, the edges or links interconnecting the input node and the output node have weights that can be variably applied by a user or an algorithm to enable the neural network to perform the desired function. For example, if one or more input nodes are interconnected to a single output node by respective links, the output node may determine its value based on the values input to the input nodes connected to the output node and the weights set on the links corresponding to each input node.
[0071] As described above, a neural network is formed in which two or more nodes are interconnected through one or more links to create input-output node relationships within the network. The characteristics of a neural network can be determined by the number of nodes and links within the network, the relationships between the nodes and links, and the weight values assigned to each link. For example, if two neural networks exist with the same number of nodes and links but different weight values between the links, the two neural networks can be recognized as being different from each other.
[0073] FIG. 3 is an overall configuration diagram of a system for error-preserving language learning through context-based automatic correction exclusion according to an embodiment of the present invention.
[0074] The embodiments of the present disclosure relate to a technology for evaluating pronunciation by analyzing the spoken voice of a language learner. Specifically, the technology relates to a pronunciation evaluation technology that fundamentally blocks the phenomenon of automatic correction for a learner's speech errors that may occur due to a language model during the speech recognition process, thereby preserving the errors as they are, and derives precise phonological-level evaluation results by setting a variable decoding path depending on the presence or absence of a script.
[0075] The embodiments of the present disclosure can provide a pronunciation evaluation system and method for language learning that can recognize phonological errors committed by a language learner by reflecting them directly in the text without correction, by strategically excluding a context-based automatic correction function by a language model during the process of recognizing the utterance of a language learner.
[0076] In addition, embodiments of the present disclosure can provide a system and method capable of precisely determining the validity of a pseudo-label serving as an evaluation criterion based on a reliability index calculated by non-linearly combining the prediction sharpness of an acoustic model, the contextual likelihood of a language model, and the degree of physical signal distortion in a scriptless free speech environment.
[0077] In addition, embodiments of the present disclosure can provide a system and method capable of securing a high-precision correct answer baseline even in free speech situations by variably switching the decoding path according to the calculated confidence index and performing hypothesis cross-validation through a large language model (LLM) when the confidence is below a threshold value.
[0078] The embodiments of the present disclosure can provide a system and method capable of presenting an objective and accurate pronunciation correction guide to a learner by classifying the learner's detailed phonological error types by force alignment of the recognized text with the error preserved and the confirmed correct text at the consonant-character level, and processing and providing this as visualized feedback.
[0079] According to the embodiments, the system can obtain precise phonological analysis data by blocking arbitrary grammar and spelling correction functions by the language model, thereby preserving phoneme sequences containing speech errors in their original form without distorting and recognizing the learner's incomplete pronunciation as standard language.
[0080] According to the embodiments, the system can technically prove the gap between the sentence intended by the learner and the actual utterance and enhance the objectivity of the evaluation by verifying the validity of evaluation criteria through a non-linear reliability analysis combining acoustic sharpness and contextual likelihood even in a free speech environment where no script exists.
[0081] According to the embodiments, the system can minimize misrecognition by the system even in complex speech situations and maximize the reliability of feedback provided to the learner by dynamically switching the decoding path according to the calculated reliability index and performing a cross-validation process using a large language model (LLM).
[0082] According to the embodiments, the system can provide a customized educational solution capable of practical pronunciation correction beyond simple score-based evaluation by visualizing and presenting specific phonological rules or articulation positions where the learner is weak through detailed forced alignment technology at the consonant-character level.
[0083] According to the embodiments, the system can build a highly efficient digital education infrastructure that can replace or assist professional evaluators in a language speaking assessment environment where a large number of people participate simultaneously, through an automated error-preserving recognition engine.
[0084] Referring to FIG. 3, a system according to one embodiment may include a voice data receiving unit (301), an error-preserving voice recognition unit (302), a speech text and acoustic alignment unit (303), a pronunciation error analysis unit (304), a customized feedback generation unit (305), and an evaluation result providing unit (306). The system is a physical and logical combination designed to generate educationally meaningful precision diagnostic data by fundamentally capturing subtle phonological variations and pronunciation errors appearing in a learner's voice without artificial intervention by a language model.
[0085] First, the voice data receiver (301) is connected to a user terminal via a communication network and receives a raw voice signal generated by a learner and a script presence indicator that indicates whether the voice is reading a predetermined script or is a free response. The receiver (301) is not merely a channel for receiving data, but processes the input analog-type digital voice signal into a form optimized for analysis. This includes sampling rate conversion, silent interval removal (VAD), and a noise canceling algorithm that suppresses ambient noise. In particular, the script presence indicator acts as a key parameter for determining the subsequent decoding path, defining whether the system will refer to a predefined text list or rely purely on the output of the acoustic model.
[0086] The error-preserving speech recognition unit (302) is a module that implements the core differentiation of the present invention. While a general STT engine automatically corrects grammatical errors by utilizing a language model (LM) to improve readability, this module strategically excludes such automatic correction functions. Instead, it extracts acoustic features, such as Mel-spectrograms, from a preprocessed speech signal and passes them through a neural network encoder to derive a probability distribution at the consonant-vowel level. This probability distribution is a probability value that corresponds purely to the 'sound' itself uttered by the learner, rather than a prediction based on context. Through this, when the learner pronounces "milk" not as "milk" but as an ambiguous sound similar to "milk" but mixed with a final consonant, it does not forcibly convert it into the standard word "milk" but decodes it into the consonant-vowel sequence exactly as it was heard to generate the final ASR error text. In addition, it preserves a high-dimensional acoustic tensor for precise time-axis analysis and derives the final correct text for comparison, which serves as the baseline for evaluation, and transmits it to a subsequent module.
[0087] The speech text and sound alignment unit (303) performs the role of physically combining text data and time-series sound data. The error text and correct text received from the recognition unit (302) are completely decomposed into consonants and vowels (initial, medial, and final consonants), which are the smallest phonological units beyond the morpheme unit. Subsequently, the edit distance between two consonant and vowel sequences is calculated through a Levenstein distance calculation layer. This calculation is performed based on dynamic programming and generates a precise node edit map indicating where the substitution, insertion, or deletion of consonants and vowels occurred. At the same time, phoneme-level forced alignment technology is applied to map, on a frame-by-frame basis, which time-stamp interval of the audio signal each consonant and vowel node is located in. As a result, the obtained primary text-sound merge map has a composite data structure in which text information and the actual physical speech signal are perfectly synchronized on the time axis.
[0088] The speech text and acoustic alignment unit receives the final ASR error text and the final correct text for comparison confirmed by the system, and can parse each text into initial, medial, and final consonant units through a character unit decomposition layer. Through this processing, the recognition result and the correct data can be generated as a first character sequence and a second character sequence, respectively, which can form a data basis for precisely capturing Korean-specific final consonant pronunciation errors or subtle phonological variation phenomena at the smallest unit.
[0089] In this process, the consonant-vowel unit decomposition layer decomposes syllable-unit data through a Unicode-based normalization process, reflecting the combinatorial characteristics of the Korean language. Going beyond simply splitting the text, it performs a process of standardization into individual phoneme sequences—the smallest phonetic units—so that they can be compared with acoustic tensors in subsequent layers. For example, when the words '사과' (apple) and '사콰' (sakwa) are input, the system establishes an environment where they can be separated into independent nodes for comparison. This reflects the characteristics of the Korean language, where the articulation positions of the initial consonants 'ㄱ' and 'ㅋ' in the second syllable are similar, yet their meaning changes depending on the presence or absence of aspiration. This enables the identification of subtle errors caused by learners failing to control the amount of air during pronunciation at the data level and provides an essential data preprocessing effect for analyzing the correlation between the probability distribution values output by the acoustic model and the text units. Consequently, consonant-vowel unit parsing functions as a medium combining deep learning-based acoustic feature extraction data with linguistic text data, serving as a technical foundation that improves evaluation resolution by more than three times from the syllable level to the consonant-vowel level.
[0090] In addition, the Levenstein distance calculation layer within the speech text and acoustic alignment unit can calculate the minimum edit distance between the first and second character sequences using dynamic programming. In this process, a character-unit node edit map can be generated by defining the state of each character node as one of insertion, deletion, substitution, or match; through this, it is possible to logically classify and structure the types of errors committed by the learner's pronunciation when compared to the correct answer.
[0091] The Levenstein distance calculation layer derives qualitative indicators capable of analyzing a learner's articulation patterns, going beyond simply calculating the distance between texts in the process of searching for the optimal path that minimizes the transformation cost between two sequences. Insertion refers to cases where a sound not present in the correct answer is unnecessarily uttered; deletion refers to cases where a specific phoneme is omitted without being pronounced; and substitution refers to cases where a specific phoneme is mispronounced as another. This node editing map serves as foundational data for statistically identifying which parts of the Korean initial, medial, or final consonants show weaknesses for the learner. In particular, by analyzing sections where substitution nodes occur intensively, it is possible to quantify confusion between consonants with similar articulation positions or pronunciation deviations within the vowel quadrant. This enables the AI to systematically classify learner errors in a manner similar to how human teachers diagnose them, thereby providing a logical basis for generating intelligent diagnostic reports that go beyond simple score calculation.
[0092] The phoneme-level forced alignment layer of the speech text and acoustic alignment section can operate to determine the speech segment corresponding to each node of the character-unit node editing map. Specifically, a first text-acoustic merge map can be generated by mapping the timestamp of the acoustic tensor for final alignment and the first character sequence on a frame-by-frame basis, and through this forced alignment processing, information in which the speech signal and text data are precisely synchronized with respect to the time axis can be obtained.
[0093] This layer utilizes the Viterbi algorithm or attention weights to identify the optimal temporal boundary between feature vector sequences, such as Mel-spectrograms extracted from speech signals, and recognized consonant-vowel sequences. To overcome the limitations of existing speech recognition systems, which provide recognition results only at the sentence or word level making it difficult to determine exactly when a learner committed an error, the start and end frames in which each consonant-vowel unit was uttered are specified in milliseconds (ms). The resulting primary text-acoustic merge map combines 'error nodes' in the text with 'time intervals' of the physical speech signal on a one-to-one basis. This provides physical coordinate information that can highlight the precise location of errors on the learner's voice waveform when generating visual feedback. Consequently, the forced alignment process synchronizes abstract text diagnosis with concrete physical signals, thereby exerting a key synchronization effect that enables learners to immediately perceive specific parts of their voice requiring correction, both auditorily and visually.
[0094] The pronunciation error analysis unit (304) diagnoses the learner's pronunciation from various angles based on the aligned data. It classifies types of phonological errors of specific patterns beyond simple matching by combining a primary merge map and time-axis aligned acoustic feature information. For example, it detects cases where the unique Korean linking rules or nasalization phenomena are not properly implemented, or analyzes whether the duration or energy distribution of specific phonemes deviates from the standard range. In addition, it quantitatively calculates the character error rate (CER) for the entire speech to quantify the learner's current level. The final pronunciation error analysis dataset generated at this stage includes deep metadata that goes beyond simple 'Right / Wrong' judgments and includes the learner's articulation habits and repetitive weaknesses.
[0095] The pronunciation error analysis unit can calculate the Scripted Error Rate (CER) of a learner's pronunciation relative to the correct consonant and vowel based on time-aligned acoustic feature information extracted from the primary text-acoustic merge map. At the same time, by intensively analyzing nodes identified as substitutions in the node editing map, it can classify specific patterns of phonological error types appearing in the learner's utterance; this can serve as a basis for quantifying pronunciation accuracy and systematically identifying the causes of errors.
[0096] The pronunciation error analysis unit analyzes the energy distribution of each phoneme segment, the trajectory of frequency formants, and the duration of articulation points from various angles based on aligned frame data. This process goes beyond simple text matching to re-verify the fluency and accuracy of actual pronunciation from an acoustic perspective. When calculating the Scripted Error Rate (CER), weights are assigned based on the total sentence length to differentiate between critical errors in short utterances and minor mistakes in long utterances. In particular, during substitution node analysis, phonetic segments of consonants and vowels identified as incorrect are compared with a database of standard speakers to quantify the degree of pronunciation distortion in decibels (dB) or frequency deviation units. This ensures that the evaluation indicators provided to learners are based on scientific acoustic analysis rather than mere speculation, thereby securing the credibility of the evaluation. Simultaneously, it offers a data mining effect that allows for tracking learners' chronic articulation habits based on a large volume of speech data.
[0097] Furthermore, the pronunciation error analysis unit can generate a final pronunciation error analysis dataset by determining whether one or more of the Korean phonological change rules—such as linking, nasalization, liquidization, gemination, and palatalization—have been violated. This analysis process enables a clear distinction, from a pedagogical perspective, between whether a learner simply mispronounced individual letters or failed to understand the unique Korean pronunciation rules resulting from the interaction between adjacent letters.
[0098] The analysis unit morphologically analyzes the consonant-vowel combination relationships within the text to prioritize the identification of specific environments where phonological changes are required. For instance, when the word 'gukmul' is input, the system recognizes in advance that nasalization occurs at the point where the final consonant 'ㄱ' meets the initial consonant 'ㅁ', resulting in the pronunciation [gungmul]. It then verifies whether nasal characteristics are captured in the learner's actual speech data through a primary text-acoustic merge map. If the learner ignores the rule and pronounces it as [gukmul], the system classifies this not as a simple initial / final consonant error, but as a high-level error known as "non-application of nasalization rules," and records it in the dataset. This enables precise measurement of the degree of internalization of phonological rules—the biggest barrier foreign learners face during the Korean language learning process—and realizes the effect of professional educational diagnosis by identifying deficiencies in linguistic knowledge beyond simple mispronunciation, allowing for customized prescriptions.
[0099] The customized feedback generation unit (305) reinterprets the analysis data set from a pedagogical perspective to construct a user-customized correction guide. Depending on the nature of the error committed by the learner, it selects or dynamically generates an appropriate feedback message. For example, if a misrecognition occurs due to insufficient airflow release when pronouncing a specific stop consonant, it matches a specific articulation position and methodological guide, such as "Try pronouncing it while bursting the air inside your mouth more strongly." In addition, it maximizes learning efficiency by selecting the weakest section with the lowest score among the analyzed data and constructing content that induces focused learning.
[0100] The customized feedback generation unit can search for timestamps of error points recorded in the final pronunciation error analysis dataset within the primary text-acoustic merge map. Based on the searched results, it can generate final pronunciation evaluation and feedback data that includes visualization information of the speech waveform of the corresponding section along with customized articulation guides corresponding to the types of phonological errors. Consequently, it can provide learners with the actual pronunciation error sections and specific guidelines for correcting them in a three-dimensional manner.
[0101] The customized feedback generator plays the role of converting all analyzed data into visual and linguistic content that is easy for users to understand. It displays the exact time intervals where errors occurred on the voice waveform using color-coding techniques, allowing learners to intuitively identify which parts of their pronunciation are problematic visually. Simultaneously, if the error type is identified as "failure to apply nasalization," it dynamically combines and generates articulation guides—such as "Nasalization did not occur. Place the back of your tongue against the roof of your mouth and try making the sound through your nose"—instead of simply displaying a "That is incorrect" message. This feedback data is transmitted to the user's terminal to form a real-time interactive learning screen, encouraging learners to listen to their own voices and compare them with the correct pronunciation segment by segment and practice repeatedly. This completes the final technological service effect, enabling AI to perfectly replace the professional counseling process of human teachers, allowing users to receive personalized, high-density language learning guidance anytime and anywhere.
[0102] Finally, the evaluation result providing unit (306) converts all processed evaluation and feedback data into visualized UI components that the learner can intuitively understand. It highlights error sections with color by comparing the correct sentence with the learner's speech content, or visualizes the accuracy of pronunciation in the form of a radar chart or graph. In particular, by utilizing forcibly aligned timestamp information, it provides an interactive learning environment where the learner can perform self-correction by selecting only their mispronunciation sections to listen to again or comparing waveforms with standard voice (TTS or native speaker voice). These UI components are packaged into structured data formats such as JSON or XML and transmitted to the user terminal, and are finally rendered on the learner's device screen to complete the real-time feedback loop.
[0103] Definitions of terms are described below.
[0104] Raw voice data is a digitized voice signal received from a user terminal before undergoing preprocessing. Essentially, it refers to time-domain waveform information (Raw Waveform) that includes microphone characteristics, ambient noise, and electrical distortion, and the data format generally has a Little-endian 16-bit PCM structure with a sampling rate of 16 kHz or 44.1 kHz.
[0105] The script presence indicator is flag information transmitted to distinguish whether the received voice data is a recitation of a specific predefined sentence or free-talking by the learner. It functions as a decoding path switching control signal that determines whether the system uses a fixed reference text or infers a pseudo-correct answer on its own, and the data format has an integer or Boolean value of 0 (no script) or 1 (script present).
[0106] Preprocessed user speech signals are signals refined from raw speech data into a state optimized for acoustic analysis through noise removal, DC offset correction, pre-emphasis, and normalization processes. They refer to time-series data with an enhanced signal-to-noise ratio (S / N ratio) that enables acoustic models to clearly identify phoneme boundaries.
[0107] Acoustic features are a set of feature vectors extracted from a preprocessed user voice signal through time and frequency domain analysis. It is data obtained by compressing and transforming the voice waveform into numerical features by mimicking the characteristics of human auditory filters, and it mainly takes the form of a Mel-spectrogram or MFCC (Mel-frequency Cepstral Coefficients) matrix (e.g., a sequence of 80-dimensional vectors).
[0108] The character unit probability distribution is a numerical set calculated from acoustic features passed through a neural network encoder to determine the probability of individual characters, such as initial consonants, medial vowels, and final consonants, appearing for each timestamp frame. It is tensor data of the softmax output layer that probabilistically represents which phoneme corresponds to the sound at a specific point in time, and the data format is a 2D floating-point matrix of [number of time frames x character dictionary size].
[0109] Text with preserved speech errors is text derived based on character-unit probability distributions extracted from acoustic features, excluding context-based spelling correction operations by a language model (LM). Essentially, it refers to a recognition result that transcribes the heard sounds exactly as they are, without forcibly substituting phonological errors committed by the learner into standard language. (e.g., when "사과" is pronounced as "사콰," standard STT outputs "사과," whereas the present invention outputs "사콰.")
[0110] The final ASR error text is the final output of the speech recognition unit and is a sequence of letters containing the learner's speech errors without correction. The final alignment acoustic tensor refers to high-dimensional multidimensional array data (Hidden States) that is preserved so as not to be lost during intermediate operations in order to precisely align the text and speech signals on the time axis.
[0111] The final correct answer text for comparison serves as the standard text for pronunciation evaluation and is a text confirmed as either the original script or an inferred pseudo-correct answer according to the script indicators. The sequence of data generated by completely decomposing this into individual letters is defined as the first letter sequence (ASR result) and the second letter sequence (correct answer standard), respectively.
[0112] The first text-acoustic merging map is a data structure obtained by mapping the timestamp frames of the acoustic tensor one-to-one with the consonant-character unit nodes. As a synchronization map that records "where (time) a specific character corresponds to a voice file," it takes the form of a list of JSON objects such as {"char": "ㄱ", "start_time": 1.25, "end_time": 1.35}.
[0113] Time-axis aligned acoustic feature information is a set of feature values physically synchronized with each character unit through a forced alignment process, and the character unit node edit map is an alignment map that records the state of each node by classifying the difference between two character sequences into insertion, deletion, and substitution. (e.g., [Index 3: Substitution, 'ㄱ' -> 'ㅋ'])
[0114] Specific patterns of phonological error types refer to pedagogically defined error categories, such as violations of Korean phonological rules (nasalization, liquidization, gemination, etc.) or discrepancies in place of articulation. The Character Error Rate (CER) is an indicator that quantifies the ratio of the number of erroneous characters, calculated via edit distance, to the number of correct characters.
[0115] The final pronunciation error analysis dataset is comprehensive diagnostic data containing error type and location information of individual phonemes derived by combining the aforementioned merged map and CER information, and the final pronunciation evaluation and feedback data is a combination of learner scores, weakness diagnosis results, and visual correction guides generated based on this.
[0116] Acoustic Sharpness is a measure indicating how highly the prediction probability distribution of an acoustic model is concentrated on specific consonants and vowels. Essentially, it signifies the inference confidence of the acoustic model, and the lower the entropy of the probability distribution, the higher the sharpness is considered to be.
[0117] A pseudo-label is a temporary reference text that is primarily inferred from the recognition results of an acoustic model in free speech situations where a correct answer script is absent, and is the subject of reliability verification. Contextual log-likelihood is a value that quantifies the probability of a recognized sentence appearing in the grammatical and statistical structure of the language on a log scale.
[0118] The degree of physical acoustic distortion refers to the level of degradation of the acoustic signal measured through the signal-to-noise ratio (SNR) of the input signal, reverberation, and whether clipping occurs, and the final reliability index is calculated through scaling weights and non-linear operations to combine these.
[0119] The multiple (N-best) hypothesis refers to a list of N recognition candidates with the highest probabilities during the decoding process. The conditional probability of a token sequence is a probabilistic combination value in which a Large Language Model (LLM) predicts the next word based on previous words, and this serves as the basis for selecting the optimal hypothesis in low-confidence situations.
[0120] The character-based decomposition layer is a layer that parses recognition results and correct text into sequences of initial, medial, and final consonants, which are the smallest phonological units of Korean, rather than into morpheme or syllable units. This enables the analysis of minute differences, such as final consonant pronunciation errors, at the atomic level.
[0121] The Levenstein distance calculation layer is an algorithm layer that calculates the minimum edit distance between the recognized character sequence and the correct character sequence using dynamic programming. It generates an edit map by precisely identifying the locations where insertions, deletions, and substitutions occurred.
[0122] The phoneme-level forced alignment layer is a synchronization layer that matches the timestamp frames of text units (characters) and acoustic tensors. It utilizes algorithms such as the Viterbi algorithm to determine the exact start and end intervals where each phoneme is spoken.
[0123] The non-linear reliability evaluation second branch judgment unit is a module that makes a final judgment on the validity of pseudo-correct answers generated in a free speech situation based on a reliability index calculated by non-linearly combining multidimensional indicators such as acoustic sharpness, contextual likelihood, and acoustic distortion.
[0124] The high-reliability-based pseudo-correct answer confirmation layer is a processing path that, if the calculated reliability index is above a preset threshold, immediately confirms the current recognition result as a reliable correct answer and proceeds to the subsequent evaluation process.
[0125] The multiple hypothesis cross-validation decoding layer operates when confidence is below a threshold and is an advanced layer that re-selects the optimal hypothesis with the highest conditional probability from the Large Language Model (LLM) as the correct answer among multiple recognition candidates (N-best hypotheses) derived through Beam Search, etc.
[0126] The pseudo-correct answer merging layer is an interface layer that integrates correct answer text, error text, and preserved acoustic data generated from different decoding paths (high confidence confirmation path or LLM cross-validation path) into a single data format and transmits it to the analysis unit.
[0128] FIG. 4 is an overall flowchart of a method for evaluating language learner pronunciation including speech error preservation and non-linear reliability analysis according to one embodiment of the present invention.
[0129] According to one embodiment, in operation (S401), the voice data receiving unit can generate a preprocessed user voice signal by removing noise based on raw voice data received from a user terminal and a script presence indicator. This process does not merely collect acoustic signals but also serves to improve the precision of subsequent neural network operations by removing various background noises that may be introduced from the learner's speech environment (e.g., cafe, classroom, outdoors, etc.). In particular, the script presence indicator serves as a criterion for dynamically adjusting the strength of the preprocessing filter depending on whether the voice data is reading a predetermined sentence or is an improvisational speech. For example, since hesitation or pause periods may be long during free speech, a digital signal optimized for analysis is formed by converting the sampling rate or performing normalization while maintaining signal integrity without forcibly cutting these periods.
[0130] According to one embodiment, in operation (S402), the error-preserving speech recognition unit excludes context-based spelling correction operations by a language model (LM) based on a preprocessed user voice signal and a script presence indicator, and decodes the text with preserved speech errors according to the probability distribution of grammatical units extracted from the acoustic features of the preprocessed user voice signal, thereby deriving the final ASR error text, the final alignment acoustic tensor, and the final correct text for comparison. While existing speech recognition technology focused on "how accurately it is restored to a standard sentence," the core of this embodiment lies in "how accurately the learner's pronunciation is converted into text without distortion." To this end, during the decoding process, the weight of the probability value of the language model affecting the result is set to near zero or completely removed, thereby listing the sound components recognized by the acoustic model (AM) exactly as they are in grammatical units, even if the pronunciation is grammatically incorrect. This fundamentally blocks the 'cognitive bias' phenomenon where the system automatically corrects a learner's pronunciation of 'school' to 'hagyo' based on context when the learner pronounces it as 'hagyo,' and consequently provides a physical foundation for converting learners' articulation errors into data.
[0131] According to one embodiment, an acoustic feature extraction layer can generate an initial acoustic feature tensor having time-series features by converting the raw waveform of a preprocessed user voice signal into the frequency domain. This layer performs physical operations to convert the analog signal in the time domain into multidimensional matrix data such as a Mel-Spectrogram or Mel-Frequency Cepstral Coefficients (MFCC). It non-linearly filters frequency bands by reflecting the characteristics of how the human auditory structure perceives sound, thereby extracting feature vectors that project the shape of the vocal tract or the movement of the vocal organs. The generated initial acoustic feature tensor is not simple audio data, but a high-dimensional feature map containing information on the learner's speech rate, pitch change, and formant, which is subsequently used as an input value for a deep learning model to enable precise phoneme identification.
[0132] According to one embodiment, an encoder neural network layer can generate an acoustic encoding hidden vector by compressing the spatiotemporal contextual information of an initial acoustic feature tensor, and output an acoustic feature tensor for transmission by synchronizing the initial acoustic feature tensor for transmission to a lower layer. The encoder may take a structure combining a convolutional neural network (CNN) and a transformer block, and captures abstract patterns hidden within the input acoustic data. The hidden vector generated here is a high-level representation for calculating the probability of individual phonemes appearing, in a state where only the core information of the acoustic signal is compressed. At the same time, the encoder manages the data transmission path in a dual manner so that 1:1 mapping can be performed without losing time-axis information during the later forced alignment process by synchronizing and transmitting the uncompressed original feature tensor to a lower decoder.
[0133] According to one embodiment, the CTC decoder layer can calculate an independent initial character probability distribution for each time frame based on acoustic encoding hidden vectors without intervention of probability values from the language model. The Connectionist Temporal Classification (CTC) technique solves the problem of different lengths between input and output sequences while having the characteristic of independently calculating the probability of character occurrence at each time step. In this embodiment, this characteristic is utilized in reverse to force the generation of characters indicated solely by the acoustic features of the corresponding time period, without being subject to contextual constraints from the previous or next word. Through this, 'abnormal phonological changes' or 'incomplete pronunciation' committed by Korean language learners are not diluted by the language model but are clearly recorded in the probability distribution, which serves as the most basic source data for evaluation.
[0134] The CTC decoder layer can output a decoding state vector, a preserved acoustic feature tensor, and a third branch control signal, along with the initial character probability distribution. These outputs are not merely simple text results, but rather a type of 'snapshot' data containing all the state information within the speech recognition system. The decoding state vector embodies the numerical confidence the model held when deriving the result, while the preserved acoustic feature tensor serves as a control group for comparison with the raw speech signal when generating a pronunciation correction guide later. In particular, the third branch control signal acts as a trigger determining the system's control flow, containing metadata information that determines whether the subsequent logic circuit sets the current data processing path for reading evaluation or for free speech.
[0135] The script presence / absence primary branching determination unit interprets the third branching control signal containing the script presence / absence indicator to determine whether the evaluation is a reading evaluation with pre-entered correct answer text or a free speaking evaluation without a script, and can switch the data transmission logic circuit to a first path or a second path. This step is a branching point responsible for the intelligent flexibility of the system. In the case of a reading evaluation (first path), since a predetermined correct answer script already exists within the system, data is sent to an algorithm that precisely compares the degree of agreement between the recognition result and the correct answer at the character level. On the other hand, in the case of free speaking (second path), since it is unknown what the correct answer is, data is transmitted to a reliability judgment module that first verifies how reliable the sentence just recognized by the system is. Through this physical switching, a multiplexed service structure is realized that can simultaneously perform structured tests and unstructured conversation tests within a single engine.
[0136] According to one embodiment, a restricted search decoding layer performs a beam search within a vocabulary network of pre-loaded correct texts based on an initial character probability distribution and an acoustic feature tensor for transmission, thereby determining the final ASR error text, the final alignment acoustic tensor, and the final comparison correct text that map the highest probability path of the sound exactly as heard. Unlike a general beam search, this layer, operating in a reading evaluation mode, limits the search range to the 'script that the user must read.' However, instead of forcing the search to only find words in the script, it includes in the vocabulary network the number of possible errors that the learner may commit based on the phonological structure of the script. Through this, the system simultaneously compares the probability that the learner read the script accurately with the probability that a specific phoneme was mispronounced, and tracks the path with the highest cumulative probability. The resulting error text is a textual representation of the learner's actual pronunciation, and the alignment acoustic tensor precisely indicates at what point in time that pronunciation occurred.
[0137] When the first branching decision unit determines that the reading evaluation is the case and switches to the first path, the restricted search decoding layer can perform the beam search within the lexical network of the pre-loaded correct text using the first path consonant-vowel probability distribution and the first path preservation acoustic tensor assigned to the first path. The data entering the first path is strongly coupled with predefined text resources. The system aligns the consonant-vowel probabilities extracted from the learner's voice with the correct sequence of the script and captures any 'discrepancy' that occurs during this process. For example, if the script is 'apple' but the learner pronounces it as 'sago', the beam search algorithm matches up to 'sa' with high probability, but then discovers a point in the last syllable where the acoustic feature corresponding to 'o' is overwhelmingly higher than 'wa', and confirms this as the final error text. Since this is a comparison within a defined framework, it guarantees very high computational speed and accuracy and serves as the basis for calculating the pronunciation score immediately.
[0138] When the first branching decision unit determines that the free speech evaluation is being conducted and switches to the second path, the first branching decision unit bypasses the second path character probability distribution, the second path decoding state vector, and the second path preservation acoustic tensor to the non-linear confidence evaluation second branching decision unit, thereby triggering a subsequent branching operation that computes the confidence of the pseudo-label text generated by an external model. In a free speech environment, since there is no fixed correct answer, the system's recognition result itself becomes the pseudo-label, which serves as the baseline for evaluation. If the system misrecognizes the learner's mumbling or ambient noise and generates an incorrect pseudo-label, all subsequent pronunciation evaluations become meaningless. Therefore, the system does not start the evaluation immediately but bypasses the bundle of data just recognized to the non-linear confidence judgment module, which acts as a 'verification waiting room.' Here, 'triggering' implies not merely sending data, but activating a precise verification computation process based on the physical feature values of the data.
[0139] According to one embodiment, a non-linear reliability evaluation secondary branch decision unit can calculate a final reliability index for a pseudo-correct text by combining the acoustic sharpness (pronunciation probability per time frame derived from the second path character probability distribution), the contextual log-likelihood of the pseudo-correct text evaluated through an external model, and the degree of physical acoustic distortion extracted from the second path-preserving acoustic tensor as scaling weights. The reliability index is not simply a single indicator, but an advanced numerical value that non-linearly combines three variables of different natures. Sharpness refers to how 'confidently' the acoustic model selected a specific phoneme, and contextual likelihood indicates how 'natural' the sentence is linguistically. By multiplying the degree of distortion, such as microphone noise or reverberation, by weights, a higher score is assigned as the text is clear to the machine, linguistically valid, and the physical environment is clean. Since these three indicators are calculated in combination, it ensures fairness in evaluation by precisely distinguishing between cases where the pronunciation is clear but the context is poor, or cases where the context is good but the pronunciation is slurred.
[0140] For example, the second branch of the non-linear reliability evaluation judgment unit can calculate the final reliability index for the pseudo-correct text using the following mathematical formula.
[0141] [Mathematical Formula]
[0142]
[0143] The strict meaning, definition, function, and system configuration method of each identification factor constituting this formula are as follows.
[0144] first Acoustic Sharpness is a metric that quantifies how concentrated the probability distribution is on a single candidate when the acoustic model's CTC decoder predicts a specific phoneme. If a speaker pronounces clearly, the probability for a specific consonant or vowel concentrates sharply, resulting in a peaked distribution; conversely, if pronunciation is unclear or mumbled, the probability is dispersed across multiple consonants or vowels, resulting in a flatter distribution. Therefore, sharpness directly reflects pronunciation clarity and acts as a positive (+) contribution term that increases final confidence. This value is calculated for each time frame output from the CTC layer. Softmax probability distribution Calculate Shannon entropy for, and the total number of frames It is calculated as the reciprocal after averaging. The formula is and, here The lower the entropy, the greater the sharpness, and the higher the entropy, the lower the sharpness.
[0145] Next Normalized Contextual Log-Likelihood (NLOGL) is a language model-based metric that measures how well acoustically generated pseudo-correct text conforms to the syntactic and semantic rules of the Korean language. It is introduced to prevent the hallucination phenomenon where acoustic models generate incorrect sentences based on clear signals. This value represents the text sequence of the language model on an external universal recognition server The sum of the log probabilities calculated for is the number of tokens It is a normalized value obtained by dividing to remove length bias. The formula is The more contextually natural a sentence is, the higher its log probability value becomes, and it acts as a positive (+) contribution term in the final confidence calculation.
[0146] Thirdly The Non-linear Acoustic-Environmental Distortion Penalty is a penalty term that quantifies the degree of environmental distortion contained in the input speech signal. The value increases as signal degradation, such as background noise, reverberation, and microphone clipping, becomes more severe, contributing a negative (-) contribution to the final reliability. This value is calculated using an exponential decay function based on the HNR (Harmonical-to-Noise Ratio) or SNR (Signal-to-Noise Ratio) derived during the speech preprocessing stage. The formula is As HNR decreases, the distortion penalty increases, and reliability decreases exponentially.
[0147] coefficient term Dynamic scaling weights are adjustment parameters designed to integrate values of different dimensions—sharpness, contextuality, and distortion penalty—into a single scalar confidence value. These are not fixed constants but tensor values that are optimized by backpropagation during the learning process. Initially, they are set randomly using methods such as He initialization, and they self-update via the Adam optimizer to minimize the loss function for large-scale pronunciation evaluation data. These weights dynamically adjust the relative influence of each term based on environmental characteristics.
[0148] Stabilization constant The Stabilization Constant is a minimum value designed to prevent numerical instability that may occur when the denominator becomes zero as the average entropy converges to zero. This is determined during the system configuration phase. It is statically allocated as a very small floating-point scalar value such as.
[0149] finally The Non-linear Decay Rate Constant is a gradient control constant that determines how steeply the distortion penalty responds to changes in HNR. The higher the value, the more rapidly the penalty increases even with small noise changes. This value is set to the empirical optimal value derived through a grid search technique and is statically loaded into system memory.
[0150] Since the entire formula has a sigmoid function structure, the larger the internal linear combination value, the higher the final reliability It converges to 1 and converges to 0 as the internal value decreases. Consequently, sharpness and contextual fit act to increase confidence, while environment distortion acts to attenuate confidence, and these three terms are balanced and integrated by the learned weights.
[0151] According to one embodiment, if the calculated final reliability index is above a preset threshold, the high-reliability-based pseudo-correct answer confirmation layer can immediately confirm the pseudo-correct answer text as the correct answer text serving as the reference point for evaluation and derive a second-type error text by decoding the second-path consonant-vowel probability distribution. The fact that the final reliability index exceeds the threshold is a technical guarantee that the system has very clearly recognized the learner's utterance and that the result sufficiently reflects the learner's intention. In this case, the text is promoted to the correct answer without additional verification, and pronunciation analysis is immediately initiated based on it. Since the second-type error text derived here contains only detailed phonological variation information in high-reliability situations, it is utilized as data to precisely measure specific articulation accuracy rather than the learner's overall language proficiency.
[0152] According to one embodiment, when the calculated final confidence index is below a threshold, the multi-hypothesis cross-validation decoding layer generates multiple candidate text hypotheses from a second-path character probability distribution, selects the optimal hypothesis as the correct text for each generated candidate text hypothesis in which the conditional probability of a Korean sentence construction token sequence calculated by an external large language model (LLM) is maximized, and then derives a third-type error text by re-decoding the second-path character probability distribution based on the selected correct text. Low confidence indicates that the speech signal is unclear or that the learner's sentence construction is fragmented. At this time, the system does not reach a single conclusion but transmits multiple hypotheses, called the 'N-best list,' to the LLM. The LLM probabilistically calculates how likely each candidate sentence is to be realized in actual human language conventions and backtracks the most plausible 'true intention.' The optimal hypothesis selected in this way becomes a complete sentence that the learner 'presumed they wanted to say,' and by comparing the raw speech signal again based on this, Type 3 error data is extracted to clearly identify what the learner tried to pronounce but failed to do.
[0153] The threshold functions as a key numerical baseline that determines the subsequent decoding path by identifying the technical validity of the recognition results derived by the system. In particular, in a scriptless free speech environment, it serves as a branching point to quantitatively evaluate whether the pseudo-correct text primarily generated by the acoustic model matches the learner's actual speech intention, and this is typically set as a floating-point value between 0.0 and 1.0.
[0154] The threshold is compared in real-time with the final confidence index calculated by the second-order branching judgment unit of the non-linear confidence evaluation, and if the calculated index is greater than or equal to the preset threshold, the system considers the corresponding pseudo-correct answer as the definitive correct answer. Essentially, this implies technical confidence that the acoustic sharpness is high and the contextual likelihood is sufficient to serve as a reference point for evaluation without additional verification by an external model. In this case, the system reduces processing delay time by skipping the heavy computational process of LLM cross-validation and entering the immediate pronunciation analysis stage.
[0155] Conversely, if the final confidence index is calculated to be below the threshold, the system defines the recognition result as having high uncertainty and activates the multi-hypothesis cross-validation decoding layer. This serves as a safeguard to prevent misrecognition that may occur due to severe acoustic signal distortion or pronunciation ambiguity; it compels the re-selection of the optimal sequence with the highest linguistic probability by transmitting multiple recognition hypotheses derived through beam search to an external large language model. In other words, the threshold acts as a pivotal parameter for computational control to simultaneously achieve the two objectives of efficient system resource allocation and ensuring the accuracy of evaluation data.
[0156] From a practical data processing perspective, the threshold value may be operated as a fixed constant, but it can also be dynamically varied depending on the noise level of the learning environment or the speaker's proficiency. For example, in environments with heavy background noise, the degree of physical acoustic distortion increases, so the threshold value may be relatively lowered to increase the frequency of LLM correction intervention; conversely, in evaluation test environments requiring high precision, the threshold value may be set high to apply strict criteria for determining correct answers.
[0157] According to one embodiment, a pseudo-correct answer merging layer can merge a second-type error text or a third-type error text output from a high-reliability-based pseudo-correct answer confirmation layer or a multi-hypothesis cross-validation decoding layer, along with the confirmed correct answer text and preserved acoustic tensor corresponding to each, into a single data format and transmit it to a speech text and acoustic alignment unit. This layer is a terminal for integrating various forms of analysis data generated through different paths (high-reliability path or LLM cross-validation path) into a standardized format. This is because, regardless of which path was taken, the subsequent acoustic alignment unit must input three elements—'sound produced by the learner (error text)', 'sound that should have been produced (correct text)', and 'actual waveform of the sound (acoustic tensor)'—in a consistent structure. The merged data pack is passed to the final alignment stage with time information and character information precisely combined, thereby enabling the system to perform ultra-precise frame-by-frame analysis from the beginning to the end of the speech.
[0158] According to one embodiment, in operation (S403), the speech text and acoustic alignment unit can obtain a first-order text-acoustic merge map and time-axis aligned acoustic feature information by decomposing the final ASR error text and the final correct text for comparison into individual characters, calculating the Levenshtein distance, and forcing alignment with the acoustic tensor for final alignment based on the time axis.
[0159] This operation is a core alignment process that synchronizes discrepancies between recognized results and correct answers with time-series speech feature data. The system does not stop at simply comparing text-based ASR results with correct scripts as strings, but combines them with the actual time intervals of acoustic signals. The forced alignment technique used here utilizes Hidden Markov Models (HMMs) or recent neural network-based alignment algorithms to specify the start and end times of each phoneme's utterance at the frame level (e.g., 10ms). The primary text-acoustic merge map obtained through this process contains information on which segment of the speech data a specific text node was uttered in, while the time-aligned acoustic feature information maintains the Mel-spectrogram or MFCC feature vector of that segment in an aligned state. This serves as fundamental physical evidence for future analysis of not only pronunciation accuracy but also speech rate and the appropriateness of pauses.
[0160] According to one embodiment, in the operation of obtaining a primary text-acoustic merging map and time-axis aligned acoustic feature information, the character unit decomposition layer parses the final ASR error text and the final correct text for comparison into minimum phonological sequences including initial consonant, medial vowel, and final consonant, respectively, rather than morpheme units, converts them into a first character sequence and a second character sequence, and delay-synchronizes the acoustic tensor for final alignment for time-axis operation.
[0161] To ensure the precision of Korean pronunciation evaluation, the system performs decomposition at the grammatical level beyond the syllable unit. For example, by separating the syllable 'gak' into three independent units, 'ㄱ, ㅏ, ㄱ', it enables the capture of specific errors made by learners in the pronunciation of final consonants. This parsing at the smallest phonological unit is combined with a morphological analyzer to generate a pure sound sequence from which grammatical semantic information has been removed. Through this step, the first grammatical sequence (ASR result) and the second grammatical sequence (correct answer) acquire the same phase. Additionally, the process of delay-synchronizing the acoustic tensor refers to performing buffering and window shifting operations to compensate for the time delay (latency) incurred while passing through the deep learning encoder, so that the physical speech signal and the text sequence can be computed at exactly the same timestamp.
[0162] According to one embodiment, the Levenstein distance calculation layer calculates the edit distance between a first character sequence and a second character sequence using dynamic programming to generate a character-unit node edit map containing accurate location information of nodes where insertion, deletion, and substitution have occurred.
[0163] The Levenstein distance algorithm is applied to the consonant-vowel sequences to quantify the difference between two sequences. In this process, dynamic programming generates a two-dimensional matrix to calculate the minimum editing cost; this approach tracks not only the simple distance score but also the location where the error occurred through a backtracking process. Specifically, if a learner adds a sound not present in the correct answer, it is labeled 'Insertion'; if a sound that should have been present is omitted, it is labeled 'Deletion'; and if a sound is mispronounced, it is labeled 'Substitution'. The generated consonant-vowel unit node editing map has a data structure tagged with error states (normal, insertion, deletion, substitution) for each consonant-vowel index, serving as a landmark that allows the pronunciation analysis unit to logically determine which phonological rule violations have occurred.
[0164] According to one embodiment, a phoneme-level forced alignment layer performs forced alignment by mapping text nodes included in a character-unit node edit map based on timestamp frames of a delayed-synchronized acoustic tensor, and can output a combination of a first-order text-acoustic merge map and time-axis aligned acoustic feature information.
[0165] This layer is the final interface stage that combines text-based editing information with actual physical speech data. Each node derived from the character-unit node editing map is connected to a specific time frame interval of the acoustic tensor. For example, if a 'substitution' error occurs at a specific point, speech feature information for that interval is extracted and processed into a state where it can be compared and analyzed with feature values of a standard acoustic model. In this process, the Viterbi algorithm or an attention map is utilized to determine the optimal path between the text unit and the audio frame. The resulting primary merge map becomes a complex dataset in which 'time-character-error type-acoustic feature' is combined four-dimensionally, containing in-depth pronunciation evaluation information that a simple recognition engine cannot provide.
[0166] According to one embodiment, in operation (S404), the pronunciation error analysis unit may combine a primary text-acoustic merge map and time-axis aligned acoustic feature information to classify phonological error types of specific patterns and calculate a Character Error Rate (CER) to generate a final pronunciation error analysis dataset.
[0167] Pronunciation error analysis combines statistical and rule-based analysis based on merged data. The system uses a classification algorithm to determine whether errors committed by learners are simple mistakes or stem from a lack of understanding of specific phonological rules in Korean, such as linking, nasalization, and liquidization. For example, if 'gukmul' is pronounced [gukmul] instead of [gungmul], retaining the 'ㄱ', it is tagged as an error resulting from the non-application of the nasalization rule. Simultaneously, the CER (Acoustic Echo Reliability), representing the proportion of incorrect consonants and vowels within the total utterance, is calculated to quantify the learner's overall pronunciation accuracy. The final analysis dataset, which includes Acoustic Reliability scores at each error location and the pedagogical types of the errors, serves as a source for constructing detailed guides during the subsequent feedback generation phase.
[0168] According to one embodiment, in operation (S405), the customized feedback generation unit can generate final pronunciation evaluation and feedback data including the user's weaknesses and correction guides based on the final pronunciation error analysis data set.
[0169] Based on the analyzed dataset, it diagnoses the learner's pronunciation habits and generates intelligent guides. The feedback generation module does not merely provide information indicating that the pronunciation is "incorrect," but analyzes the repetitiveness of the errors to identify phoneme pairs that learners frequently confuse (e.g., 'ㅓ' and 'ㅗ', 'ㄱ' and 'ㄲ'). Additionally, utilizing timestamps obtained during the forced alignment process, it matches correction phrases with loop repetition data to enable learners to visually recognize which points in their pronunciation are incorrect. At this stage, by referencing large language models or predefined educational rule sets, it dynamically generates text guides regarding tongue position and articulation methods for correctly pronouncing the corresponding phoneme, thereby enhancing the educational utility of the feedback.
[0170] According to one embodiment, in operation (S406), the evaluation result providing unit can process the final pronunciation evaluation and feedback data into a visualized UI component and transmit it to a user terminal.
[0171] The generated feedback data is ultimately delivered to the learner through a user interface (UI). The evaluation result delivery unit converts the data stream into visual objects; for example, it highlights erroneous characters in red above correct sentences, creates pronunciation accuracy graphs, and configures text viewers aligned with voice waveforms. Additionally, it processes the response data into JSON format optimized for mobile app or web browser environments and transmits it to the user's terminal via a communication module. Through the provided interface, learners can listen to their own voices again to immediately identify the error sections pointed out by the system and achieve real-time pronunciation correction by performing repetitive learning according to the provided guides.
[0173] FIG. 5 is a structural diagram of a software layer performed in a system according to one embodiment of the present invention.
[0174] The voice data receiving unit (501) is a layer that collects raw voice data transmitted from a user terminal and a script presence indicator at the forefront of the system. This layer prevents data packet loss that may occur in a network environment and standardizes the sampling rate or bit rate of the received voice file to match the standard specifications of the analysis server. In particular, by receiving metadata such as a script presence indicator together, it provides basic information for determining whether the voice recognition strategy to be performed in a subsequent step is a narration type or a free speech type. In addition, it performs basic preprocessing to improve the recognition rate of the acoustic model by filtering out unnecessary low-frequency noise included in the input signal.
[0175] The error-preserving speech recognition unit (502) is a core engine layer that extracts the learner's actual pronunciation from the received speech signal without interference from the language model. Unlike a general speech recognition unit that automatically corrects "incorrect pronunciation" to "correct word" for the naturalness of the context, this layer performs decoding by relying solely on the probability distribution of the letter-character unit of the acoustic model. By doing so, it preserves the minute phonological errors committed by the learner in the text, thereby ensuring that the exact error points can be captured in the subsequent alignment and analysis stages. The data derived at this time includes the final ASR error text along with a high-dimensional acoustic tensor for precise analysis.
[0176] The effect achieved through this layer is that subtle phonological variations made by learners (e.g., final consonant deletion, vowel changes, etc.) do not disappear but remain intact in the final ASR error text. This plays a pivotal role in maximizing the reliability of assessments by fundamentally preventing distorted corrections toward standard language in educational systems where it is essential to accurately identify "what is wrong."
[0177] The non-linear reliability evaluation secondary branching judgment unit (503) is an intelligent judgment layer that verifies the stability of the system, particularly in a scriptless free speech mode. This layer calculates a final reliability index by combining acoustic sharpness, contextual log likelihood, and physical acoustic distortion degree as scaling weights. This is a process in which the system itself quantifies whether "the currently recognized result is qualified as the correct text." It determines whether the calculated index is above or below a threshold value and dynamically branches the data processing path, thereby preventing fatal errors in which misrecognized text becomes the standard for evaluation.
[0178] The benefit of introducing this layer is that it can technically filter out misrecognitions by the system that may occur during free speech situations. By pre-selecting cases with low reliability, it prevents fatal errors in evaluating a learner's pronunciation based on incorrect answers and dramatically improves the stability of the evaluation process.
[0179] The multiple hypothesis cross-validation decoding layer (504) is a layer that operates when the confidence index is below a threshold to resolve the uncertainty of recognition. Instead of a single recognition result, the system generates multiple (N-best) hypotheses derived through beam search and transmits them to an external large language model (LLM). The LLM calculates the conditional probability of the Korean sentence construction for each hypothesis and re-selects the most contextually valid optimal hypothesis as the correct text. This is a process that complements the limitations of the acoustic model with the intelligence of the language model, and finally determines the correct text and the text containing errors and transmits them to the next stage.
[0180] The benefit of this approach is that it can identify the most appropriate correct answer text likely intended by the learner by considering the contextual flow, even in cases of severe acoustic noise or extremely inaccurate pronunciation. This expands the system's universality and serves as an example of effectively integrating the computational power of large-scale language models into the process of determining correct answers for pronunciation evaluation.
[0181] The speech text and sound alignment unit (505) is a layer that physically synchronizes the confirmed text data with the time-series sound signal. Through a character unit decomposition function, the text is parsed into initial, medial, and final consonant units, and the editing history between the correct answer and the recognition result is mapped at the node level using the Levenstein distance algorithm. Additionally, by applying a forced alignment technique, the time stamp interval in which each character unit was spoken in the audio signal is determined at the frame level. After this process, a three-dimensional data map is completed regarding "which character the voice signal at a specific time was recognized as and what error occurred when compared with the correct answer."
[0182] The primary effect of this layer is that it combines qualitative 'incorrect answer' information with quantitative 'temporal and physical information.' Learners can obtain not only the result of which word is incorrect, but also physical evidence regarding which phonological rules were violated at which point in the audio, thereby improving the precision of visual feedback.
[0183] The pronunciation error analysis unit (506) is an analysis engine layer that precisely diagnoses and statistically analyzes the learner's pronunciation based on aligned data. It calculates the character error rate (CER) by conducting a full survey of text-acoustic merge maps and classifies patterns of violations of phonological rules unique to the Korean language or discrepancies in the articulation positions of specific consonants / vowels. Beyond simply determining whether they match, it extracts pedagogical causes from the data as to why the learner committed such errors and generates a final pronunciation error analysis dataset.
[0184] The benefit of this process is that it derives pedagogically meaningful analytical data beyond a mere listing of misrecognition results. By enabling the statistical identification of specific phonological phenomena where learners are primarily weak, it provides a systematic language education infrastructure based on causal analysis rather than simple repetitive learning.
[0185] The customized feedback generation unit (507) is a layer that converts the analyzed data set into an educational message that is easy for the learner to understand. This layer tracks the learner's repetitive error patterns to identify weaknesses that need to be corrected first, and matches them with articulation guides to correctly pronounce the corresponding phonemes. Rather than simply listing scores, it generates specific and action-oriented feedback text, such as "When pronouncing the final consonant 'ㄱ', press the back of your tongue closer to the roof of your mouth," thereby increasing the practical effectiveness of learning.
[0186] Finally, the evaluation result providing unit (508) is an output layer that processes all generated evaluation and feedback data into visualized UI components and transmits them to a user terminal. It includes a text viewer with highlight effects applied, a pronunciation accuracy chart, and a waveform comparison component with standard speech so that the learner can intuitively recognize their errors. The processed data is packaged in a JSON format or similar and transmitted to a user terminal via a network, and finally, a pronunciation evaluation session is concluded by visualizing it on the learner's app screen.
[0188] FIG. 6 is a structural diagram of a neural network used in a system according to one embodiment of the present invention.
[0189] The input signal feature extraction layer (601) receives the preprocessed speech signal at the bottom of the neural network and converts it into a high-dimensional numerical vector. This layer analyzes frequency components from the speech waveform, which is time-series data, to generate a feature map such as a Mel-spectrogram. The key is to extract multidimensional features that reflect the texture and articulation characteristics of the sound, rather than simply looking at the amplitude of the sound.
[0190] A raw speech signal in the time domain received from a user terminal is input, and energy in the frequency domain is calculated through frame-unit windowing of approximately 25ms and Fast Fourier Transform, and passed through a Mel filter bank. As a result, a Mel-spectrogram feature matrix, which is a floating-point tensor of size [T, 80], is output. This excludes speaker-specific variability and secures a low-dimensional feature space optimized for language recognition, thereby providing the effect of maximizing the computational efficiency of subsequent layers.
[0191] The primary effect of this layer is that it enables the neural network to exclude noise and speaker characteristics, such as gender and age, that are unnecessary for pronunciation evaluation from the learner's voice, while preserving only the core acoustic cues regarding which phonemes were purely pronounced. This lays the foundation for subsequent layers to perform precise computations based on stable input data.
[0192] The deep learning encoder layer (602) is a layer that receives a vector sequence generated by the feature extraction layer and encodes it into an abstract hidden state vector. This layer includes a transformer or conformer structure to capture contextual information before and after the utterance, while converting acoustic characteristics at each frame into highly concentrated data.
[0193] The benefit of this layer is the ability to clearly identify transition characteristics between phonemes by analyzing correlations between short speech frames. In particular, by capturing subtle acoustic changes at the points where Korean-specific assimilation rules or phonological shifts occur, it provides powerful potential expressive capabilities that allow for the determination of pronunciation accuracy in subsequent stages.
[0194] It receives an extracted feature matrix as input and performs processing to learn the time-series context dependencies of the entire utterance through a self-attention mechanism of a conformer or transformer structure. Through this, it outputs a high-dimensional hidden state vector sequence in which the temporal boundaries and articulatory characteristics of each phoneme are concentrated, and by internalizing transition rules between phonemes and contextual information beyond short-term acoustic features, it achieves the effect of providing abstract key clues capable of capturing subtle distortions in pronunciation.
[0195] The error-preserving CTC decoder layer (603) is a layer that directly calculates the probability distribution of individual characters without the intervention of a language model, based on the hidden vector derived from the encoder. Unlike existing recognition systems that automatically correct "incorrect sounds" into "correct words" by considering the context, this layer outputs the sound that the learner actually uttered as a character probability value based solely on acoustic grounds.
[0196] The decisive effect of this layer is its ability to transcribe learners' pronunciation errors into text without distortion. For example, when a learner omits a final consonant or mispronounces a vowel, it outputs the result as a text sequence containing the error rather than 'forcibly correcting' it to standard language, thereby enabling the securing of source data for 'analysis of the causes of errors' necessary in actual educational settings.
[0197] Here, processing is performed to calculate the probability distribution of consonants and vowels for each timestamp through a linear transformation based on the CTC loss function, without the intervention of an external language model. Finally, a log probability matrix of the form [T, Vocabulary Size] is output; this exerts a key effect in realizing an error preservation function that prevents the automatic correction of language models that enforce contextual naturalness, thereby preserving the phonological errors actually committed by the learner in the text.
[0198] The restricted search decoding layer (604) operates in a script-based evaluation mode (reading evaluation) and is a layer that performs a beam search by referencing the vocabulary range of the pre-loaded correct text. Instead of searching an infinite language space, this layer maximizes recognition accuracy by searching for the optimal path only within the consonant-vowel combinations of the correct sentence that the learner must read.
[0199] The effect of this layer is that it dramatically increases the reliability of recognition when a fixed script is available. It blocks the possibility of being recognized as a wrong word and enables precise one-on-one pronunciation comparison evaluation by allowing focus on how closely the learner's pronunciation matches or deviates from the standard pronunciation of the correct script.
[0200] It takes the probability distribution of a CTC decoder and a pre-loaded correct text as input, and performs processing to infer the optimal path by limiting the search range to the correct vocabulary network during the beam search process. By doing so, it outputs the optimal consonant-vowel sequence within the limited vocabulary, thereby minimizing misrecognition of incorrect words and establishing an environment where pronunciation deviations can be precisely compared with the correct answer.
[0201] The non-linear reliability index calculation layer (605) operates particularly importantly in free speech mode and quantifies the reliability of the recognition result by calculating the pronunciation probability sharpness (Acoustic Sharpness), etc., from the extracted acoustic features and probability distribution. It is a judgment layer that derives a single reliability index by combining multiple indicators.
[0202] The effect of this layer is that it allows for the verification of the validity of answer keys when the system is required to generate them autonomously in free-response situations. By preemptively blocking "false feedback" that arises from uncritically adopting unreliable recognition results as the standard for correct answers, it serves as a safeguard to secure the educational credibility of the entire system.
[0203] It receives inputs such as character probability distributions and hidden vectors, and performs a non-linear weighted combination of pronunciation probability sharpness, acoustic energy distortion, and contextual log-likelihood using a multilayer perceptron. It outputs a final confidence index, which is a single scalar value between 0.0 and 1.0. This quantitatively evaluates the quality of pseudo-correct answers generated by the system itself, preventing misdiagnosis of unreliable data and providing a safeguard effect to ensure the credibility of the system.
[0204] The multiple hypothesis generation and LLM linkage layer (606) is a layer that generates multiple candidate sentences (N-best hypotheses) for recognition results with low confidence and re-selects the optimal sentence by comparing them with an external large language model (LLM). It is responsible for an advanced decoding process that resolves the confusion experienced by the acoustic model through an external model with vast linguistic knowledge.
[0205] The effect of this layer is that it can derive the most contextually valid pseudo-label even in situations where the acoustic environment is poor or the learner's pronunciation is extremely ambiguous. This expands the universality of the system and results in raising the accuracy of free speech evaluation beyond the level of commercial services by integrating the latest artificial intelligence technology.
[0206] It receives probability distribution data as input, generates an N-best hypothesis with an expanded beam size, and performs processing to compare the conditional probabilities of sentences by contrasting this with an external large language model. By outputting text containing confirmed correct answers and errors through cross-validation, it achieves an advanced recognition effect that accurately identifies the learner's speech intent by supplementing acoustically ambiguous pronunciations with vast linguistic intelligence.
[0207] The text-acoustic forced alignment layer (607) is a layer that synchronizes the finalized text sequence and the preserved high-dimensional acoustic tensor on the time axis. It forms a one-to-one correspondence map between text and sound by mapping the location of each character unit to which timestamp (ms) of the speech data on a frame-by-frame basis.
[0208] The effect of this layer is that it combines physical temporal information regarding "when the mistake occurred" with information regarding "what went wrong." Through this, it provides learners with spatiotemporal synchronized data that can visually show them exactly which section of their voice waveform is incorrectly pronounced.
[0209] Processing is performed to specify frame-unit segments of each consonant and vowel unit using the Viterbi algorithm or attention map. A primary text-sound merge map with synchronized time information is output, which has the effect of generating base data that can analyze 'where and how' the error occurred by mapping error points in the text to actual physical speech segments in milliseconds.
[0210] The final evaluation data output layer (608) is the final stage of the neural network that synthesizes the computation results of all preceding layers and outputs a final analysis set including character error rate (CER), phonological error patterns, and aligned time information. This data is passed to a feedback generation module and processed into a guide in a form that the learner can understand.
[0211] The effect of this layer is that it standardizes complex internal neural network operations into educationally meaningful diagnostic data. Rather than simply producing numerical results from artificial intelligence, it completes the system's ultimate goal of learning support by structuring and providing the source materials for error analysis and feedback necessary for actual language learning.
[0212] It receives merge maps, character error rates, and phonological error pattern data as input and performs processing to package all computation results into a single format for transmission to the feedback generation unit. It outputs a structured final pronunciation error analysis dataset, which standardizes the numerical computations within the complex neural network into a source for human-understandable educational content, thereby enabling the generation of practical and concrete learning guides.
[0214] The embodiments described above may be implemented as hardware components, software components, and / or combinations of hardware and software components. For example, the devices, methods, and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.
[0215] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.
[0216] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software devices to perform the operation of the embodiment, and vice versa.
[0217] Although the embodiments have been described above with reference to the limited drawings, those skilled in the art can apply various technical modifications and variations based on the above. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or if the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
[0218] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below. Explanation of the symbols
[0220] 301: User terminal 302: Network 303: Analysis Server 304: External Server
Claims
Claim 1 An operation to generate a preprocessed user voice signal by removing noise based on raw voice data received from a user terminal and a script presence indicator by a voice data receiving unit; an operation to derive a final ASR error text, a final alignment acoustic tensor, and a final comparison correct text by an error-preserving speech recognition unit, by decoding into text preserving speech errors according to a probability distribution at the character level extracted from the acoustic features of the preprocessed user voice signal, excluding context-based spelling correction operations by a language model (LM) based on the preprocessed user voice signal and the script presence indicator, and decoding according to the probability distribution at the character level extracted from the acoustic features of the preprocessed user voice signal; an operation to obtain a first-order text-acoustic merge map and time-axis aligned acoustic feature information by decomposing the final ASR error text and the final comparison correct text into character units, respectively by a speech text and acoustic alignment unit, calculating the Levenshtein distance, and forcing alignment with the final alignment acoustic tensor based on the time axis; and an operation to obtain the first-order text-acoustic merge map and time-axis aligned acoustic feature information by a pronunciation error analysis unit A method for evaluating the pronunciation of a language learner, comprising: an operation of classifying phonological error types of a specific pattern by combining a merge map and time-axis aligned acoustic feature information, and generating a final pronunciation error analysis dataset by calculating a Character Error Rate (CER); an operation of generating final pronunciation evaluation and feedback data including a user's weaknesses and correction guide based on the final pronunciation error analysis dataset by a customized feedback generation unit; and an operation of processing the final pronunciation evaluation and feedback data into a visualized UI component and transmitting it to the user terminal by an evaluation result providing unit. Claim 2 In claim 1, the operation of deriving the final ASR error text, the final acoustic tensor for alignment, and the final correct text for comparison comprises: an operation by an acoustic feature extraction layer to convert the raw waveform of the preprocessed user voice signal into the frequency domain to generate an initial acoustic feature tensor having time-series features; an operation by an encoder neural network layer to compress the spatiotemporal contextual information of the initial acoustic feature tensor to generate an acoustic encoding hidden vector, and to synchronize the initial acoustic feature tensor for transmission to a lower layer and output it as a transmission acoustic feature tensor; and an operation by a CTC decoder layer to calculate an independent initial consonant-vowel probability distribution for each time frame without intervention of probability values of the language model based on the acoustic encoding hidden vector. A method for evaluating a language learner's pronunciation, comprising: performing a beam search within a vocabulary network of pre-loaded correct texts based on the initial consonant-vowel probability distribution and the acoustic feature tensor for transmission, by means of a restricted search decoding layer, to determine the final ASR error text, the final alignment acoustic tensor, and the final comparison correct text that map the highest probability path of the sound exactly as heard. Claim 3 In claim 1, the operation of acquiring the primary text-acoustic merging map and time-axis aligned acoustic feature information comprises: an operation of parsing the final ASR error text and the final correct text for comparison into minimum phonological sequences including initial consonants, medial vowels, and final consonants, respectively, rather than morpheme units, by means of a character unit decomposition layer, converting them into a first character sequence and a second character sequence, and delay-synchronizing the final alignment acoustic tensor for time-axis computation; and an operation of calculating the edit distance between the first character sequence and the second character sequence by means of a Levenstein distance calculation layer using dynamic programming to generate a character unit node edit map including accurate location information of nodes where insertion, deletion, and substitution occurred. A method for evaluating a language learner's pronunciation, comprising: performing a forced alignment by means of a phoneme-level forced alignment layer to map text nodes included in the consonant-character unit node editing map based on the timestamp frame of the delayed-synchronized acoustic tensor, and outputting a combination of the first text-acoustic merge map and the time-axis aligned acoustic feature information. Claim 4 A language learner pronunciation evaluation system according to claim 1, wherein the utterance text and acoustic alignment unit comprises a character unit decomposition layer that parses the final ASR error text and the final correct text for comparison into initial, medial, and final consonant units, respectively, to generate a first character sequence and a second character sequence. Claim 5 A language learner pronunciation evaluation system according to claim 4, wherein the utterance text and acoustic alignment unit comprises a Levenstein distance calculation layer that calculates the minimum edit distance between the first letter sequence and the second letter sequence using dynamic programming and defines the state of each letter node as one of insertion, deletion, substitution, and match to generate a letter unit node edit map. Claim 6 A language learner pronunciation evaluation system according to claim 5, wherein the speech text and acoustic alignment unit comprises a phoneme-level forced alignment layer that generates a primary text-acoustic merge map by mapping the timestamp of the final alignment acoustic tensor and the first consonant sequence in frame units to determine a speech segment corresponding to each node of the consonant-character unit node editing map. Claim 7 A language learner pronunciation evaluation system according to claim 6, wherein the pronunciation error analysis unit calculates the character error rate (CER) of the learner's pronunciation relative to the correct consonant and vowel based on time-axis aligned acoustic feature information extracted from the primary text-acoustic merging map, and classifies a specific pattern of phonological error types by analyzing nodes identified as substitutions in the node editing map. Claim 8 A language learner pronunciation evaluation system according to claim 7, wherein the pronunciation error analysis unit determines whether one or more of the Korean phonological change rules, namely linking, nasalization, liquidization, gemination, and palatalization, have been violated to generate a final pronunciation error analysis dataset. Claim 9 A language learner pronunciation evaluation system according to claim 8, wherein the customized feedback generation unit searches for timestamps of error points recorded in the final pronunciation error analysis dataset in the primary text-acoustic merge map and generates final pronunciation evaluation and feedback data including a customized articulation method guide corresponding to the phonological error type along with speech waveform visualization information of the corresponding section.