System for recognizing pronunciation based on artificial intelligence
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- GONGTEO ENGLISH CO LTD
- Filing Date
- 2025-02-18
- Publication Date
- 2026-08-03
Smart Images

Figure 112025018420107-PAT00014_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to an artificial intelligence-based pronunciation recognition system, and more specifically, to an artificial intelligence-based pronunciation recognition system that recognizes a learner's voice input in response to text output through a display module or sound output through a speaker module, measures the accuracy of the pronunciation through an artificial neural network, and distinguishes and displays the pronunciation accuracy for each word according to preset evaluation criteria. Background Technology
[0003] Generally, offline learning has secured its place in the educational service sector through schools, private academies, private tutoring, and the study material market. However, collective education in institutions such as schools, based on paper-based training (PBT), involves studying the same curriculum in the same space; consequently, it does not consider individual learning levels and fails to take learner achievement into account due to the instructor's fragmented time constraints and the one-way delivery of instruction. In other words, while some students understand the lesson content after it ends, most students who do not understand simply move on from the session without reviewing the material.
[0004] The biggest problem with this constructivist learning is that when a hierarchy of superiority arises among peer groups, members with feelings of inferiority face significant difficulties in requesting a retake from the instructor regarding their shortcomings after the class ends.
[0005] Consequently, there are cases where people try to reacquire this knowledge through private academies, tutoring, and workbooks, which incurs significant opportunity costs during the learning process, leading to increased private education expenses and becoming a major burden on households.
[0006] Meanwhile, as the IT industry develops, this offline learning is expanding into online learning (e-Learning) through the internet, which allows learning through a new method called online services, starting from 2001.
[0007] This online learning method has the advantage of allowing for repeated learning, is not restricted by time or place, and enables advance learning because it involves taking pre-recorded lectures.
[0008] For this reason, through such continuous growth, online learning has come to form a pillar of the education industry as a learning method on par with offline learning, serving as both a complement and a substitute.
[0009] However, unlike offline education, guidance for learners attending a class is extremely limited because conventional real-time online education must rely solely on a monitor to monitor the learner's status.
[0010] In particular, when the relationship between the educator and the learner is one-to-many, controlling the learners is nearly impossible. Consequently, even if learners doze off or engage in behaviors unnecessary for learning, it is difficult for the educator to detect this and guide them to focus, as is done in an offline setting. This presents a critical problem in that the same learning effects as in an offline environment cannot be achieved.
[0011] Meanwhile, the aforementioned background technology is technical information that the inventor possessed for the derivation of the present invention or acquired during the process of deriving the present invention, and it cannot necessarily be considered publicly known technology disclosed to the general public prior to the filing of the present invention. Prior art literature
[0013] Korean Patent Publication No. 10-2022-0126412 The problem to be solved
[0014] One aspect of the present invention provides an artificial intelligence-based pronunciation recognition system that recognizes a learner's voice input in response to text output through a display module or sound output through a speaker module, measures the accuracy of pronunciation through an artificial neural network, and distinguishes and displays the pronunciation accuracy for each word according to preset evaluation criteria.
[0015] The technical problems of the present invention are not limited to those mentioned above, and other unmentioned technical problems will be clearly understood by those skilled in the art from the description below. means of solving the problem
[0017] An artificial intelligence-based pronunciation recognition system according to one embodiment of the present invention includes: a learner terminal that recognizes the voice of a learner speaking in response to text output through a display module or sound output through a speaker module; and an educational content management server that analyzes voice data received from the learner terminal through a pre-trained artificial neural network and distinguishes and displays the pronunciation accuracy for each word according to pre-set evaluation criteria.
[0018] The above-mentioned educational content management server is,
[0019] A learner analysis unit that calculates and scores the current level based on the learner's pre-test results to evaluate the learner's ability according to pre-set evaluation criteria; and
[0020] It includes a learning content management unit that generates foreign language sentences for foreign language pronunciation practice based on the evaluation results of the learner analysis unit mentioned above, and
[0021] The above-mentioned learning content management department is,
[0022] Voice data received from the learner terminal in correspondence with the above foreign language sentence is analyzed through a pre-trained artificial neural network, and the pronunciation accuracy of each word is distinguished and displayed in different colors according to pre-set evaluation criteria.
[0023] The above learner analysis unit is,
[0024] Calculate the evaluation score for each learner terminal using the following mathematical formula 1.
[0026] [Mathematical Formula 1]
[0027]
[0028] Here, St is the final evaluation score of the learner terminal, S1 is the first evaluation score calculated through the following mathematical formula 2, S2 is the second evaluation score calculated through the following mathematical formula 3, S3 is the third evaluation score calculated through the following mathematical formula 4, α is the first weight, β is the second weight, and γ is the third weight.
[0030] [Mathematical Formula 2]
[0031]
[0032] Here, S1 is the first evaluation score, N_ia is the number of words that match the correct answer among the response words received from the learner terminal for the i-th question, N_ic is the number of words that do not match the correct answer among the response words received from the learner terminal for the i-th question, t_i is the time (seconds) taken to solve the i-th question, n is the total number of questions, l_a is the average number of correct answers for other learners, N_a is the average number of words in the problem data, wi is the weight assigned to each question, and wa is the average weight.
[0034] [Mathematical Formula 3]
[0035]
[0036] Here, S_2 is the second evaluation score, t_s is the response threshold time per item (seconds), t_r is the actual response time per item (seconds), w is the first correction constant, t_rn is the average response time (seconds), t_an is the average response time of other learners (seconds), and P is the average time (seconds) taken from the time of the previous item's response to the completion of the current item's response.
[0038] [Mathematical Formula 4]
[0039]
[0040] Here, S3 is the third evaluation score, DA_i is the distance between the embedding vector of the correct word for the i-th question and the embedding vector of the response word, j_i is the score for the i-th question, DT_i is the distance between the embedding vector of the trap word set for the i-th question and the embedding vector of the response word, e is the natural constant, and K is the second correction constant. Effects of the invention
[0042] According to one aspect of the present invention described above, learning efficiency can be improved by recognizing the voice of a learner input in response to text output through a display module or sound output through a speaker module, measuring the accuracy of pronunciation through an artificial neural network, and distinguishing and displaying the pronunciation accuracy for each word according to preset evaluation criteria. Brief explanation of the drawing
[0044] FIG. 1 is a diagram showing the schematic configuration of an artificial intelligence-based pronunciation recognition system according to one embodiment of the present invention. Figure 2 is a diagram showing the specific configuration of the educational content management server illustrated in Figure 1. Specific details for implementing the invention
[0045] The following detailed description of the invention refers to the accompanying drawings, which illustrate specific embodiments in which the invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the invention. It should be understood that various embodiments of the invention are different but need not be mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the spirit and scope of the invention in relation to one embodiment. It should also be understood that the location or arrangement of individual components within each disclosed embodiment may be changed without departing from the spirit and scope of the invention. Accordingly, the following detailed description is not intended to be limiting, and the scope of the invention is limited only by the appended claims, including all equivalents to those claimed therein, provided appropriately described. Similar reference numerals in the drawings refer to the same or similar functions across various aspects.
[0046] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the drawings.
[0047] FIG. 1 is a diagram showing the schematic configuration of an artificial intelligence-based pronunciation recognition system according to one embodiment of the present invention.
[0048] The artificial intelligence-based pronunciation recognition system according to the present invention aims to analyze learners through artificial intelligence and big data analysis, and to provide customized educational content based on the analysis results.
[0049] In particular, the artificial intelligence-based pronunciation recognition system according to the present invention aims to improve the effectiveness of foreign language pronunciation education or pronunciation training for learners by recognizing the voice of a learner input in response to text output through a display module or sound output through a speaker module, measuring the accuracy of pronunciation through an artificial neural network, and distinguishing and displaying the pronunciation accuracy for each word according to preset evaluation criteria.
[0050] Specifically, the artificial intelligence-based pronunciation recognition system according to the present invention includes a learner terminal (100) and an educational content management server (200).
[0051] The learner terminal (100) is an electronic device possessed by a learner who wishes to receive online learning services, and may be in the form of a smartphone, PC, laptop, tablet PC, wearable device, kiosk, TV, speaker, etc., capable of wired and wireless communication with external devices and input / output and processing of information.
[0052] The learner accesses the educational content management server (200) through the learner terminal (100) and receives the online-based multi-learning service according to the present invention from the educational content management server (200).
[0053] The learner terminal (100) recognizes the voice of a learner speaking in response to text output through a display module or sound output through a speaker module and transmits it to an educational content management server (200).
[0054] For example, when a learner pronounces foreign language text output through a display device such as a TV or monitor provided in a foreign language academy, or pronounces a foreign language sentence output through the speaker of the display device, the learner terminal (100) transmits voice data recording the learner's voice to the education content management server (200).
[0055] The education content management server (200) is connected to multiple learner terminals (100) via wired or wireless communication and provides customized online education services for each learner terminal (100).
[0056] FIG. 2 is a diagram showing the specific configuration of such an educational content management server (200).
[0057] As described, the educational content management server (200) includes a learner analysis unit (210) and a learning content management unit (220).
[0058] The learner analysis unit (210) calculates evaluation scores for each learner terminal based on learner data stored in the educational content management server (200).
[0059] Here, the learner data stored in the educational content management server (200) includes all information input by a learner to use the online-based multi-learning system according to the present invention, such as task performance data, test data, and learning history data, as well as information generated during the process of using the online-based multi-learning system.
[0060] Specifically, the learner analysis unit (210) generates pre-evaluation content consisting of a first test question group, a second test question group, and a third test question group and transmits it to the learner terminal, and calculates an evaluation score based on response data received from the learner terminal regarding the pre-evaluation content.
[0061] At this time, the learner analysis unit (210) organizes each item group based on item response theory.
[0062] It is a test theory that focuses on each individual item constituting a test and estimates the characteristics of the item in relation to the subject's latent traits or abilities through the unique item characteristic curve of each item. The item characteristic curve, a core element of Item Response Theory, represents the functional relationship between a subject's latent traits or abilities and the probability of correctly answering an item. This functional relationship allows for the prediction of item scores based on the subject's latent traits or abilities, and the theory is developed based on the fundamental assumption of local independence, which holds that a subject with a certain ability will have a response to one item that is mutually independent from another.
[0063] For example, the first group of test items is composed of basic questions, the second group of test items is composed of questions where you read the question and choose the correct answer, and the third group of test items is composed of questions where you listen to the word in the question and choose the first sound of the word.
[0064] The learner analysis unit (210) generates pre-evaluation content composed of these items and transmits it to the learner terminal, and calculates an evaluation score based on response data received from the learner terminal regarding the pre-evaluation content. The specific process for calculating the evaluation score will be described later.
[0065] In this way, the educational content management server (200) can provide user-customized educational content based on the evaluation score calculated by the learner analysis unit (210). The learner's ability is evaluated according to the pre-set evaluation criteria of the learner analysis unit (210), and the process of calculating the specific evaluation score will be described later.
[0066] The learning content management unit (220) generates foreign language sentences for foreign language pronunciation practice based on the evaluation results of the learner analysis unit. That is, the learning content management unit (220) extracts a foreign language sentence suitable for the learner's level test result from among a plurality of sentences of different difficulty levels stored in advance and transmits it to the learner terminal (100). Accordingly, the learner terminal (100) directly displays the foreign language sentence received from the learning content management unit (220) as shown in FIG. 3, or allows the foreign language sentence to be output through another display device linked to the learner terminal (100).
[0067] The learning content management unit analyzes voice data received from the learner terminal in response to the foreign language sentence output through the display device, speaker device, or learner terminal (100) using a pre-trained artificial neural network.
[0068] In one embodiment, the learning content management unit separates the received voice data by word and extracts a first feature vector representing the meaning of the word and a second feature vector representing the tone features of the word for each word.
[0069] The learning content management department inputs the first and second feature vectors extracted from the correct answer word and the first and second feature vectors extracted from the words corresponding to the correct answer word among the learner's voice data into a pre-trained artificial neural network, and calculates the similarity between the correct answer word and the pronunciation word from the artificial neural network.
[0070] Here, the artificial neural network may be in the form of a deep neural network consisting of an input layer, a hidden layer, and an output layer.
[0071] Afterwards, the learning content management unit (220) displays the pronunciation accuracy of each word in different colors based on the analysis results (similarity) using an artificial neural network and according to pre-set evaluation criteria.
[0072] For example, as illustrated in FIGS. 4 and 5, the learning content management unit (220) displays words with a pronunciation accuracy (similarity) of 90% or more, which is the first standard range, in the first color, blue; words with a pronunciation accuracy between 71% and 89%, which is the second standard range, in the second color, green; words with a pronunciation accuracy between 51% and 70%, which is the third standard range, in the third color, yellow; words with a pronunciation accuracy between 31% and 49%, which is the fourth standard range, in the fourth color, red; and words with a pronunciation accuracy of less than 30%, which is the fifth standard range, in the fifth color, gray. Accordingly, learners can distinguish their pronunciation accuracy for each word by color alone, thereby improving learning efficiency.
[0073] In addition to this, the education content management server (200) may further include configurations such as a live class management department, a quiz content management department, a gamification content management department, a project management department, and an online mall management department, although these are not shown.
[0074] The learning content management unit (220) provides online learning content corresponding to the learner's competency evaluated by the learner analysis unit to the learner's terminal.
[0075] The learning content management department (220) provides an optimized educational course with a variety of digital content that stimulates interest according to the learner's level and stage, such as E-Library Channel, US School Channel, Game Channel, Toon Channel, and Media Channel.
[0076] The live class management unit (230) provides a real-time video conferencing service to enable a live class to take place between the instructor terminal that teaches the online learning content and the learner terminal.
[0077] The quiz content management unit (240) provides a quiz for verifying achievement corresponding to the learner's competency evaluated by the learner analysis unit to the learner terminal.
[0078] The quiz content management department (240) can maximize the efficiency of problem management by digitizing all problems created and accumulating them in a database along with meta-information.
[0079] The gamification content management unit (250) provides game content related to the above online learning content to the learner's terminal.
[0080] The gamification content management unit (250) controls the output of game content suitable for any one of five learning types, such as Unscramble Words, Unscramble Sentence, Sentence Speaking, Sentence Typing, and Word Typing, from the learner terminal (100).
[0081] The project management department (260) creates learner-led learning content.
[0082] Specifically, the project management department is characterized by generating the learner-led learning content through the steps of: setting a learning goal through the learner terminal; setting a weekly project to achieve the set learning goal; providing background knowledge information related to the weekly project to the learner terminal so that the learner can acquire background knowledge for each weekly project; and receiving presentation materials generated through an application pre-installed on the learner terminal.
[0083] That is, the project management department (260) assigns a weekly learning project to the learner for each stage of the course through the G+ learning process developed by the applicant, so that when the learner starts a week's lesson, they can set a project direction and set study strategy goals suitable for the project.
[0084] The project management department (260) sets learning goals (G), sets the direction of weekly project tasks (P), learns background knowledge required for the project (learn, L), incorporates background knowledge into the project (U), and completes and presents (show, S), thereby encouraging learners to move away from the existing passive and one-way method of acquiring knowledge and to approach learning with an active attitude.
[0085] The online mall management department manages the accumulation of rewards (points) generated during the learning process for each learner, enables the purchase of various products using the accumulated rewards, and allows products directly planned and produced by the learner to be registered on a webpage linked to the educational content management server (200) so that they can be sold and purchased through the online mall.
[0087] Meanwhile, the educational content management server (200) recommends customized learning by performing data mining on the big data of learners accumulated in the database, and recommends customized educational content suitable for the student by analyzing the student's interest level, tendencies, and learning content.
[0088] To this end, the learner analysis unit transmits problem data for learner evaluation to the learner terminal and receives response data regarding the problem data from the learner terminal. The learner analysis unit analyzes the response data and calculates evaluation scores for each learner terminal.
[0089] In one embodiment, the learner analysis unit calculates an evaluation score for each learner terminal using the following mathematical formula 1.
[0091] [Mathematical Formula 1]
[0092]
[0093] Here, St is the final evaluation score of the learner terminal, S1 is the first evaluation score calculated through the following mathematical formula 2, S2 is the second evaluation score calculated through the following mathematical formula 3, S3 is the third evaluation score calculated through the following mathematical formula 4, α is the first weight, β is the second weight, and γ is the third weight.
[0094] Here, the first evaluation score is calculated based on the characteristics of the correct answer word among the response data of learners who solved the evaluation item, the second evaluation score is calculated based on the characteristics of the time taken by learners who solved the evaluation item, and the third evaluation score is calculated based on the semantic similarity between the words included in the learners' response data and the words that are easily mistaken for the correct answer.
[0095] The specific process for calculating the first weight, the second weight, and the third weight will be described later.
[0096] In this way, the learner analysis unit according to the present invention evaluates the learner's learning from various perspectives using the above mathematical formula 1 and calculates the final evaluation score, thereby having the effect of improving the reliability of the calculated evaluation score.
[0097] Meanwhile, in the mathematical formulas described in the present invention, the variables of each term can be calculated by considering only the magnitude value of each variable. That is, for each variable according to the above-described mathematical formula, only the magnitude value of the variable is considered, and an importance score can be set using the magnitude value of the result of the left-hand term, which is the result calculated in this way.
[0098] In the mathematical formulas described in the present invention, the application of a logarithmic function to the calculated value is because the logarithmic function can express large values by relatively reducing them. That is, as the value of the logarithmic function increases, the value changes in a form that is less affected by units or scales. This allows the formula to respond less sensitively to changes in units, enabling the comparison of data of various scales on the same scale. Furthermore, the logarithmic function expresses the change in the x value as a linear value when the change in the x value has an exponential difference. This allows the range of change in large values to be appropriately reduced and expressed, making it useful when the change in the x value has a proportional nature. In the aforementioned mathematical formula according to the present invention, a logarithmic function is applied to the calculated value to prevent the result from changing significantly due to excessive reliance on the value of a specific term.
[0099] It is self-evident that by using this mathematical formula, when a technician of the ordinary sends consumer data to the analysis unit, the analysis unit generates variable values using the input information and substitutes the generated variable values into the already completed mathematical formula 2 to immediately utilize them in the relevant industry.
[0101] [Mathematical Formula 2]
[0102]
[0103] Here, S1 is the first evaluation score, N_ia is the number of words that match the correct answer among the response words received from the learner terminal for the i-th question, N_ic is the number of words that do not match the correct answer among the response words received from the learner terminal for the i-th question, t_i is the time (seconds) taken to solve the i-th question, n is the total number of questions, l_a is the average number of correct answers for other learners, N_a is the average number of words in the problem data, wi is the weight assigned to each question, and wa is the average weight.
[0104] For example, if ∑N_ia is 50, ∑N_ic is 20, ∑t_i is 600, ∑w_i is 150, n is 10, l_a is 4.5, N_a is 40, and wa is 5, the first evaluation score S_1 can be calculated as approximately 28.1 points.
[0105] Meanwhile, the number of words that match the correct answer and the number of words that do not match among the response words received from the learner terminal means, for example, that when the correct answer to question 10 is 'I go to school' and the correct answer to question 10 received from the learner terminal is 'You go to school', N_ia is set to 2 and N_ic is set to 1.
[0106] In addition, a lower evaluation score is calculated as the time taken to solve the questions increases, and the evaluation score is calculated in proportion to the number of questions answered correctly by the learner relative to the average number of correct answers.
[0107] In other words, the learner analysis unit can provide the effect of enabling a more accurate evaluation of the learner by using the aforementioned mathematical formula 2 to calculate an evaluation score based on how close the answer is to the correct answer, even if it is incorrect, rather than simply whether it is correct or incorrect.
[0108] In addition, the second evaluation score can be calculated as follows.
[0110] [Mathematical Formula 3]
[0111]
[0112] Here, S_2 is the second evaluation score, t_s is the response threshold time per item (seconds), t_r is the actual response time per item (seconds), w is the first correction constant, t_rn is the average response time (seconds), t_an is the average response time of other learners (seconds), and P is the average time (seconds) taken from the time of the previous item's response to the completion of the current item's response. For example, if it took 20 seconds to complete the response to item 2 after the response to item 1 was completed, and 30 seconds to complete the response to item 3 after the response to item 2 was completed, then P is 25.
[0114] For example, if ∑t_s is 100, ∑t_r is 80, ∑P is 40, t_an is 25, t_rn is 15, and w is 50, the second evaluation score S_2 can be calculated as approximately 15.5 points.
[0115] In other words, the reliability of the evaluation results can be improved by calculating a second evaluation score that takes into account cases where the learner writes the correct answer insincerely, such as guessing the answer or solving the problem carelessly.
[0117] [Mathematical Formula 4]
[0118]
[0119] Here, S3 is the third evaluation score, DA_i is the distance between the embedding vector of the correct word for the i-th question and the embedding vector of the response word, j_i is the score for the i-th question, DT_i is the distance between the embedding vector of the trap word set for the i-th question and the embedding vector of the response word, e is the natural constant, and K is the second correction constant.
[0120] The term "word embedding vector" refers to the magnitude (length) value of a vector excluding direction values, obtained by embedding the correct answer word into a vector space. By calculating the evaluation score based on the distance between the embedding vector of the correct answer word and the embedding vector of the response word, the evaluation score is calculated based on the semantic similarity of the response word, even if it is not the correct answer word intended by the question setter.
[0121] Similarly, the learner analysis unit calculates a third evaluation score by considering the distance between the embedding vector of the response word and the trap word, which is a word easily mistaken for the correct answer, so that even if the response word is semantically similar to the correct answer word, a lower evaluation score is produced the closer it is to the trap word intended by the questioner.
[0122] In this way, the learner analysis department can improve the reliability of the evaluation by evaluating the third-order response question area by comprehensively considering the similarity between the word set as the correct answer and the word responded by the learner, and the similarity between the trap word, which is a word with a pronunciation that is easily mistaken for the correct answer, and the word responded by the learner.
[0123] In this case, the final evaluation score St of the learner terminal can be calculated as approximately 72.7 points by adding the first evaluation score of 28.1 points, the second evaluation score of 15.5 points, and the third evaluation score of 29.1 points, according to the embodiments described above, when the first weight (α), the second weight (β), and the third weight (γ) are all 1.
[0125] In one embodiment, the first weight (α) described in Equation 1 may be any one integer value between 1 and 10 set by the manager.
[0126] In another embodiment, the first weight (α), the second weight (β), and the third weight (γ) can be calculated by the following mathematical formula.
[0128] [Mathematical Formula 5]
[0129]
[0130] Here, α is the first weight, n is the total number of items, qi is the score of the i-th item, and qv is the average score.
[0131] For example, n is 20 and If α is 5, the first weight a is calculated as 15.
[0133] [Mathematical Formula 6]
[0134]
[0135] Here, β is the second weight, n is the total number of items included in the problem data, and DIF_i is the difficulty score of the i-th item set by the administrator managing the education content management server (200).
[0136] For example, when ∑DIF_i is 15 and n is 10, the second weight β can be set to a value of 1.5.
[0138] [Mathematical Formula 7]
[0139]
[0140] Here, γ is the third weight, u is the time (in seconds) taken from the time of response to the first question of the problem data until the time of completion of the response to the last question, and g is the insincere response index, and the insincere response index is a value set by an administrator managing the education content management server (200), for example, it can be set as the time (in seconds) taken to read all questions included in the problem data.
[0141] For example, when u is 450 and g is 300, γ is set to a value of 1, and when u is 150 and g is 300, γ can be set to a value of 0.
[0142] In this way, the learner analysis unit calculates a first evaluation score for the first test result, a second evaluation score for the second test result, and a third evaluation score for the third test result using the mathematical formulas described above, and calculates the final evaluation score by performing calculations on the calculated evaluation scores.
[0143] In addition, the education content management server (200) manages the education content by assigning evaluation scores in advance, and then provides the online education content having a level score corresponding to the calculated evaluation score to the learner's terminal, thereby providing customized education content for the learner and improving the reliability of the analysis process.
[0144] In some other embodiments, the educational content management server (200) may further include an instructor evaluation unit (not shown) that analyzes instructor-specific review data received from a learner terminal (100) and calculates an instructor-specific rating.
[0145] For example, the instructor evaluation department can construct an artificial neural network that extracts contextual information from input data.
[0146] Here, the input data can be review data about the instructor.
[0147] The instructor evaluation unit can classify and accumulate review data received from the learner terminal (100) by instructor and can extract the stored review data as learning data.
[0148] The instructor evaluation department can construct a neural network that extracts contextual information from input data by training the data using the Word2Vec algorithm.
[0149] The Word2Vec algorithm may include a Neural Network Language Model (NNLM). A Neural Network Language Model is fundamentally a neural network composed of an Input Layer, a Projection Layer, a Hidden Layer, and an Output Layer. The Neural Network Language Model is used as a method for vectorizing words. Since the Neural Network Language Model is a well-known technology, a more detailed explanation will be omitted.
[0150] The Word2vec algorithm is designed for text mining and determines proximity based on the preceding and succeeding relationships between words. It is an unsupervised learning algorithm. As its name suggests, Word2vec is a quantitative technique that represents the meaning of words in vector form. The Word2vec algorithm can represent each word as a vector in a space of approximately 200 dimensions. By utilizing the Word2vec algorithm, a vector corresponding to each word can be obtained.
[0151] The Word2vec algorithm can enable a dramatic improvement in precision in the field of natural language processing compared to other conventional algorithms. Word2vec learns the meaning of words by utilizing the relationships between words and adjacent words within sentences of an input corpus. Based on artificial neural networks, the Word2vec algorithm starts from the premise that words with the same context carry similar meanings. The algorithm learns through text documents, training the neural network to identify related words by identifying other words that appear nearby (approximately 5 to 10 words before and after) a given word. Since words with related meanings are highly likely to appear close together within a document, the two words can gradually acquire closer vectors as the learning process is repeated.
[0152] There are two training methods for the Word2vec algorithm: CBOW (Continuous Bag Of Words) and skip-gram. The CBOW method predicts a target word by utilizing the context created by surrounding words. The skip-gram method predicts words that may follow a single word. The skip-gram method is known to be more accurate in large-scale datasets.
[0153] Accordingly, in the embodiments of the present invention, a Word2vec algorithm using the skip-gram method is used. For example, if training is successfully completed through the Word2vec algorithm, similar words can be located nearby in a high-dimensional space. According to the Word2vec algorithm described above, the closer the distribution of surrounding words within a training document, the more similar the calculated vector values can be, and words with similar calculated vector values can be considered similar. Since the Word2vec algorithm is a known technology, a more detailed explanation regarding the calculation of vector values will be omitted.
[0154] The instructor evaluation unit can divide the rating levels into multiple stages and set evaluation criterion text corresponding to each rating level. For example, the instructor evaluation unit can access an external server where ratings have been assigned by experts based on evaluation results derived from review data to obtain evaluation criterion text for multiple rating levels.
[0155] The instructor evaluation department inputs multiple evaluation criterion texts for each rating level into a neural network, and can extract a rating level criterion vector value representing contextual information for the evaluation criterion texts for each rating level.
[0156] The instructor evaluation unit can input text of the evaluation result for the review data received from the learner terminal (100) into the neural network to extract an evaluation result vector value representing contextual information.
[0157] The instructor evaluation department calculates the similarity between the evaluation result vector value and each of the multiple reference vector values, and can extract the reference vector value with the highest similarity to the evaluation result vector value among the multiple reference vector values. In this case, Euclidean distance, cosine similarity, Tanimoto coefficient, etc., may be adopted as methods for calculating similarity.
[0158] The instructor evaluation department can calculate the rating level corresponding to the reference vector value with the highest similarity to the evaluation result vector value as the instructor's rating.
[0159] The instructor evaluation department can store the calculated scores for each instructor and provide the scores for each instructor through an application run by the learner (100).
[0160] In some other embodiments, the instructor evaluation unit may calculate the instructor-specific rating using the following mathematical formula 8.
[0162] [Mathematical Formula 8]
[0163]
[0164] Here, w_t is the total number of words extracted from review data received from the learner terminal (100), w_p is the number of words included in the pre-learned positive evaluation word list among the words extracted from the review data, w_n is the number of words included in the pre-learned negative evaluation word list among the words extracted from the review data, w_i is the number of words included in the pre-learned inappropriate word list among the words extracted from the review data, and S_t is the final evaluation score of the learner terminal calculated by the above-described mathematical formula 1.
[0165] Meanwhile, in the aforementioned mathematical formula, the variables of each term can be calculated by considering only the magnitude value of each variable. That is, according to the aforementioned mathematical formula, only the magnitude value of each variable is considered, and the rating can be set using the magnitude value of the result on the left side, which is the result calculated in this way.
[0166] In the mathematical formula described above, the reason for applying a logarithmic function to the operation value of each term is that the logarithmic function can express large values by relatively reducing them. That is, as the value of the logarithmic function increases, the value changes in a form that is less affected by units or scales. This allows the formula to respond less sensitively to changes in units, enabling the comparison of data of various scales on the same scale. Furthermore, the logarithmic function expresses changes in x values as linear values when they have exponential differences. This allows the range of change in large values to be appropriately reduced and represented, making it useful when the change in x values has a proportional nature. In the mathematical formula described above according to the present invention, a logarithmic function was applied to the operation value of each term to prevent the magnitude of the evaluation score from changing significantly due to the operation value of each term.
[0167] In this way, the instructor evaluation unit calculates an instructor-specific rating using the aforementioned mathematical formula 8, thereby calculating a rating proportional to the ratio (difference value) of negative words to positive words in the review data. In this process, the more inappropriate words (w_i), such as profanity and sexual harassment, there are, the more negligibly negative words are applied to the rating. Additionally, by calculating an instructor-specific rating in proportion to the final evaluation score of the learner's terminal, the tendency of learners with low final evaluation scores to maliciously give negative evaluation scores can be prevented in advance. Furthermore, by calculating an instructor-specific evaluation score using mathematical formula 8, which is designed to calculate a rating proportional to the total number of words, as a larger amount of review data (number of words) is judged to be a more reliable evaluation, a reliable evaluation can be achieved.
[0168] Meanwhile, the aforementioned pre-learned positive evaluation word list, negative evaluation word list, and inappropriate word list can be set by an administrator managing the educational content management server (200). For example, the positive evaluation word list may include words or word combinations such as kindness, detail, enthusiasm, goodness, and easy to understand, while the negative evaluation word list may include words or word combinations such as difficulty to understand, carelessness, poor quality, and distraction. In this way, the administrator can classify words that directly affect the instructor evaluation into a positive evaluation word list, a negative evaluation word list, an inappropriate word list, etc., and pre-learn them.
[0169] The technology according to the present invention, as described above, may be implemented in the form of program instructions that can be executed through various computer components or implemented as an application, and may be recorded on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., either individually or in combination.
[0170] The program instructions recorded on the above-mentioned computer-readable recording medium are those specifically designed and configured for the present invention, but may also be those known and available to those skilled in the art of computer software.
[0171] Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions such as ROM, RAM, and flash memory.
[0172] Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware device may be configured to operate as one or more software modules to perform processing according to the present invention, and vice versa.
[0173] Although the invention has been described above with reference to embodiments, those skilled in the art will understand that various modifications and changes can be made to the invention without departing from the spirit and scope of the invention as set forth in the following claims. Explanation of the symbols
[0175] 210: Learner Analysis Department 220: Customized Problem Management Department
Claims
Claim 1 An artificial intelligence-based pronunciation recognition system comprising: a learner terminal that recognizes the voice of a learner speaking in response to text output through a display module or sound output through a speaker module; and an educational content management server that analyzes voice data received from the learner terminal through a pre-trained artificial neural network and distinguishes and displays pronunciation accuracy for each word according to pre-set evaluation criteria, wherein the educational content management server comprises: a learner analysis unit that calculates a current level by scoring it based on the learner's pre-test results in order to evaluate the learner's ability according to pre-set evaluation criteria; and a learning content management unit that generates foreign language sentences for foreign language pronunciation practice based on the evaluation results of the learner analysis unit.It includes an instructor evaluation unit that calculates an instructor rating by analyzing instructor-specific review data received from a learner terminal, and the learning content management unit analyzes voice data received from the learner terminal corresponding to the foreign language sentence through a pre-trained artificial neural network to distinguish and display pronunciation accuracy for each word in different colors according to pre-set evaluation criteria, and the learner analysis unit calculates the final evaluation score of the learner terminal by summing a first evaluation score calculated based on response words received from the learner terminal, a second evaluation score calculated based on response time per question, and a third evaluation score calculated based on the difference value of embedding vectors between the correct word and the trap word, wherein a first weight is reflected in the first evaluation score, a second weight is reflected in the second evaluation score, and a third weight is reflected in the third evaluation score, and the first, second, and third evaluation scores with different weights are summed to calculate the final evaluation score of the learner terminal, and the first weight is the total number of questions and An AI-based pronunciation recognition system characterized by being calculated based on the point value for each item, wherein the second weight is calculated based on the total number of items and the difficulty score for each item, and the third weight is calculated based on the time elapsed from the time of response to the first item to the time of completion of response to the last item and the insincere response index, wherein the insincere response index is set as the time required to read all items included in the problem data, and wherein the instructor evaluation unit calculates an instructor-specific rating using the following mathematical formula. [Mathematical Formula]; Here, w_t is the total number of words extracted from review data received from the learner terminal, w_p is the number of words included in a pre-learned positive evaluation word list among the words extracted from the review data, w_n is the number of words included in a pre-learned negative evaluation word list among the words extracted from the review data, w_i is the number of words included in a pre-learned inappropriate word list among the words extracted from the review data, and S_t is the final evaluation score of the learner terminal calculated by the above-described mathematical formula 1, wherein the positive evaluation word list includes words or word combinations such as kindness, detail, diligence, good, and easy to understand, the negative evaluation word list includes words or word combinations such as difficult to understand, carelessness, not very good, and distracting, and the inappropriate word list includes words or word combinations corresponding to profanity and sexual harassment. Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 delete Claim 6 delete