System and method for visual symbolization of phonetic signal for english learning

The system addresses the limitations of existing English learning systems by visually representing stress, pitch, and duration in English pronunciation, enabling intuitive user feedback and improved learning outcomes.

WO2026155432A1PCT designated stage Publication Date: 2026-07-23LEE GI HUN
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
LEE GI HUN
Filing Date
2025-12-26
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing English learning systems fail to intuitively convey the multidimensional phonetic features of English pronunciation, such as stress, intonation, and rhythm, limiting user understanding and correction of pronunciation errors.

Method used

A system that analyzes voice signals to visually represent stress, pitch, and pronunciation duration, using symbols with varying thickness, angle, and color to intuitively convey these features, and generates metadata for personalized learning feedback.

Benefits of technology

Enables users to intuitively understand and correct English pronunciation and intonation by providing customized visual and text feedback, enhancing learning efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025022908_23072026_PF_FP_ABST
    Figure KR2025022908_23072026_PF_FP_ABST
Patent Text Reader

Abstract

According to an embodiment of the present invention, a system for visual symbolization of a phonetic signal for English learning comprises: a phonetic signal input unit for receiving a phonetic signal; a phonetic signal feature extraction unit for extracting feature data including stress, pitch, and pronunciation duration from the input phonetic signal; a symbol attribute value determination unit for determining a symbol attribute value on the basis of the feature data; a visual symbolization unit for expressing the phonetic signal by using a visual symbol on the basis of the determined symbol attribute value; and an output unit for outputting visual symbol data to a user device, wherein the visual symbol is expressed by a combination of stress shading and a reference line, the stress shading having a different thickness on the basis of a stress magnitude and being displayed overlappingly with a base character symbol, and the reference line crossing a center point of the base character symbol at a predetermined angle while cutting the stress shading.
Need to check novelty before this filing date? Find Prior Art

Description

System and method for visual encoding of speech signals for English learning

[0001] The present invention relates to a system and method for visual encoding of speech signals for English learning, and more specifically, to speech signal processing and encoding technology that supports users in intuitively perceiving and learning by visually representing speech features such as English pronunciation, intonation, stress, and rhythm.

[0002] In modern society, English has established itself as a major tool for international communication, and the importance of learning English is increasing day by day. In particular, phonetic characteristics such as pronunciation accuracy, intonation, stress, and rhythm are considered key elements for smooth communication.

[0003] Language is essentially a pattern-matching process characterized by being acquired through long-term learning and experience. In particular, English is a non-phonetic and stress-timed language, where combinations of stress, pitch, and length play a crucial role in conveying meaning. On the other hand, Korean is a phonetic and syllabic language, which causes Korean speakers to face fundamental difficulties in understanding and mastering the English stress system.

[0004] Figure 1 is a reference diagram illustrating an example of voice visualization of a system for English learning according to the prior art.

[0005] Existing English learning systems adopt a method of expressing stress in speech signals through simple symbols or text emphasis, as shown in Fig. 1. In Fig. 1-(1), stress is indicated by the size of a dot, and in Fig. 1-(2), stress is visually indicated through word emphasis processing (boldness, color, etc.). This method has limitations in enabling users to learn English.

[0006] Specifically, existing methods express only the single characteristic of stress and fail to reflect other phonetic features, such as intonation, duration, and pitch, which are important elements in English learning. Since the naturalness of English pronunciation stems from the harmony of multidimensional features, including not only stress but also intonation and rhythm, the learning effect is limited when emphasizing a single element.

[0007] Furthermore, methods that express stress through dot size or text emphasis may fail to intuitively convey the differences in stress to users. For example, while users may recognize that the stress in “play” and “sports” differs, it is difficult for them to grasp how this difference relates to pitch or pronunciation duration. Rather than users simply learning visual patterns through repetition, a more specific and multidimensional method of expression is needed to help them understand the relationship between stress and intonation.

[0008] Furthermore, English learning tools currently on the market primarily focus on functions such as converting speech to text or scoring pronunciation accuracy. These systems have limitations in helping users intuitively understand phonetic characteristics like stress, pitch, and duration. In particular, correcting pronunciation and intonation errors solely by listening to audio is a very difficult task for users.

[0009] To address this, some technologies attempt to analyze speech signals and represent them visually; however, existing technologies often analyze only partial elements such as stress or pitch, or are limited to simply providing feedback. Furthermore, there are still limitations in normalizing analyzed data for use in learning systems or in providing customized learning environments where users can receive personalized feedback.

[0010] Therefore, there is an urgent need to develop technology that can provide intuitive visual feedback by integrally analyzing and encoding the stress, pitch, and pronunciation timing of speech signals, while simultaneously supporting the correction of users' pronunciation and intonation using normalized data.

[0011] The technical problem that the present invention aims to solve is to provide a system and method that analyzes a voice signal to visually represent voice characteristics such as stress, pitch, and pronunciation duration, thereby enabling a user to intuitively understand and correct English pronunciation and intonation.

[0012] Furthermore, the technical objective of the present invention is to implement a system that normalizes and patterns visually encoded voice data to utilize it as user-customized learning data.

[0013] Furthermore, the technical objective of the present invention is to quantitatively analyze voice feature data and generate metadata to provide a platform that can be utilized in learning systems and various application fields.

[0014] To achieve the above technical objective, a visual encoding system for a voice signal for English learning according to an embodiment of the present invention comprises: a voice signal input unit that receives a voice signal; a voice signal feature extraction unit that extracts feature data including stress, pitch, and duration from the input voice signal; a symbol attribute value determination unit that determines a symbol attribute value based on the feature data; a visual encoding unit that expresses a visual symbol based on the determined symbol attribute value; and an output unit that outputs the visual symbol data to a user device. The visual symbol may be expressed as a combination of a stress shade that is displayed overlapping a basic character symbol as a shade having a different thickness based on the stress size, and a reference line that cuts the stress shade and crosses the center point of the basic character symbol at a predetermined angle.

[0015] Here, the angle of the reference line can be used to visually represent the pitch of the tone and the change in the pitch of the tone over time.

[0016] In addition, it may further include a normalization unit that converts the feature data of the extracted voice signal into normalized data.

[0017] In addition, it may further include a metadata generation unit that generates data quantifying the structure and characteristics of the above symbol data.

[0018] Additionally, it may further include a pattern providing unit that provides a similar pattern sentence based on the above-mentioned generated metadata.

[0019] In addition, regarding the feature data in the above symbol attribute value determination unit, stress can be classified into general stress, main stress, and weak / unstress patterns, and the pitch level and changes over time can be classified into no change, rise, gradual rise, extend, gradual fall, and fall patterns.

[0020] In addition, the number of symbols can be determined as a symbol attribute value in the above symbol attribute value determination unit to correspond to the pronunciation duration.

[0021] To achieve the above technical objective, a method for visualizing a voice signal using a visual symbolization system for English learning according to an embodiment of the present invention comprises: (a) receiving a voice signal; (b) extracting feature data including stress, pitch, and duration of the voice signal from the input voice signal; (c) determining a symbol attribute value based on the feature data; (d) expressing it as a visual symbol based on the determined symbol attribute value; and (e) outputting the visual symbol data to a user device. The visual symbol may be expressed as a combination of a stress shade that is displayed overlapping a basic character symbol as a shade having a different thickness based on the stress size, and a reference line that cuts the stress shade and crosses the center point of the basic character symbol at a predetermined angle.

[0022] Here, the angle of the reference line can be used to visually represent the pitch of the tone and the change in the pitch of the tone over time.

[0023] In addition, the method may further include a step of converting the feature data of the extracted voice signal into normalized data.

[0024] In addition, it may further include a step of generating metadata that quantifies the structure and characteristics of the above symbol data.

[0025] In addition, it may further include a step of providing a similar pattern sentence based on the above-mentioned generated metadata.

[0026] In addition, in step (c) above, regarding the feature data, stress can be classified into general stress, main stress, and weak / unstress patterns, and the pitch level and changes over time can be classified into no change, rise, gradual rise, extend, gradual fall, and fall patterns.

[0027] In addition, in step (c) above, the number of symbols can be determined as a symbol attribute value to correspond to the pronunciation duration.

[0028] According to an embodiment of the present invention, the visual symbolization system for speech signals for English learning and the speech signal visualization system and method for English learning have the effect of enabling users to intuitively understand the characteristics of English pronunciation by integrally analyzing stress, pitch, and pronunciation duration and visually symbolizing them.

[0029] In addition, according to an embodiment of the present invention, a visual encoding system for speech signals for English learning can improve the efficiency of pronunciation correction and learning by generating metadata for the visually encoded data and providing customized feedback.

[0030] The effects of the present invention are not limited to the effects described above, and should be understood to include all effects that can be inferred from the configuration of the invention described in the detailed description of the invention or the claims.

[0031] Figure 1 is a reference diagram illustrating an example of voice visualization of a system for English learning according to the prior art.

[0032] FIG. 2 is a schematic diagram illustrating the overall operation process of a voice signal visual encoding system according to an embodiment of the present invention.

[0033] FIG. 3 is a block diagram of a voice signal visual encoding system according to an embodiment of the present invention.

[0034] FIG. 4 is a reference diagram illustrating the visual symbolization of the alphabet “O” according to an embodiment of the present invention.

[0035] FIG. 5 is a reference diagram illustrating the encoding data and visual encoding results of a voice signal according to an embodiment of the present invention.

[0036] FIG. 6 is a block diagram illustrating detailed components of a voice signal feature extraction unit according to an embodiment of the present invention.

[0037] FIG. 7 is a block diagram illustrating detailed components of a symbol attribute value determination unit according to an embodiment of the present invention.

[0038] FIG. 8 is a flowchart illustrating the processing process of a visual encoding system for voice signals according to an embodiment of the present invention in steps.

[0039] FIG. 9 is a flowchart illustrating the detailed operation steps of a voice signal feature extraction unit according to an embodiment of the present invention.

[0040] FIG. 10 is a flowchart illustrating the detailed operation steps of a symbol attribute value determination unit according to an embodiment of the present invention.

[0041] FIG. 11 is a reference diagram illustrating the sentence-unit encoding process of a voice signal visual encoding system according to an embodiment of the present invention in steps.

[0042] FIG. 12 is a reference diagram illustrating the encoding process of a voice signal according to an embodiment of the present invention.

[0043] The present invention will be described below with reference to the attached drawings. However, the present invention may be implemented in various different forms and is therefore not limited to the embodiments described herein. Furthermore, in order to clearly explain the present invention in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification have been given similar reference numerals.

[0044] Throughout the specification, when it is stated that a part is “connected (connected, in contact, combined)” with another part, this includes not only cases where they are “directly connected” but also cases where they are “indirectly connected” with other members interposed between them. Furthermore, when it is stated that a part “includes” a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but rather allows for the inclusion of additional components.

[0045] The terms used herein are merely for describing specific embodiments and are not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as “comprising” or “having” are intended to indicate the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0046] In this specification, the term “module” includes a unit composed of hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be a component formed as a whole, or a minimum unit or part thereof that performs one or more functions. For example, a module may be composed of an application-specific integrated circuit (ASIC).

[0047] Embodiments of the present invention will be described in detail below with reference to the attached drawings.

[0048]

[0049] FIG. 2 is a schematic diagram illustrating the overall operation process of a voice signal visual encoding system according to an embodiment of the present invention. FIG. 2 illustrates the structural flow in which the voice signal visual encoding system (100) of the present invention receives various input data, such as user voice, text, and video data, analyzes them, and visually encodes and outputs them.

[0050] The voice signal input unit (20) collects various input data, including the user's voice. The input data is processed by providing real-time voice data through a microphone, or by receiving a pre-stored voice file or text script. Additionally, voice signals can be extracted from video data and used for analysis. As shown in FIG. 2, for example, the sentence “Let's go” can be input.

[0051] The voice signal visual encoding system (100) analyzes voice signal data input from the voice signal input unit (20) and converts it into visual symbols that visually represent voice features. The voice visual encoding system extracts voice features of the voice signal and generates symbol attribute values ​​based on the extracted features. The generated symbol attribute values ​​can be used as data for performing visual encoding and are visualized in a form that a user can intuitively understand.

[0052] The voice-visual symbol output unit (30) is a module that ultimately provides visual symbolized data to the user. In the example of FIG. 2, the input sentence “Let’s go” is converted into a visual symbol and output. The visual symbol visually represents speech characteristics such as stress, tone, and pronunciation length for each word, and supports the user in easily understanding the pronunciation characteristics of the sentence. The visual symbolized result can be visually displayed on the user screen or saved and output as learning material.

[0053] The voice signal visual encoding system (100) of the present invention can be implemented as a program that is learned using artificial intelligence (AI) technology and can operate independently based on this. In particular, during the process of analyzing and encoding voice signals, a deep learning-based voice processing model can be used to automatically learn voice features (stress, pitch, pronunciation time) and derive results optimized for user data.

[0054] In addition, the system can utilize an AI model trained in a cloud environment or be implemented as a lightweight program optimized for execution on local devices (e.g., computers, smartphones, tablet PCs, etc.). This allows users to input voice signals in real time, receive visual feedback, and expand usability by integrating with various platforms.

[0055] In addition, the above implementation method can be applied to various application fields such as learning systems, pronunciation correction devices, and voice analysis software, and the accuracy and efficiency of the system can be improved through continuous AI model updates.

[0056]

[0057] FIG. 3 is a block diagram of a voice signal visual encoding system according to an embodiment of the present invention.

[0058] The voice signal visual encoding system (100) illustrated in FIG. 3 is a system that receives and analyzes a voice signal, converts it into a visual encoding form, and provides it to a user, focusing on increasing the learning efficiency and intuitiveness of the voice signal.

[0059] The above voice signal visual encoding system (100) may include a voice signal input unit (110), a voice signal feature extraction unit (120), a normalization unit (130), a symbol attribute value determination unit (140), a visual encoding unit (150), an output unit (160), a metadata generation unit (170), and a pattern providing unit (180).

[0060] First, the voice signal input unit (110) may be configured as an initial module for inputting voice pronounced by a user into the system. Analog voice signals or digitized voice data are collected through a microphone or a voice file upload method. Additionally, scripts or text documents that can be converted into voice signals may be collected.

[0061] The voice signal feature extraction unit (120) analyzes the voice signal or voice data input from the voice signal input unit (110) to extract voice signal feature data such as stress, pitch, and duration. The features of the voice signal can be expressed in the frequency domain. For example, the fundamental frequency (F0) of the voice can be selected to express the stress, pitch, and duration of the voice.

[0062] The normalization unit (130) normalizes the voice signal feature data extracted from the voice signal feature extraction unit (120) to correct differences in speech between users. The normalization unit (130) can convert the voice data into a certain range or scale to increase the comparability of the data and can adjust the data for user-customized data processing. For example, it can correct the fundamental frequency (F0) of the voice according to age group or adjust the duration of pronunciation to a constant level. In addition, it can reduce differences in speech characteristics between men and women, and between adolescents and adults, and make adjustments so that users can understand English pronunciation more accurately. The normalization unit (130) can ensure data consistency by converting the characteristics of different voice data into a unified standard.

[0063] The symbol attribute value determination unit (140) assigns a symbol attribute value corresponding to each voice feature based on the features of the voice signal or normalized data. For example, stress can be converted into the size or thickness of the symbol, pitch into the direction or angle of the symbol, and pronunciation duration into the length or number of repetitions of the symbol.

[0064] The detailed process of the symbol attribute value determining unit (140) determining the symbol attribute value will be explained in detail in FIG. 7 and FIG. 10.

[0065] The visual encoding unit (150) integrates the symbol attribute values ​​converted by the symbol attribute value determination unit (140) and implements them in a visual form. The visual encoding unit (150) combines attributes corresponding to speech signal characteristics such as stress, pitch, and pronunciation time to construct a visual layout or screen element. The arrangement of symbols is adjusted to match the sentence flow and rhythm so that the user can intuitively understand the characteristics of the speech signal.

[0066] In addition, the visual symbolization unit (150) aligns the symbols of words and syllables along the time axis by time axis-based visualization, thereby clearly indicating the timing and order of pronunciation of the voice signal. As a result, the user can visually confirm how the pronunciation of each syllable and word is arranged and changes over time.

[0067] In addition, the visual symbolization unit (150) generates a multidimensional visual symbol by combining stress, pitch, and pronunciation time through multidimensional symbolization. For example, stress is expressed as the size or color of the symbol, pitch as the slant or position of the symbol, and pronunciation time as the length or repetition of the symbol. The multidimensional visual symbol more clearly conveys the complex features of the voice signal and helps the user simultaneously recognize changes in the rhythm, intonation, and duration of the voice.

[0068] As a result, the visual encoding unit (150) integrates features such as stress, tone, and pronunciation duration of the voice signal into a temporal flow and a multidimensional visual structure to provide a user-friendly learning environment.

[0069] The output unit (160) outputs the visual symbolization data generated by the visual symbolization unit (150) to a user device. The output can be provided in various forms, such as screen output, print output, file saving, voice feedback, and integration with a learning system.

[0070] The metadata generation unit (170) generates summary information for the encoded data to quantify the structure and characteristics of the speech signal. For example, it can generate summary information corresponding to the rising and falling angles and slope values ​​of each syllable associated with pitch, and duration values ​​of syllables and words corresponding to pronunciation length. In addition, it can generate values ​​quantifying the degree of stress and values ​​corresponding to the time difference between each syllable or word, and generate values ​​quantifying the pronunciation rhythm, intonation pattern, etc. of the entire sentence.

[0071] The metadata generated above is utilized in the learning system analysis, pronunciation evaluation, and pattern provision unit (180), and can be used as basic data for comparing the user's pronunciation pattern with standard pronunciation.

[0072] In addition, the above-mentioned generated metadata can be configured as input data for training an artificial intelligence model.

[0073] The pattern providing unit (180) can provide feedback and learning materials to enhance pronunciation correction and learning effects by analyzing pronunciation patterns using the generated metadata and providing voice patterns similar to the pronunciation pattern corresponding to the metadata.

[0074] The pattern providing unit (180) can recommend other sentences or learning materials that have intonation and stress structures similar to the user's pronunciation pattern. For example, if the user has practiced “I want to play sports,” the pattern providing unit supports the user in repeatedly practicing pronunciation of similar patterns by recommending sentences such as “I need to finish work” that have similar rhythm and stress. In addition, the difficulty of the recommended sentences can be dynamically adjusted and provided according to the user's pronunciation difficulty and learning ability.

[0075] Additionally, the pattern providing unit (180) can provide visual and text feedback based on the results of the pronunciation characteristic analysis. For example, if a user pronounces “I want to play sports” and the stress on “want” is lacking, the word can be highlighted in bold or with a dark shade to visually indicate that stress is needed. At the same time, correction instructions can be provided through text feedback, indicating that “want” should be pronounced more strongly. From the visual and text feedback, the user can intuitively recognize the pronunciation error and correct it immediately.

[0076] In addition, the pattern providing unit (180) is not limited to English learning but can be extended to various languages. For example, it generates learning materials by analyzing pronunciation characteristics such as the stress structure, intonation patterns, and intervals between syllables of a specific language, and supports the user in effectively acquiring pronunciation habits suitable for the target language.

[0077] Accordingly, the pattern providing unit (180) provides immediate and accurate feedback to the user based on the analysis results of the pronunciation pattern, and recommends similar sentences or customized learning materials to maximize pronunciation correction and learning effects. In addition, it supports continuous learning tailored to learning goals by accumulating the user's pronunciation data and quantifying the improvement status.

[0078]

[0079] FIG. 4 is a reference diagram illustrating the visual symbolization of the alphabet “O” according to an embodiment of the present invention.

[0080] As illustrated in FIG. 4, the visual symbolization according to an embodiment of the present invention visually represents the corresponding syllable or word based on the shape of an alphabet character. In FIG. 4, the alphabet character “O” is used as an example of visual symbolization.

[0081] In Fig. 4, visual symbolism expresses that the emphasis is strong or weak by increasing or decreasing the size of the shading. Additionally, changes in tone can be indicated by the arrangement of the shading of the characters.

[0082] First, the rows (L1, L2, L3) of the table represent types of stress. The first row (L1) represents general stress, which is a symbolization of stress appearing in content words within words, phrases, or sentences. The second row (L2) represents syllables or words with the strongest stress among content words, i.e., the main stress, expressed by emphasizing strong phonetic features. The third row (L3) represents weak stress / unstress, which indicates cases of weak stress or almost no stress appearing mainly in function words.

[0083] The columns (C1–C6) of the table indicate the types of pitch changes over time. Column 1 (C1) represents a pronunciation that remains flat with no change in pitch. Column 2 (C2) represents a rising pitch pattern, indicating a raised intonation. Column 3 (C3) represents a gradual rising pitch pattern, indicating a change in intonation that gradually increases. Additionally, Column 4 (C4) represents an extended pitch pattern, indicating cases where pronunciation is drawn out. Column 5 (C5) represents a gradual falling pitch pattern, indicating a gradually decreasing intonation. Column 6 (C6) represents a falling pitch pattern, indicating a descending intonation.

[0084] In order to visually represent the height and changes over time of pitch, the present invention utilizes the angle of the baseline of the accent shading.

[0085] Here, the aforementioned stress shading is a visual symbol displayed overlaid on the basic character symbol as a shading of different thickness to indicate general stress and major stress.

[0086] In addition, the baseline of the above-mentioned accent shading is a line that cuts through the above-mentioned accent shading and crosses the center point of the above-mentioned basic character symbol, forming a predetermined angle with the said center point.

[0087] For example, the rise pitch is implemented at a 180-degree angle so that the shading faces the top of the baseline, thereby intuitively indicating an intonation where the pitch rises in the corresponding base character. For example, cells (L1, C2) and (L2, C2) in Fig. 4 represent the rise pitch in the general stress and main stress, respectively.

[0088] Fall pitch is implemented at a 180-degree angle so that the shading faces the bottom of the baseline, thereby intuitively expressing the falling intonation of the corresponding base character. For example, cells (L1, C6), (L2, C6), and (L3, C6) in Fig. 4 represent the falling pitch in normal stress, main stress, and weak or unstressed stress, respectively.

[0089] Gradual Rise Pitch visually represents cases where the intonation gradually rises by filling in the shading, with the baseline set at an angle tilted 45 degrees clockwise from the reference point. For example, cells (L1, C3) and (L2, C3) in Fig. 4 represent the gradual rise pitch in general stress and main stress.

[0090] Gradual Fall Pitch visually represents cases where the intonation gradually falls by filling in the shading, with the baseline set at an angle of 45 degrees counterclockwise from the reference point. For example, cells (L1, C5), (L2, C5), and (L3, C5) in Fig. 4 represent a gradual fall pitch in general stress, main stress, and weak or unstressed.

[0091] In addition, when expressing tonal changes using the angle of a baseline, gradual changes can be emphasized by setting the angle differently according to the intensity of the tonal change. Depending on the angle setting, it can be visually distinguished as a weak rise, a moderate rise, or a strong rise.

[0092] When visually representing tone, shading above a baseline can be used to indicate an increase in intonation. In this case, to visually emphasize gradual change, the shading becomes darker as it moves away from the baseline, and areas closer to the top can be expressed in darker colors. The gradient effect of the shading described above can intuitively express the pitch of the tone.

[0093] Additionally, in the case of gradually rising tones, the line can be represented in the form of a curve that starts at a low angle and gradually increases steeply. This allows for the clear visual distinction of the intensity of the intonation rise and helps users intuitively understand the gradual change in intonation.

[0094] The encoding method of the present invention provides visual intuitiveness of speech signals by integrating color, shading, and angle changes to simultaneously express the intensity, pitch, and temporal changes of speech. Furthermore, it can be utilized as a basic structure that can be effectively applied to speech learning systems.

[0095] As shown in FIG. 4, the cell (L1, C2), which is a combination of the first row (L1) and the second column (C2), is a symbol that combines normal stress and rising tone, representing a case where the intonation rises from normal stress. In this case, the symbol is implemented at an angle of 180 degrees so that the shading faces the top of the baseline, thereby emphasizing the effect of the rising tone. Thus, the user can intuitively understand that the tone of the word starts at a flat tone and rises towards the end. For example, it can express the rising intonation of “coming” at the end of the question in “Are you coming?”

[0096] Cell (L2, C4), a combination of the second row (L2) and the fourth column (C4), is a symbol that combines main stress and extended tone, representing a pronunciation pattern in which the main stressed syllable is drawn out. In this case, to indicate the main stress, the shade is enlarged or bolded from the basic character symbol, and the symbol is implemented at a 90-degree angle so that the shade faces to the right of the symbol's baseline. Additionally, to represent the extended pronunciation pattern, a shade with a different thickness from the basic symbol is superimposed on the basic character symbol. This visually represents that the stress is emphasized and the syllable is pronounced longer. For example, it can represent the case where "play" in "I want to play sports" is pronounced longer as the key emphasis of the sentence.

[0097] Additionally, cell (L3, C5), which is a combination of row 3 (L3) and column 5 (C5), is a symbol that combines weak / unstressed tone with a gradually descending tone, visually representing a pattern where the tone gradually descends while being weak or having almost no stress. In this case, the baseline is set at an angle of 45 degrees counterclockwise from the reference point, and the area below the baseline is depicted as thinner and filled with shading than the basic symbol. In particular, the top of the baseline is depicted without shading, intuitively indicating the gradual descent of the tone. By utilizing this symbolization, the characteristic of a weak intonation gradually descending can be clearly conveyed. For example, it can symbolize the case where a weak intonation, such as "to" in "I want to play sports," is pronounced with a gradually descending tone.

[0098] Cell (L1, C3), a combination of the first row (L1) and the third column (C3), is a symbol that combines general stress and a gradually rising tone, capable of expressing an intonation pattern that starts flat with general stress and then gradually rises in tone. In this case, the baseline is set at a 45-degree clockwise angle relative to the center point of the symbol, and the shading of the upper area of ​​the baseline is emphasized. Additionally, the lower part of the baseline remains unshaded to indicate a gradual rise in tone. For example, it can express the gradually rising intonation of "doing" at the end of the question in "What are you doing?".

[0099] Cell (L3, C1), which is a combination of the third row (L3) and the first column (C1), is a symbol that combines weak / unstressed and no-change tones, allowing the user to clearly recognize weak stress without tonal change. In this case, the symbol may be displayed transparently without shading, or may have a preset light color. For example, it can represent monotonous pronunciations such as “to” in “I want to play sports”.

[0100] As illustrated in Fig. 4, the method of expressing speech features using symbols allows users to intuitively understand intonation intensity and changes through angles and shading, and to clearly distinguish the differences between rising, falling, and gradual rising and falling tones. In addition, it allows for clear recognition of points where intonation naturally changes in pronunciation patterns, thereby supporting pronunciation learning and correction more effectively.

[0101]

[0102] FIG. 5 is a reference diagram illustrating the encoding data and visual encoding results of a voice signal according to an embodiment of the present invention.

[0103] FIG. 5 includes basic data (A1) of a voice signal and a result (B1) encoded using said data. A1 of FIG. 5 represents a spectrogram and RMS energy analyzed from a voice signal generated when pronouncing “Let’s go,” and B1 of FIG. 5 illustrates the result of visually representing the data of A1 by visualizing it.

[0104] A1 in Fig. 5 is basic data of a speech signal, and provides basic data obtained by analyzing the speech signal to implement visual symbolization such as B1. The basic data represents data from which features of the speech signal, including pitch, stress, and duration, can be extracted.

[0105] The frequency domain (A11), which is the upper region of A1 in Fig. 5, represents the temporal change in frequency that occurs when pronouncing “Let’s go” as “Let’s gooo.” Additionally, changes in pitch and intonation can be extracted from the frequency domain (A11) and used to determine the shading of the symbol and the angle of the baseline in visual symbolization. For example, the prolonged part of “gooo” exhibits a clearer pattern in the high-frequency band and is converted into a symbol with emphasized shading in the symbol of B1 in Fig. 5.

[0106] The time and word matching area (A12) displays the pronunciation time of each word along the time stamp axis and serves as a standard for generating visual representations for each word during the symbolization process. For example, it clearly indicates the positions where “Let,” “'s,” “go,” and “o” are pronounced over time.

[0107] The RMS energy change region (A13) indicates the intensity and pronunciation length of the speech signal. As the pronunciation of “go” in “Let’s go” continues for a longer period, a pattern of gradually decreasing RMS energy appears, which is visualized by the continuous overlapping of symbols representing “o” in B1 of Fig. 5. Additionally, the second and third “o” symbols, which represent the extended pronunciation of “o,” can be visualized as symbols such as cells (L1, C5) of Fig. 4 to indicate a gradually descending tone.

[0108] B1 in Fig. 5 represents the result of visual symbolization generated based on the data provided in A1 in Fig. 5. B1 is the result of analyzing characteristic data of a speech signal, such as tone, stress, and pronunciation length, and converting it into a visual symbol.

[0109] In B1 of Fig. 5, “Let” and “'s” are represented with weak stress and normal stress, respectively, with the weak stress appearing in light shading. On the other hand, the first “o” of “gooo” is represented with a darker upper shading of the symbol, indicating that it is pronounced with main stress and rising pitch.

[0110] In particular, the prolonged pronunciation of “go” is visualized by symbols overlapping continuously. These overlapping symbols visually emphasize the duration of the pronunciation.

[0111] The second and third “o” are implemented in the form of cells (L1, C5) in Fig. 4, reflecting a gradually descending tone.

[0112] Figure 5B1 is a diagram that clearly explains the effect of the visual symbolization technology based on the voice signal of the present invention, and expresses various characteristics of the voice signal, such as pitch, stress, and duration, in an integrated yet intuitive manner. As a result, users can easily understand complex changes in the voice signal, and it can be effectively utilized for pronunciation practice and intonation learning, particularly in application fields such as English learning.

[0113]

[0114] FIG. 6 is a block diagram illustrating detailed components of a voice signal feature extraction unit according to an embodiment of the present invention.

[0115] The voice signal feature extraction unit (120) may include a pitch analysis unit (121), a voice segment analysis unit (122), and an accent analysis unit (123).

[0116] First, the pitch analysis unit (121) analyzes the frequency band of the input voice signal to extract pitch information of the voice signal. The pitch information is used to visually represent the intonation pattern of each pronunciation segment and can be used as basic data to reflect pitch changes during the visual encoding process. For example, the trends of rise pitch and fall pitch in the pronunciation segment are analyzed.

[0117] The voice segment analysis unit (122) separates the input voice signal into segments by dividing intervals along the time axis and extracts the duration of each segment. The duration information can be used to determine the continuity of repeating symbols or shading during the encoding process. It can determine how long a segment with a constant tone lasts, and whether the change between segments is gradual or abrupt. Additionally, the start and end points of the pronunciation can be defined for each segment and used for the structural analysis of the voice signal.

[0118] The voice segment analysis unit (122) can be interconnected with the pitch analysis unit (121) to distinguish between gradual rising, extending, and gradual falling patterns of pitch. By combining the two analysis units (121, 122) to simultaneously analyze temporal persistence and frequency change rates, precise pitch classification is possible.

[0119] The stress analysis unit (123) analyzes changes in the Root Mean Square Energy (RMS) of the input voice signal to extract stress intensity information. The extracted stress intensity information is used as data to visually distinguish the parts emphasized in the pronunciation of each word or syllable. For example, by distinguishing between main stress and weak stress, sections with high pronunciation intensity can be visualized as emphasized symbols. In addition, it can be used as basic data for selecting content words by analyzing stress between syllables.

[0120] The pitch analysis unit (121), voice segment analysis unit (122), and stress analysis unit (123) operate in conjunction with each other and comprehensively analyze the tone, pronunciation length, and stress of the voice signal.

[0121] For example, the voice segment analysis unit (122) and the stress analysis unit (123) may be interconnected to perform segmentation (dividing into words and syllables) of the voice signal and extract content words. The segmentation separates the voice data on the time axis to form a basic structure in syllable units. In addition, the content words are words that have meaning, such as nouns or verbs, and are selected by excluding unnecessary prepositions or articles from the entire voice.

[0122]

[0123] FIG. 7 is a block diagram illustrating detailed components of a symbol attribute value determination unit according to an embodiment of the present invention. The symbol attribute value determination unit (140) determines an attribute value corresponding to each feature of a voice signal using the features of the voice signal or normalized data, and determines a visual symbol based on the symbolized value.

[0124] The symbol attribute value determining unit (140) may include a pitch symbol determining unit (141), a pronunciation time symbol determining unit (142), an accent symbol determining unit (143), and a character-specific symbol generation unit (144).

[0125] The pitch symbol determination unit (141) determines the direction, slope, or change in color and shade of the symbol based on pitch data extracted from the voice signal. By analyzing the pitch data, it can distinguish rising, falling, or unchanged pitches and convert them into symbol attribute values.

[0126] The height attribute value defined by the height symbol determining unit (141) can determine the baseline slope angle in response to the frequency value of the voice signal. The height attribute value may be determined as a baseline angle that is pre-set to correspond to the voice frequency, or may have an angle value proportional to the frequency. For example, a rising tone may be implemented at a 180-degree angle so that the shading is directed toward the top of the baseline.

[0127] The pronunciation duration symbol determination unit (142) analyzes the pronunciation duration and converts the duration value into an extension of the symbol's length or the number of repetitions. For example, a character with a long pronunciation visually emphasizes the pronunciation duration by overlapping multiple symbols or increasing the horizontal length of the symbol in proportion to the pronunciation duration.

[0128] The stress symbol determination unit (143) can analyze the stress extracted from the voice signal and determine the thickness, size, and shading intensity values ​​of the symbol according to the stress value. For example, the higher the stress, the thicker the symbol and the darker the shading, and the weaker the stress, the more the symbol is expressed as a relatively thin and light-colored symbol.

[0129] The character-specific symbol generation unit (144) integrates symbol attribute values ​​from the pitch symbol determination unit (141), pronunciation time symbol determination unit (142), and stress symbol determination unit (143) to generate character-specific attribute values ​​on an individual character basis. The character-specific attribute values ​​can be used as basic data for generating individual symbols by integrating pitch, pronunciation time, and stress attribute values ​​on a character basis.

[0130] The above symbol attribute value determination unit (140) performs a process of converting the phonetic characteristics of a voice signal into visual symbols through the mutual linkage of each component, thereby providing the user with an intuitive understanding of the complex changes in the voice signal.

[0131]

[0132] FIG. 8 is a flowchart illustrating the processing process of a visual encoding system for voice signals according to an embodiment of the present invention in steps.

[0133] In step (S110), the visual encoding system for the voice signal receives an analog voice signal and digitized voice data input by the user. The input voice signal and voice data are used as basic data to be analyzed and processed in subsequent steps.

[0134] In step (S120), the visual encoding system of the speech signal analyzes and extracts various features of the input speech signal. The extracted features include speech characteristics such as pitch, duration, and stress.

[0135] In step (S130), the visual encoding system of the voice signal determines a symbol attribute value based on the extracted voice signal feature data. The extracted voice feature data is quantified and analyzed to determine the symbol attribute value. Symbolized symbol attribute values ​​are assigned to correspond to voice feature data such as stress, pitch, and pronunciation duration. For example, stress may be converted into the size and thickness or color value of the symbol, pitch may be converted into the angle, direction, or position of the symbol, and pronunciation duration may be converted into the number or length of the symbol.

[0136] In step (S140), the visual encoding system of the voice signal performs visual encoding using the symbol attribute values ​​generated in step (S130) to construct a visual layout. In step (S140), the symbol attribute values ​​are visualized and arranged based on a time axis, and a multidimensional visual pattern is generated to support the user in simultaneously recognizing various voice feature changes.

[0137] In step (S150), the visual encoding system of the voice signal outputs visually represented visual encoding data to the user. At this time, the visual encoding result is provided visually using an interface with the user, but is not limited to this and supports various output methods. The output data can be utilized for various applications such as learning records, user pronunciation correction, intonation pattern analysis, and the provision of customized learning feedback.

[0138] In step (S160), the visual encoding system of the speech signal generates metadata containing information that quantifies the signal structure and features with respect to the symbol attribute values. The metadata summarizes the structure and features of the encoded result and may include the symbol attribute values ​​of each syllable, the rhythm of the entire sentence, intonation patterns, and stress distribution.

[0139] In step (S170), the visual encoding system of the voice signal provides a voice pattern similar to the user's pronunciation pattern using the generated metadata. Additionally, it can provide user-customized learning materials to enhance the user's pronunciation correction and learning effectiveness.

[0140]

[0141] FIG. 9 is a flowchart illustrating the detailed operation steps of a voice signal feature extraction unit according to an embodiment of the present invention. Here, the features of the voice signal can be expressed in the frequency domain, and the fundamental frequency of the voice can be selected to express the stress, pitch, and duration of pronunciation of the voice.

[0142] In step (S121), the voice signal feature extraction unit (120) collects a voice signal input from the voice signal input unit (110). The voice signal input unit (110) can digitize an analog voice signal input through a microphone or directly receive digital voice data, such as a voice file.

[0143] In step (S122), the voice signal feature extraction unit (120) performs voice recognition based on the input voice data. In step (S122), the word and sentence composition of the voice are analyzed.

[0144] In step (S123), the voice signal feature extraction unit (120) analyzes the entire spectrum of the input voice signal to extract the fundamental frequency (F0), which is the basic structure of the pronunciation. The fundamental frequency is used as basic data for analyzing the pitch of the voice. For example, the extraction of the fundamental frequency can be performed using frequency analysis techniques such as the Fast Fourier Transform (FFT) or the YIN algorithm.

[0145] In step (S1231), the voice signal feature extraction unit (120) calculates the entire voice segment based on the frequency and time domains of the voice frequency data. The voice segment is used as information to distinguish the time range in which words and syllables are pronounced.

[0146] In step (S1232), the voice signal feature extraction unit (120) analyzes the sound frequency band of the entire voice data based on the voice frequency data. The calculation of the sound frequency band distinguishes various frequency elements within the voice frequency data to clearly identify the characteristics of the voice signal.

[0147] In step (S124), the voice signal feature extraction unit (120) segments the recognized voice by word based on the voice. Step (S124) is a preprocessing step for separating voice data by word in a sentence and performing word-by-word processing.

[0148] In step (S125), the voice signal feature extraction unit (120) analyzes the data segmented by word to extract content words, which are meaningful words within a sentence based on the structure and meaning of the sentence. The content words are key words that have meaning and can be nouns, verbs, adjectives, adverbs, etc., and are key elements of stress and tone analysis.

[0149] In step (S1251), the voice signal feature extraction unit (120) extracts a syllable interval. The syllable interval is used as basic data to analyze pronunciation rhythm and flow by quantifying the time interval between words and syllables.

[0150] In step (S1252), the voice signal feature extraction unit (120) extracts the syllable frequency of each syllable. The syllable frequency is used to quantify the pitch of each syllable and to analyze the characteristics of intonation and pronunciation.

[0151] In step (S126), the voice signal feature extraction unit (120) extracts stress at the syllable level. The stress data is a value that quantifies the intensity of the pronunciation and can be converted into visual elements such as size, color, or boldness in visual symbolization.

[0152] In step (S127), the voice signal feature extraction unit (120) performs a quantitative calculation of normalized voice characteristics based on the data extracted from steps (S123), (S125), and (S126).

[0153]

[0154] FIG. 10 is a flowchart illustrating the detailed operation steps of a symbol attribute value determination unit according to an embodiment of the present invention.

[0155] In step (S141), the symbol attribute value determination unit (140) analyzes the pitch of the voice signal and determines the symbol attribute. Step (S141) analyzes changes in pitch and sets a slope angle or a shading position to represent pitch rise, fall, gradual rise, and gradual fall. For example, if the pitch value rises, a shading position is set at the top of the symbol. The determined pitch symbol attribute value is reflected in the character-unit symbolization in a subsequent step.

[0156] In step (S142), the symbol attribute value determining unit (140) analyzes the pronunciation duration of the voice signal and determines the associated symbol attribute value. It is quantified as the number or length of symbols corresponding to the pronunciation duration, and as the pronunciation duration increases, it can be visualized by repeating the symbols or expanding the symbols or shading in a horizontal direction. For example, if the pronunciation duration is 3 seconds, the length of the pronunciation can be emphasized by overlapping 3 identical symbols, or the length of the symbols can be expanded three times to express it.

[0157] In step (S143), the symbol attribute value determining unit (140) analyzes the stress of the voice signal and converts the intensity of the stress into the thickness, size, color, or intensity of the shading. The intensity of the stress is quantified into an attribute value of 0 to 1, and is classified into high stress, normal stress, and weak stress according to the quantified stress attribute value. For example, if the stress attribute value is high at 0.8, the shading can be set to be thick, the size increased, and the intensity intense. Also, if the stress attribute value is low at 0.3, the shading can be set to be thin and the size small, the shading not filled, or the intensity expressed lightly. Thus, the user can intuitively distinguish the intensity and change of the stress.

[0158] In step (S144), the symbol attribute value determining unit (140) combines the symbol attribute values ​​of pitch, pronunciation time, and stress determined in steps (S141 to S143) to generate a visual symbol in character units. For example, for a specific character, if the stress attribute value is 0.8, the pitch value is 300 Hz, and the pronunciation time is 2 seconds, the symbol is displayed in bold, the slope angle is set steeply in the upward direction, and the number of overlapping symbols reflecting the pronunciation time can be displayed as 2.

[0159] In conclusion, the pitch, pronunciation duration, and stress of the voice signal are converted into quantified symbolic attributes by step (S144) in step (S141), and the symbolic attributes are combined to generate a symbolic attribute value in character units. The generated symbolic attribute value is visualized by the visual encoding unit (150) and provided to the user, allowing the user to intuitively understand the complex characteristics of the voice signal.

[0160]

[0161] FIG. 11 is a reference diagram illustrating the sentence-unit encoding process of a voice signal visual encoding system according to an embodiment of the present invention in steps.

[0162] In step (S101), the voice signal visual encoding system (100) receives the sentence “I want to watch sports”. The sentence can be input in the form of a voice signal or text, and analysis begins on a sentence-by-sentence basis.

[0163] In steps (S1021, S1022, S1023), the voice signal visual encoding system (100) divides the sentence into words. Here, it is divided into “want” (S1021), “watch” (S1022), and “sports” (S1023), respectively.

[0164] In steps (S1031, S1032, S1033), the speech signal visual encoding system (100) analyzes each word in syllable units. “want” is analyzed as the syllable “a” (S1031), “watch” as the syllable “a” (S1032), and “sports” as the syllable “po” (S1033). In the above steps (S1031, S1032, S1033), data on the interval between syllables within a word and the duration of pronunciation are quantified by syllable recognition.

[0165] In steps (S1041, S1042, S1043), the voice signal visual encoding system (100) determines symbol attribute values ​​for the divided syllables based on pitch, duration, and stress. At this time, the pitch symbol determining unit (141), the duration symbol determining unit (142), and the stress symbol determining unit (143), which are detailed components of the symbol attribute value determining unit (140), are interconnected to perform the encoding operation.

[0166] Here, the attribute value “s” indicates general stress, the attribute value “ms” indicates main stress, the attribute value “r” indicates rise pitch, the attribute value “f” indicates fall pitch, and the attribute value “nc” indicates no change pitch. Additionally, the duration can be indicated by using the number of characters or a pre-set character inside parentheses, such as “(a)”.

[0167] In step (S1041), the syllable “a” of “want” is assigned a symbolic attribute value reflecting a normal stress and rising tone, such as “sr-(a).” In step (S1042), the syllable “a” of “watch” is assigned a symbolic attribute value reflecting a normal stress and rising tone, such as “sf-(a).” In step (S1043), the syllable “po” of “sports” is assigned a symbolic attribute value reflecting a main stress on “p” and a main stress and unchanging tone on “o,” such as “ms-(p)-ms-nc-(o).”

[0168] In step (S105), the voice signal visual encoding system (100) combines the symbol attribute values ​​for each syllable to finally generate a sentence-unit visual symbol. The sentence-unit visual symbol is “[<s-r-(a)><s-f-(a)><ms-(p)-ms-nc-(o)> It appears in an integrated form like ]”.

[0169] FIG. 11 illustrates, in steps, the process by which the voice signal visual symbolization system (100) assigns symbol attribute values ​​corresponding to pitch, pronunciation time, and stress based on data analyzed in sentence → word → syllable units. The result generated by the above process can be converted into an intuitive visual symbol and provided to the user.

[0170]

[0171] FIG. 12 is a reference diagram illustrating the visual encoding process of a voice signal according to an embodiment of the present invention.

[0172] FIG. 12 includes basic data of a speech signal (A2 and A3) and a result encoded based on said data (B2 and B3). A2 and A3 in FIG. 12 represent spectrograms and RMS energy data obtained by analyzing speech signals generated when pronouncing “Wow” and “I like playing sports,” respectively, and B2 and B3 in FIG. 12 illustrate the result of encoding the data of A2 and A3 and expressing it visually.

[0173] First, the frequency domain (A21) of A2 represents the temporal change in frequency that occurs when pronouncing “Wow” and can be used as basic data for analyzing intonation and pitch changes of the pronunciation. The RMS energy change domain (A23) of A2 analyzes the intensity and stress of the pronunciation over time and serves as a criterion for determining the duration of the pronunciation. Additionally, the time and word matching domain (A22) clearly indicates the location where each phoneme is pronounced and analyzes the timing of the pronunciation.

[0174] B2 in Fig. 12 represents the visual symbolization result generated based on the speech feature data extracted from A2. For example, “W” is represented with normal stress and an unchanging tone, “o” with main stress and a rising tone, and “ww” with normal stress and a gradually falling tone. Additionally, “www” is visualized in a manner where the symbols are continuously overlapping to reflect a long pronunciation duration.

[0175] Likewise, A3 in Fig. 12 is data showing a voice signal of “I like playing sports” pronounced, and B3 is the result of visualizing the data.

[0176] For example, the frequency domain of A3 represents tonal changes between pronunciations, and the RMS energy change domain analyzes the stress and pronunciation length of each word. In B3, the “i” in “like” represents normal stress and rising tone, the “pl” in “playing” and the “p” in “sports” are represented as normal stress and unchanging tone, and the “a” in “playing” and the “o” in “sports” are represented as major stress and unchanging tone. Additionally, the “I” and the remaining symbols are represented as weak stress and unchanging tone.

[0177] Figure 12 visually represents complex characteristics of a speech signal, such as stress, tone, and pronunciation duration, enabling the user to clearly understand changes in the speech signal and learn effectively.

[0178] The method according to the embodiments of the present invention described above may be implemented in the form of program instructions that can be executed through various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., either individually or in combination. The program instructions recorded on the computer-readable recording medium may be specially designed and configured for the embodiments of the present invention, or may be known and available to a person skilled in the art of computer software. The computer-readable recording medium includes hardware configured to store and execute program instructions, such as magnetic recording media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; ROMs; RAMs; and flash memory. Program instructions include machine code generated by a compiler and high-level language code that can be executed on a computer using an interpreter. The hardware may be configured to operate as one or more software modules to process the method according to the present invention, and vice versa.

[0179] The method according to an embodiment of the present invention can be executed in the form of program instructions on an electronic device. The electronic device includes portable communication devices such as smartphones or smartpads, computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, and home appliances.

[0180] The method according to an embodiment of the present invention may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable recording medium or online through an application store. In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created in a storage medium such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0181] Each component, such as a module or a program, according to an embodiment of the present invention may be composed of a single or multiple sub-components, and some of these sub-components may be omitted or additional sub-components may be included. Some components (modules or programs) may be integrated into a single entity and may perform the functions performed by each corresponding component prior to integration in the same or similar manner. Operations performed by a module, program, or other component according to an embodiment of the present invention may be executed sequentially, in parallel, iteratively, or heuristically, or at least some operations may be executed in a different order, omitted, or additional operations may be added.

[0182] The foregoing description of the present invention is for illustrative purposes only, and those skilled in the art will understand that other specific forms can be easily modified without altering the technical spirit or essential features of the present invention. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, each component described as a single unit may be implemented in a distributed manner, and components described as distributed may likewise be implemented in a combined form.

[0183] The scope of the present invention is defined by the claims set forth below, and all modifications or variations derived from the meaning and scope of the claims and equivalent concepts thereof should be interpreted as being included within the scope of the present invention.

[0184] The modes for carrying out the invention are described together in the best mode for carrying out the invention.

[0185] A system for visualizing speech signals for English learning and a method according to an embodiment of the present invention can provide a system that normalizes and patterns visually symbolized speech data and utilizes it as user-customized learning data.

[0186] In addition, the visual encoding system and method for speech signals for English learning according to an embodiment of the present invention can quantitatively analyze speech feature data and generate metadata to provide a platform that can be utilized in learning systems and various application fields.

Claims

1. In a system for visual encoding of speech signals for English learning, Voice signal input unit that receives a voice signal; A voice signal feature extraction unit that extracts feature data including stress, pitch, and duration from the input voice signal; A symbol attribute value determination unit that determines a symbol attribute value based on the above feature data; A visual encoding unit that expresses as a visual symbol based on the above-determined symbol attribute value; and An output unit that outputs the above visual symbol data to a user device; Includes, The above visual symbol is expressed as a combination of an accent shade that is displayed overlapping the basic character symbol as a shade having different thicknesses based on the accent size, and a reference line that cuts through the accent shade and crosses the center point of the basic character symbol at a predetermined angle. Speech signal visual encoding system.

2. In Paragraph 1, The angle of the above baseline is used to visually represent the pitch of the above tone and the change in the pitch of the above tone over time. Speech signal visual encoding system.

3. In Paragraph 2, The method further includes a normalization unit that converts the feature data of the extracted voice signal into normalized data. Speech signal visual encoding system.

4. In Paragraph 3, A metadata generation unit that generates data quantifying the structure and characteristics of the above-mentioned symbol data; further comprising Speech signal visual encoding system.

5. In Paragraph 4, The pattern providing unit further includes providing a similar pattern sentence based on the above-mentioned generated metadata. Speech signal visual encoding system.

6. In Paragraph 5, In the above symbol attribute value determination unit Regarding the above feature data Bullish trends are classified into general bullish (Stress), main bullish (Main Stress), and bearish / unbullish (Weak Stress / Unstress) patterns, and Changes in pitch and variations over time are classified into No Change, Rise, Gradual Rise, Extend, Gradual Fall, and Fall patterns. Speech signal visual encoding system.

7. In Paragraph 6, In the above symbol attribute value determination unit Determining the number of symbols as symbol attribute values ​​to correspond to the pronunciation duration Speech signal visual encoding system.

8. A method for a visual encoding system of speech signals for English learning to visual encode speech signals, (a) A step of receiving a voice signal; (b) a step of extracting feature data including stress, pitch, and duration of the voice signal from the input voice signal; (c) A step of determining a symbolic attribute value based on the above feature data; (d) a step of representing as a visual symbol based on the determined symbol attribute value above; and (e) a step of outputting the above visual symbol data to a user device; including, The above visual symbol is expressed as a combination of an accent shade that is displayed overlapping the basic character symbol as a shade having different thicknesses based on the accent size, and a reference line that cuts through the accent shade and crosses the center point of the basic character symbol at a predetermined angle. Method for visualizing speech signals.

9. In Paragraph 8, The angle of the above baseline is used to visually represent the pitch of the above tone and the change in the pitch of the above tone over time. Method for visualizing speech signals.

10. In Paragraph 9, The method further includes the step of converting the feature data of the extracted voice signal into normalized data. Method for visualizing speech signals.

11. In Paragraph 10, The method further includes the step of generating metadata that quantifies the structure and characteristics of the above-mentioned symbol data. Method for visualizing speech signals.

12. In Paragraph 11, The method further includes the step of providing a similar pattern sentence based on the above-mentioned generated metadata. Method for visualizing speech signals.

13. In Paragraph 12, In step (c) above Regarding the above feature data Bullish trends are classified into general bullish (Stress), main bullish (Main Stress), and bearish / unbullish (Weak Stress / Unstress) patterns, and Changes in pitch and variations over time are classified into No Change, Rise, Gradual Rise, Extend, Gradual Fall, and Fall patterns. Method for visualizing speech signals.

14. In Paragraph 13, In step (c) above Determining the number of symbols as symbol attribute values ​​to correspond to the pronunciation duration Method for visualizing speech signals.