Electronic apparatus for contructing database of pronunciation variations by type for speech recognition
Patent Information
- Application Number
- KR1020250013822
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-04
- Publication Date
- 2026-08-11
Smart Images

Figure PAT00001_ABST
Abstract
Description
Technology Field
[0001] The present disclosure relates to the operation of at least one electronic device for constructing training data for a speech recognition model, and more specifically, to the operation of constructing training data according to the type of pronunciation variation. Background Technology
[0002] - Project Name: 2024 Regional University Revitalization (Glocal University)-062
[0003] - Period: 2024.03.01~2025.02.28
[0005] Voice recognition technology has advanced dramatically and is now familiarly used in various aspects of our daily lives, such as smartphone voice assistants, in-vehicle voice control systems, and home smart devices, enhancing the value and quality of life.
[0006] In speech recognition technology, the areas where accuracy and performance currently need continuous improvement are, most notably, the speech recognition of children and individuals with speech disorders whose speech sounds are inaccurate.
[0007] Although the fields of children and speech impairment are difficult and challenging areas in terms of speech recognition technology, the impact and ripple effects on the lives of children and people with speech impairments who experience inconvenience in daily life or have difficulty communicating due to speech problems are significant, so there is a need for them to be researched and developed in the end.
[0008] One of the important factors in deep learning-based speech recognition technology is related to the quantity and quality of deep learning training data.
[0009] To enhance the performance of child speech recognition technology, it is necessary to first build rich and refined speech resources that meticulously reflect information on the speech development process from children of various age groups. Prior art literature
[0010] Published Patent Application No. 10-2022-0121183 The problem to be solved
[0011] The present disclosure provides a system or platform for building and providing a database for voice learning for normal and pathological children.
[0012] The present disclosure provides a system or platform that builds and provides a database for speech learning regarding pronunciation variations of various regions / levels.
[0013] The purposes of the present disclosure are not limited to those mentioned above, and other purposes and advantages of the present disclosure not mentioned may be understood from the following description and will be more clearly understood from the embodiments of the present disclosure. Furthermore, it will be readily apparent that the purposes and advantages of the present disclosure can be realized by the means and combinations thereof set forth in the claims. means of solving the problem
[0014] A method of operating at least one electronic device according to one embodiment of the present disclosure comprises the steps of: acquiring first audio data that matches a normal pronunciation for at least one text; acquiring second audio data that matches at least one type included in a pronunciation variant for the text; and registering first training data composed of the text and the first audio data for the normal pronunciation, and registering second training data composed of the text and the second audio data for the type included in the pronunciation variant, thereby constructing training data for training a speech recognition model.
[0015] The step of acquiring the second audio data above may acquire audio data including the voice of a speaker with a speech impediment with respect to the text.
[0016] The step of acquiring the second audio data may involve acquiring audio data of a voice that matches each of a plurality of dialects formed in each of a plurality of regions or cultures with respect to the text.
[0017] The method of operating the electronic device may further include the step of training a variant prediction model to predict whether it corresponds to a pronunciation variant and at least one of the type of the pronunciation variant, based on the text, the first audio data, and the second audio data.
[0018] In this case, the method of operation of the electronic device may include the steps of: inputting target audio data including a user's voice into the speech recognition model to obtain a target text; inputting the target audio data and the target text into the variant prediction model to identify whether the user's voice corresponds to a pronunciation variant; and, if the user's voice corresponds to a pronunciation variant, identifying the type of pronunciation variant corresponding to the user's voice according to the output of the variant prediction model.
[0019] At this time, the method of operation of the electronic device may further include the step of additionally registering the target text and the target audio data as training data for the identified type of pronunciation variant.
[0020] In this case, the additional registration step may identify the reliability of the output of the variant prediction model for the target audio data, and based on the reliability, additionally register the target text and the target audio data as training data for the identified type of pronunciation variant.
[0021] According to one embodiment of the present disclosure, in a system comprising a data building server and a speech recognition server, the speech recognition server comprises a speech recognition model and a variant prediction model. The data building server acquires first audio data that matches a normal pronunciation for at least one text, and acquires second audio data that matches at least one type included in a pronunciation variant for the text. The speech recognition server may train the speech recognition model for a normal pronunciation based on first training data composed of the text and the first audio data, train the speech recognition model for the type included in a pronunciation variant based on second training data composed of the text and the second audio data, and train the variant prediction model to predict whether it corresponds to a pronunciation variant and at least one of the types of pronunciation variants based on the text, the first audio data, and the second audio data. Effects of the invention
[0022] The method of operation of an electronic device or system according to the present disclosure constructs text and audio data (voice) as training data for each pronunciation type (normal pronunciation, or at least one type of pronunciation variant), thereby enabling a speech recognition model to perform speech recognition for texts of various pronunciation variants, including regional / cultural dialects, in addition to articulation disorders and speech sound disorders. Brief explanation of the drawing
[0023] FIG. 1 is a flowchart for explaining the operation of an electronic device according to one embodiment of the present disclosure, FIG. 2 is a diagram illustrating the operation of an electronic device according to an embodiment of the present disclosure, which collects audio data for each of various texts according to the type of pronunciation variation and constructs data according to the type of pronunciation variation. FIG. 3a is a block diagram for explaining the configuration of an electronic device according to one embodiment of the present disclosure, FIG. 3b is a block diagram illustrating the configuration and operation of a system including a data building server and a voice recognition server according to an embodiment of the present disclosure, and FIG. 4 is a diagram illustrating the process of an electronic device or system according to one embodiment of the present disclosure training a speech recognition model and a variant prediction model based on audio data according to various types of pronunciation variants. Specific details for implementing the invention
[0024] The embodiments described herein are subject to various modifications and may have various forms; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope of specific embodiments and should be understood to include various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In relation to the description of the drawings, similar reference numerals may be used for similar components.
[0025] In describing the present disclosure, if it is determined that a detailed description of related known functions or configurations could unnecessarily obscure the essence of the present disclosure, such detailed description is omitted.
[0026] Additionally, the following embodiments may be modified in various other forms, and the scope of the technical concept of the present disclosure is not limited to the following embodiments. Rather, these embodiments are provided to make the present disclosure more faithful and complete and to fully convey the technical concept of the present disclosure to those skilled in the art.
[0027] The terms used in this disclosure are used merely to describe specific embodiments and are not intended to limit the scope of the rights. The singular expression includes the plural expression unless the context clearly indicates otherwise.
[0028] In the present disclosure, expressions such as “have,” “may have,” “include,” or “may include” indicate the presence of such features (e.g., numerical values, functions, actions, or components such as parts) and do not exclude the presence of additional features.
[0029] In the present disclosure, expressions such as “A or B,” “at least one of A or / and B,” or “one or more of A or / and B” may include all possible combinations of items listed together. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” may refer to cases including (1) at least one A, (2) at least one B, or (3) both at least one A and at least one B.
[0030] Expressions such as "first," "second," "first," or "second" used in this disclosure may modify various components regardless of order and / or importance, and are used only to distinguish one component from another and do not limit said components.
[0031] Where it is stated that a component (e.g., Component 1) is "(operatively or communicatively) coupled with / to" or "connected to" another component (e.g., Component 2), it should be understood that the component may be directly connected to the other component or connected through the other component (e.g., Component 3).
[0032] On the other hand, when it is stated that a certain component (e.g., a first component) is "directly connected" or "directly coupled" to another component (e.g., a second component), it may be understood that no other component (e.g., a third component) exists between the certain component and the other component.
[0033] As used in this disclosure, the expression “configured to” may be replaced, depending on the context, with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.” The term “configured to” may not necessarily mean only “specifically designed to” in hardware.
[0034] Instead, in some situations, the expression “device configured to do something” may mean that the device is “able to do something” together with other devices or components. For example, the phrase “processor (130) configured (or set) to perform A, B, and C” may mean a dedicated processor (130) for performing the said operations (e.g., an embedded processor (130)), or a general-purpose processor (130) (e.g., a CPU or an application processor) capable of performing said operations by executing one or more software programs stored in a memory device.
[0035] In the embodiment, the 'module' or 'part' performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Additionally, a plurality of 'modules' or a plurality of 'parts' may be integrated into at least one module and implemented as at least one processor (130), except for the 'module' or 'part' that needs to be implemented in specific hardware.
[0036] Meanwhile, the various elements and areas in the drawings are depicted schematically. Accordingly, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.
[0037] Hereinafter, embodiments according to the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement them.
[0038] First, the electronic device according to the present disclosure may correspond to a device or system composed of at least one computer. The electronic device may be implemented as a server to build training data. For example, the electronic device may be implemented as a database management server. Alternatively, the electronic device (100) may be implemented as a terminal device such as a desktop PC, a laptop PC, a tablet PC, or a smartphone. Meanwhile, the operation of the electronic device described below may be performed through a plurality of electronic devices.
[0039] FIG. 1 is a flowchart for explaining the operation of an electronic device according to one embodiment of the present disclosure.
[0040] Referring to FIG. 1, the electronic device can acquire first audio data that matches normal pronunciation for at least one text (S110). The text may consist of at least one syllable, word, morpheme, keyword, phrase, sentence, paragraph, etc., but is not limited thereto.
[0041] The first audio data may refer to audio data containing a voice uttered by a speaker corresponding to a normal person (without a speech impediment) of the text. Additionally, the first audio data may include a voice with pronunciation corresponding to standard language.
[0042] Additionally, for the same text, the electronic device may acquire second audio data that matches at least one type included in the pronunciation variant (S120). The types of pronunciation variants may include the pronunciation of a speaker with a speech sound disorder or pronunciation disorder. In this case, the speech sound / pronunciation disorder may be classified into detailed types according to specific conditions such as hearing impairment, cleft palate, cerebral palsy, intellectual disability, etc.
[0043] Types of pronunciation variations may include the pronunciation of speakers using dialects from various regions or cultures. For example, an electronic device may acquire audio data of speech that matches each of multiple dialects formed in multiple regions or cultures.
[0044] The electronic device can acquire second audio data containing the voice of a speaker for each of the multiple types corresponding to pronunciation variants. For example, audio data containing the voice of a speaker using the dialect of region A, audio data containing the voice of a speaker using the dialect of region B, audio data containing the voice of a speaker with a hearing impairment, etc., can be acquired respectively.
[0045] In addition, the electronic device can construct training data for training a speech recognition model based on each of the above-described audio data (S130). Specifically, the electronic device can register first training data consisting of the text and first audio data for normal pronunciation, and register second training data consisting of the text and second audio data for at least one type included in the pronunciation variant.
[0046] A speech recognition model is an artificial intelligence model for converting audio data into text (Speech-to-Text), and may include, but is not limited to, an acoustic module for extracting acoustic features from audio data to identify individual syllables, and a language module for combining the identified individual syllables.
[0047] The speech recognition model may be included in an electronic device, or it may be included on a separate speech recognition server capable of communicating with the electronic device.
[0048] In this regard, FIG. 2 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure to collect audio data for each of various texts according to the type of pronunciation variation and to construct data according to the type of pronunciation variation.
[0049] Referring to FIG. 2, the electronic device can acquire audio data of speech corresponding to normal pronunciation as well as audio data of speech corresponding to various pronunciation variant types (various disorders, dialects, etc.) for each text (e.g., Text 1, 2, 3, …). Each text may be different keywords, phrases, sentences, etc., but is not limited thereto.
[0050] At this time, the electronic device can construct training data for training a speech recognition model by classifying the audio data of each text corresponding to normal pronunciation into one group (Group 1) and also grouping the audio data of each text according to the type of pronunciation variation (Group 2, 3, …).
[0051] In addition, each audio data constituting the training data can be linked with metadata such as gender and age, in addition to the speaker's pronunciation type (normal pronunciation, types of pronunciation variations, etc.), and the electronic device can support functions such as searching, classification, and downloading in addition to storing and managing audio data.
[0052] FIG. 3a is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present disclosure.
[0053] Referring to FIG. 3a, the electronic device (100) may include a memory (110), a communication interface (120), a processor (130) (130), etc.
[0054] The memory (110) stores various programs or data temporarily or non-temporarily and transmits the stored information to the processor (130) upon the call of the processor (130). Additionally, the memory (110) can store various information required for the operation, processing, or control operation of the processor (130) in an electronic format.
[0055] The memory (110) may include, for example, at least one of a main memory and an auxiliary memory. The main memory may be implemented using a semiconductor storage medium such as ROM and / or RAM. The ROM may include, for example, a conventional ROM, EPROM, EEPROM and / or MASK-ROM. The RAM may include, for example, a DRAM and / or SRAM. The auxiliary memory may be implemented using at least one storage medium capable of storing data permanently or semi-permanently, such as a flash memory device, an SD (Secure Digital) card, a solid state drive (SSD), a hard disk drive (HDD), an optical recording medium such as a magnetic drum, a compact disc (CD), a DVD, or a laser disc, a magnetic tape, a magneto-optical disc and / or a floppy disk.
[0056] Referring to FIG. 3, the memory (110) may include a speech recognition model (111), a variant prediction model (112), etc.
[0057] The speech recognition model (111) corresponds to an artificial intelligence model for recognizing speech included in audio data and converting it into text. The speech recognition model (111) may include at least one acoustic module that recognizes phonemes, syllables, and words, which are elements necessary for constructing a sentence, based on feature information obtained from audio data. Additionally, the speech recognition model (111) may include a language module that performs an operation to obtain words, sentences, etc., by reconstructing phonemes, syllables, and words for vector values output from the acoustic module. The speech recognition model (111) may verify whether the reconstructed words / sentences, etc., are accurate by comparing them with words in a pre-stored pronunciation dictionary.
[0058] The electronic device (100) can train a speech recognition model (111) based on training data constructed according to the embodiment of FIG. 1 described above.
[0059] The variant prediction model (112) is a model for identifying the type of pronunciation variant corresponding to a voice within audio data that is the subject of speech recognition. For example, the variant prediction model (112) can identify whether the voice within the audio data is normal pronunciation or corresponds to a certain type of pronunciation variant.
[0060] The variant prediction model (112) can estimate the type of input pronunciation variant together with the text generated as the audio data is recognized by the speech recognition model (111), in addition to the audio data.
[0061] To this end, the variant prediction model (112) may be trained based on audio data containing voices of various pronunciation types (normal pronunciation, pronunciation variant, etc.) that utter each of a plurality of texts. In one embodiment, the electronic device (100) may train the variant prediction model (112) to predict whether a pronunciation variant corresponds to a pronunciation variant and at least one of the types of pronunciation variants based on text, first audio data, and second audio data.
[0062] The communication interface (120) may include a wireless communication interface, a wired communication interface, or an input interface.
[0063] A wireless communication interface can perform communication with various external devices using wireless communication technology or mobile communication technology. Such wireless communication technologies may include, for example, Bluetooth, Bluetooth Low Energy, CAN communication, Wi-Fi, Wi-Fi Direct, ultrawide band (UWB), Zigbee, infrared data association (IrDA), or near field communication (NFC), and mobile communication technologies may include 3GPP, Wi-Max, LTE (Long Term Evolution), 5G, etc. The wireless communication interface may be implemented using an antenna, a communication chip, a substrate, etc., capable of transmitting electromagnetic waves to the outside or receiving electromagnetic waves transmitted from the outside.
[0064] A wired communication interface can communicate with various devices based on a wired communication network. Here, the wired communication network can be implemented using physical cables, such as, for example, a pair cable, a coaxial cable, a fiber optic cable, or an Ethernet cable. Additionally, the wired communication interface may be a Universal Serial Bus (USB) terminal, and may also be any one of the following interfaces: HDMI (High Definition Multimedia Interface), MHL (Mobile High-Definition Link), USB (Universal Serial Bus), DP (Display Port), Thunderbolt, VGA (Video Graphics Array) port, RGB port, D-SUB (D-subminiature), or DVI (Digital Visual Interface).
[0065] Depending on the embodiment, either the wireless communication interface or the wired communication interface may be omitted. Accordingly, the electronic device may include only a wireless communication interface or only a wired communication interface. Furthermore, the electronic device may be equipped with an integrated communication interface that supports both wireless connection via the wireless communication interface and wired connection via the wired communication interface.
[0066] The electronic device is not limited to including a single communication interface that performs a communication connection in one manner, but may include multiple communication interfaces that perform communication connections in multiple manners.
[0067] The electronic device (100) can communicate with various terminal devices, databases, service servers, etc. through a communication interface (120). For example, the electronic device (100) can obtain audio data containing the spoken voices of users regarding various texts by communicating with various terminal devices or voice recognition service servers that provide voice recognition services. Additionally, the electronic device (100) may store audio data containing various spoken voices in a database.
[0068] The processor (130) controls the overall operation of the electronic device. Specifically, the processor (130) is connected to the configuration of the electronic device including memory as described above, and can control the overall operation of the electronic device by executing at least one instruction stored in the memory as described above. In particular, the processor (130) can be implemented as a single processor as well as as a plurality of processors.
[0069] The processor (130) may be implemented in various ways. For example, one or more processors (130) may include one or more of a CPU (Central Processing Unit), GPU (Graphics Processing Unit), APU (Accelerated Processing Unit), MIC (Many Integrated Core), DSP (Digital Signal Processor), NPU (Neural Processing Unit), hardware accelerator, or machine learning accelerator. One or more processors (130) may control one or any combination of other components of an electronic device and may perform operations or data processing related to communication. One or more processors (130) may execute one or more programs or instructions stored in memory. For example, one or more processors (130) may perform a method according to one embodiment of the present disclosure by executing one or more instructions stored in memory.
[0070] In embodiments of the present disclosure, the processor (130) may mean a system-on-chip (SoC) in which one or more processors (130) and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, GPU, APU, MIC, DSP, NPU, hardware accelerator or machine learning accelerator, etc., but the embodiments of the present disclosure are not limited thereto.
[0071] When a method according to one embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one processor (130) or by a plurality of processors (130). For example, when a first operation, a second operation, and a third operation are performed by a method according to one embodiment, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first operation and the second operation may be performed by a first processor (e.g., a general-purpose processor) and the third operation may be performed by a second processor (e.g., an artificial intelligence dedicated processor).
[0072] One or more processors (130) may be implemented as a single-core processor including one core, or as one or more multicore processors including multiple cores (e.g., homogeneous multicore or heterogeneous multicore). When one or more processors (130) are implemented as multicore processors, each of the multiple cores included in the multicore processor may include internal processor memory such as on-chip memory, and a common cache shared by multiple cores may be included in the multicore processor (130). Additionally, each of the multiple cores included in the multicore processor (130) (or some of the multiple cores) may independently read and execute program instructions for implementing a method according to one embodiment of the present disclosure, or all (or some) of the multiple cores may be linked together to read and execute program instructions for implementing a method according to one embodiment of the present disclosure.
[0073] When a method according to one embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one of the plurality of cores included in a multi-core processor, or may be performed by a plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to one embodiment, the first operation, the second operation, and the third operation may all be performed by a first core included in a multi-core processor, or the first operation and the second operation may be performed by a first core included in a multi-core processor and the third operation may be performed by a second core included in a multi-core processor.
[0074] Referring to FIG. 3, the processor (130) can control a data construction module (131), a speech recognition management module (132), a variant prediction module (133), etc. These modules may correspond to functional unit modules implemented in software and / or hardware.
[0075] The data construction module (131) is a module for constructing training data based on audio data containing speech voices of users of various pronunciation types (normal pronunciation, pronunciation variant, etc.) for multiple texts. The data construction module (131) may also be structured to facilitate the management or updating of training data by grouping audio data by pronunciation type as shown in FIG. 2.
[0076] The speech recognition management module (132) is a module for controlling the training or utilization of the speech recognition model (111). The speech recognition management module (132) can train the speech recognition model (111) based on the aforementioned training data built by the data construction module (131), and can also perform speech recognition by loading the trained speech recognition model (111) and inputting at least one target audio data.
[0077] The variant prediction module (133) is a module for predicting the pronunciation type of target audio data that is the subject of speech recognition through the speech recognition model (111). The variant prediction module (133) can identify the pronunciation type by inputting text corresponding to the speech recognition result for the target audio data into the variant prediction module (112) together with the target audio data.
[0078] In one embodiment, an electronic device (100) can input target audio data including a user's voice into a speech recognition model (111) to obtain target text. Then, the electronic device (100) can input the target audio data and target text into a variant prediction model (112) to identify whether the user's voice corresponds to a pronunciation variant. In this regard, the variant prediction model (112) can output a probability value that the input target audio data corresponds to each pronunciation type (each type of normal pronunciation and pronunciation variant).
[0079] Here, when the user's voice is identified as corresponding to a pronunciation variant (e.g., when the probability of corresponding to normal pronunciation is below a threshold or the probability of corresponding to at least one pronunciation variant is above a threshold), the electronic device (100) can identify the type of pronunciation variant corresponding to the user's voice according to the output of the variant prediction model (112).
[0080] At this time, the electronic device (100) may additionally register target text and target audio data as training data for the identified type of pronunciation variant. The additionally registered training data may also be used for training the speech recognition model (111) and / or for training the variant prediction model (112).
[0081] As a related embodiment, new training data including the aforementioned target audio data, target text, and information on the pronunciation type of the target audio data may be additionally registered as training data after verification by an administrator or expert. To this end, the electronic device (100) may transmit the new training data to the administrator's or expert's terminal and additionally register the new training data on the premise that information indicating verification is received from the terminal.
[0082] In this regard, in one embodiment, the electronic device (100) identifies the reliability of the pronunciation type (normal pronunciation or type of pronunciation variant) output by the variant prediction model (112) into which the target audio data is input, and if the reliability is above a threshold, additional training data can be registered without verification by an expert or manager.
[0083] Specifically, the electronic device (100) can identify whether there exists any previously registered audio data with a similarity value greater than or equal to a first value by comparing the feature value (value of each vector) of the target audio data extracted by the speech recognition model (111) with at least one audio data constituting the previously registered training data. Here, the previously registered training data refers to training data that has already been used for training the speech recognition model (111) and the variant prediction model (112), respectively.
[0084] To this end, with each audio data constituting the previously registered training data constructed by pronunciation type, the average feature value of the audio data may be obtained by pronunciation type and by text. In this case, the electronic device (100) can calculate the similarity by comparing the average feature value of the previously registered audio data by pronunciation type with the feature value of the previously registered audio data for text matching the target text of the aforementioned target audio data. Here, the pronunciation type with the highest similarity of the average feature value (or above a certain value) is selected, and for the selected pronunciation type, the individual similarity can be calculated again for each previously registered audio data.
[0085] And, if there is audio data within the previously registered training data that has a similarity greater than or equal to a first value, the electronic device (100) can identify that the reliability of the output of the variant prediction model (112) for the target audio data is greater than or equal to a threshold.
[0086] On the other hand, if there is no audio data with a similarity greater than or equal to a first value within the previously registered training data, but there is audio data with a similarity greater than or equal to a second value (< first value), the electronic device (100) can calculate the reliability of the output of the variant prediction model (112) (for the target audio data) according to the number of audio data with a similarity greater than or equal to the second value. Here, the reliability may increase as the number of audio data with a similarity greater than or equal to the second value increases. In this case, depending on the number, the reliability may be greater than or equal to a threshold.
[0087] Meanwhile, if there is no audio data with a similarity of the second value or higher within the previously registered training data, the electronic device (100) can identify that the reliability of the output of the variant prediction model (112) for the target audio data is below the threshold.
[0088] Meanwhile, in another embodiment, the electronic device (100) may be implemented as a separate data building server that does not include a voice recognition model (111) and a variant prediction model (112).
[0089] FIG. 3b is a block diagram illustrating the configuration and operation of a system including a data building server and a voice recognition server according to one embodiment of the present disclosure.
[0090] Referring to FIG. 3b, the data building server (100-1) can build a training DB (101) based on audio data composed of speech voices of users of various pronunciation types who utter multiple texts. At this time, the training DB (101) may be stored directly on the data building server (100-1) or may be stored on at least one database managed by the data building server (100-1).
[0091] Referring to FIG. 3b, the speech recognition server (100-2) may include a speech recognition model (111) and a variant prediction model (112).
[0092] The speech recognition server (100-2) can train a speech recognition model (111) and a variant prediction model (112) based on individual text and audio data constituting a training DB (101) managed by a data construction server (100-1). Then, the speech recognition server (100-2) can perform speech recognition on at least one target audio data through the trained speech recognition model (111), and at the same time, input the recognized result (text) and the target audio data into the variant prediction model (112) to predict the pronunciation type.
[0093] Meanwhile, FIG. 4 is a diagram illustrating the process of training a speech recognition model and a variant prediction model based on audio data according to various types of pronunciation variants in an electronic device or system according to one embodiment of the present disclosure. The process of FIG. 4 may be performed by an electronic device (100) configured as in FIG. 3a, or it may be performed on a system (data construction server (100-1), speech recognition server (100-2)) constructed as in FIG. 3b.
[0094] For example, in the operation of the system of FIG. 3b, first, the data building server (100-1) can obtain first audio data that matches the normal pronunciation for at least one text, and for the text, can obtain second audio data that matches at least one type included in the pronunciation variant.
[0095] At this time, the speech recognition server (100-2) can train a speech recognition model (111) for normal pronunciation based on first training data consisting of the text and first audio data, and can also train a speech recognition model (111) for types included in pronunciation variants based on second training data consisting of the text and second audio data.
[0096] In this case, the speech recognition server (100-2) can train a variant prediction model (112) to predict whether a pronunciation variant corresponds to a pronunciation variant and at least one of the types of pronunciation variants based on the text, the first audio data, and the second audio data.
[0097] Meanwhile, the various embodiments described above may be implemented by combining two or more embodiments, provided that they do not conflict or contradict each other.
[0098] Meanwhile, according to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., downloadable app) may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0099] Meanwhile, computer instructions or computer programs for performing processing operations in the various embodiments of the present disclosure described above may be stored in a non-transitory computer-readable medium. When such computer instructions or computer programs stored in the non-transitory computer-readable medium are executed by the processor (130) of a specific device, the specific device described above performs processing operations according to the various embodiments described above.
[0100] A non-transient computer-readable medium refers to a medium that stores data semi-permanently and can be read by a device, unlike media that store data for a short period of time such as registers, caches, and memory. Specific examples of non-transient computer-readable media include CDs, DVDs, hard disks, Blu-ray discs, USBs, memory cards, and ROMs.
[0101] Although preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above. It is understood that various modifications can be made by those skilled in the art without departing from the essence of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure. Explanation of the symbols
[0102] 100: Electronic device 110: Memory 120: Communication interface 130: Processor 100-1: Data Construction Server 100-2: Speech Recognition Server
Claims
Claim 1 A method of operating at least one electronic device, comprising: a step of acquiring first audio data that matches a normal pronunciation for at least one text; a step of acquiring second audio data that matches at least one type included in a pronunciation variant for the text; and a step of registering first training data composed of the text and the first audio data for a normal pronunciation, and registering second training data composed of the text and the second audio data for a type included in a pronunciation variant, thereby constructing training data for training a speech recognition model. Claim 2 In claim 1, the step of acquiring the second audio data is a method of operation of an electronic device for acquiring audio data including the voice of a speaker with a speech impediment with respect to the text. Claim 3 In claim 1, the step of acquiring the second audio data is a method of operation of an electronic device for acquiring audio data of a voice that matches each of a plurality of dialects formed in a plurality of regions or a plurality of cultures with respect to the text. Claim 4 A method of operation of an electronic device according to claim 1, further comprising the step of training a variant prediction model to predict whether a pronunciation variant corresponds to a pronunciation variant and at least one of the type of the pronunciation variant based on the text, the first audio data and the second audio data. Claim 5 In claim 4, the method of operation of the electronic device comprises: a step of inputting target audio data including a user's voice into the voice recognition model to obtain a target text; a step of inputting the target audio data and the target text into the variant prediction model to identify whether the user's voice corresponds to a pronunciation variant; and a step of identifying the type of pronunciation variant corresponding to the user's voice according to the output of the variant prediction model when the user's voice corresponds to a pronunciation variant. Claim 6 In claim 5, the method of operation of the electronic device further comprises the step of additionally registering the target text and the target audio data as training data for the identified type of pronunciation variant. Claim 7 A method of operation of an electronic device according to claim 6, wherein the additional registration step involves identifying the reliability of the output of the variant prediction model for the target audio data, and, based on the reliability, additionally registering the target text and the target audio data as training data for the identified type of pronunciation variant. Claim 8 A method of operation of an electronic device comprising: a memory in which at least one instruction is stored; and a processor (130) that executes the instruction to perform the method of operation of claim 1. Claim 9 A non-transient computer-readable medium storing at least one instruction that is executed by a processor (130) of an electronic device to cause the electronic device to perform the method of operation of claim 1. Claim 10 A system comprising a data building server and a speech recognition server, wherein the speech recognition server comprises a speech recognition model and a variant prediction model; the data building server acquires first audio data that matches a normal pronunciation for at least one text, acquires second audio data that matches at least one type included in a pronunciation variant for the text, and the speech recognition server trains the speech recognition model for a normal pronunciation based on first training data composed of the text and the first audio data, trains the speech recognition model for the type included in a pronunciation variant based on second training data composed of the text and the second audio data, and trains the variant prediction model to predict whether it corresponds to a pronunciation variant and at least one of the types of pronunciation variants based on the text, the first audio data, and the second audio data.