Electronic device and control method for electronic device
Patent Information
- Application Number
- PCT/KR2024/000350
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-09
- Filing Date
- 2024-01-08
- Publication Date
- 2025-05-22
AI Technical Summary
Existing speech recognition technologies face limitations in accurately recognizing foreign words due to bias from context-dependent encoders and decoders, which fail to sufficiently reflect the context of the voice signal, leading to inaccurate recognition results.
The proposed solution involves dividing the encoder in the speech recognition model into a context-independent encoder and a context-dependent encoder, allowing for the separation of phoneme and subword sequences, with a text information acquisition module for spelling and entity name corrections based on phoneme sequences.
This approach improves speech recognition accuracy by focusing on phonemes to represent actual pronunciation and corrects subword sequences, enabling more accurate recognition of user intentions, especially for foreign words.
Smart Images

Figure KR2024000350_22052025_PF_FP_ABST
Abstract
Description
Electronic devices and methods for controlling electronic devices
[0001] The present disclosure relates to an electronic device and a method for controlling the electronic device, and more particularly, to an electronic device capable of obtaining text information corresponding to a voice signal and a method for controlling the same.
[0002] Recently, with the advancement of technologies related to artificial intelligence (AI), the development of technologies for performing accurate voice recognition of the user's voice and obtaining text information that matches the user's speech intention is accelerating.
[0003] However, among conventional technologies, there are technologies that can consider the context of speech signals received before a specific point in time, or the context of the entire speech signal received before and after a specific point in time, to ensure that speech recognition models (e.g., automatic speech recognition models, ASR models) faithfully reflect the linguistic information contained in the speech signal. However, these conventional technologies utilize context-dependent encoders and decoders, which can be strongly biased by previous words, and thus have limitations in that accurate recognition may not be achieved, especially for foreign words.
[0004] Meanwhile, among the conventional technologies, there is a technology that performs voice recognition using a limited section of the voice signal. However, according to this conventional technology, there is a limitation that recognition results that do not match the user's speech intention may be obtained because the context of the voice signal is not sufficiently reflected.
[0005] The present disclosure is intended to overcome the limitations of the prior art as described above, and an object of the present disclosure is to provide an electronic device and a control method thereof capable of improving the accuracy of speech recognition by dividing an encoder included in a speech recognition model into a context-independent encoder and a context-dependent encoder.
[0006] According to one embodiment of the present disclosure for achieving the above-described object, an electronic device includes a memory storing at least one instruction and at least one processor executing the at least one instruction, wherein when a voice signal is acquired, the at least one processor inputs the voice signal to a common encoder to acquire a first vector corresponding to each of a plurality of sections of the voice signal, inputs the first vector to a first individual encoder to acquire a second vector corresponding to each of the plurality of sections and independent of the context of the voice signal, inputs the second vector to a first decoder to acquire a phoneme sequence corresponding to the second vector, inputs the first vector to a second individual encoder to acquire a third vector corresponding to at least two sections among the plurality of sections and dependent on the context of the voice signal, inputs the third vector to a second decoder to acquire a sub-word sequence corresponding to the third vector, and, through a text information acquisition module, based on the phoneme sequence. By correcting the above subword sequence, text information corresponding to the above multiple sections is obtained.
[0007] Meanwhile, the text information acquisition module includes a spell correction module, and the spell correction module can acquire the text information by correcting the spelling of the subword sequence based on the phoneme sequence when the subword sequence is identified as being in violation of a specific spelling.
[0008] Meanwhile, the text information acquisition module includes a named entity correction module, and the named entity correction module can acquire the text information by correcting the named entity of the subword sequence based on the phoneme sequence when the subword sequence is identified as not being included in a specified plurality of named entities.
[0009] Meanwhile, the common encoder can be trained to obtain the first vector suitable for both the first individual encoder and the second individual encoder without any specific constraints.
[0010] Meanwhile, the first individual encoder may be trained to obtain the second vector representing the characteristics of the phoneme sequence based on unlabeled training data, and the first decoder may be trained to obtain the phoneme sequence based on labeled training data.
[0011] Meanwhile, the second individual encoder may be trained to obtain the third vector representing the characteristics of the subword sequence based on unlabeled training data, and the second decoder may be trained to obtain the subword sequence based on labeled training data.
[0012] Meanwhile, at least two of the plurality of sections may include all sections received before a specific point in time among the plurality of sections or all sections received before and after the specific point in time among the plurality of sections.
[0013] According to one embodiment of the present disclosure for achieving the above-described object, a method for controlling an electronic device includes the steps of: when a voice signal is acquired, inputting the voice signal to a common encoder to acquire a first vector corresponding to each of a plurality of sections of the voice signal; inputting the first vector to a first individual encoder to acquire a second vector corresponding to each of the plurality of sections and independent of the context of the voice signal; inputting the second vector to a first decoder to acquire a phoneme sequence corresponding to the second vector; inputting the first vector to a second individual encoder to acquire a third vector corresponding to at least two sections among the plurality of sections and dependent on the context of the voice signal; inputting the third vector to a second decoder to acquire a sub-word sequence corresponding to the third vector; and obtaining text information corresponding to the plurality of sections by correcting the sub-word sequence based on the phoneme sequence through a text information acquisition module.
[0014] Meanwhile, if the subword sequence is identified as being in violation of a specific spelling, the step of obtaining the text information by correcting the spelling of the subword sequence based on the phoneme sequence may be further included.
[0015] Meanwhile, if the subword sequence is identified as not being included in a specific plurality of entity names, the step of obtaining the text information may further include correcting the entity name of the subword sequence based on the phoneme sequence.
[0016] Meanwhile, the step of learning to obtain the first vector suitable for both the first individual encoder and the second individual encoder without any specific constraints may be further included.
[0017] Meanwhile, the method may further include a step of learning to obtain the second vector representing the characteristics of the phoneme sequence based on unlabeled learning data and a step of learning to obtain the phoneme sequence based on labeled learning data.
[0018] Meanwhile, the method may further include a step of learning to obtain the third vector representing the characteristics of the subword sequence based on unlabeled learning data, and a step of learning to obtain the subword sequence based on labeled learning data.
[0019] Meanwhile, at least two of the plurality of sections may include all sections received before a specific point in time among the plurality of sections or all sections received before and after the specific point in time among the plurality of sections.
[0020] According to one embodiment of the present disclosure for achieving the above-described object, there is provided a non-transitory computer-readable recording medium including a program for executing a control method of an electronic device, wherein the control method of the electronic device comprises the steps of: when a voice signal is acquired, inputting the voice signal to a common encoder to acquire a first vector corresponding to each of a plurality of sections of the voice signal; inputting the first vector to a first individual encoder to acquire a second vector corresponding to each of the plurality of sections and independent of the context of the voice signal; inputting the second vector to a first decoder to acquire a phoneme sequence corresponding to the second vector; inputting the first vector to a second individual encoder to acquire a third vector corresponding to at least two sections among the plurality of sections and dependent on the context of the voice signal; inputting the third vector to a second decoder to acquire a sub-word sequence corresponding to the third vector; and inputting the phoneme sequence and the sub-word sequence to a text information acquisition module to acquire a text information sequence corresponding to the plurality of sections. It includes a step of obtaining corresponding text information.
[0021] Other aspects, features and advantages of one or more embodiments according to the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings.
[0022] FIG. 1 is a block diagram briefly showing the configuration of an electronic device according to an embodiment of the present disclosure;
[0023] FIG. 2 is a diagram showing a plurality of modules and input / output data for each of the plurality of modules according to one embodiment of the present disclosure;
[0024] FIG. 3 is a block diagram showing detailed modules of a text information acquisition module according to an embodiment of the present disclosure;
[0025] FIG. 4 is a flowchart for explaining a control method according to an embodiment of the present disclosure;
[0026] FIG. 5 is a diagram for explaining a plurality of sections of a voice signal according to an embodiment of the present disclosure;
[0027] FIG. 6 is a block diagram showing in detail the configuration of an electronic device according to an embodiment of the present disclosure, and
[0028] FIG. 7 is a flowchart illustrating a control method of an electronic device according to an embodiment of the present disclosure.
[0029] The present embodiments may be modified and have various embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope to specific embodiments, but should be understood to encompass various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In connection with the description of the drawings, similar reference numerals may be used for similar components.
[0030] In describing the present disclosure, if it is determined that a specific description of a related known function or configuration may unnecessarily obscure the gist of the present disclosure, a detailed description thereof will be omitted.
[0031] Additionally, the following embodiments may be modified in various other forms, and the scope of the technical concepts of the present disclosure is not limited to the following embodiments. Rather, these embodiments are provided to further faithfully and completely convey the technical concepts of the present disclosure to those skilled in the art.
[0032] The terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the scope of the rights. Singular expressions include plural expressions unless the context clearly dictates otherwise.
[0033] In this disclosure, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a corresponding feature (e.g., a component such as a number, function, operation, or part), and do not exclude the presence of additional features.
[0034] In this disclosure, expressions such as “A or B,” “at least one of A and / or B,” or “one or more of A or / and B” can include all possible combinations of the listed items. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” can all refer to (1) including at least one A, (2) including at least one B, or (3) including both at least one A and at least one B.
[0035] The expressions “first,” “second,” “first,” or “second,” etc., used in this disclosure can describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.
[0036] When it is said that a component (e.g., a first component) is “(operatively or communicatively) coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that said component may be directly coupled to said other component, or may be coupled via another component (e.g., a third component).
[0037] On the other hand, when it is said that a component (e.g., a first component) is "directly connected" or "directly connected" to another component (e.g., a second component), it can be understood that no other component (e.g., a third component) exists between said component and said other component.
[0038] The expression "configured to" as used in the present disclosure may be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" may not necessarily mean only "specifically designed to" in terms of hardware.
[0039] Instead, in some contexts, the phrase "a device configured to" may mean that the device, in conjunction with other devices or components, is "capable of" performing the operations described herein. For example, the phrase "a processor configured (or set) to perform A, B, and C" may refer to a dedicated processor (e.g., an embedded processor) for performing the operations, or a general-purpose processor (e.g., a CPU or application processor) that can perform the operations by executing one or more software programs stored in a memory device. The term "processor" may encompass various processing circuits (as used herein, including in the claims, the term "processor" may encompass various processing circuits, including at least one processor, wherein one or more of the at least one processor is configured to perform the various functions described herein).
[0040] In the embodiment, a 'module' or 'part' performs at least one function or operation, and may be implemented by hardware or software, or by a combination of hardware and software. In addition, a plurality of 'modules' or a plurality of 'parts' may be integrated into at least one module and implemented by at least one processor, except for a 'module' or 'part' that needs to be implemented by a specific hardware, and may include executable program instructions that can be executed by one or more processors among the at least one processor.
[0041] Meanwhile, the various elements and areas in the drawings are schematically drawn. Therefore, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.
[0042] Hereinafter, with reference to the attached drawings, embodiments according to the present disclosure will be described in detail so that a person having ordinary knowledge in the technical field to which the present disclosure pertains can easily implement the present disclosure.
[0043] FIG. 1 is a block diagram schematically illustrating the configuration of an electronic device (100) according to an embodiment of the present disclosure. FIG. 2 is a diagram illustrating a plurality of modules and input / output data for each of the plurality of modules according to an embodiment of the present disclosure. Hereinafter, various embodiments of the present disclosure will be described with reference to FIGS. 1 and 2 together.
[0044] The electronic device (100) according to the present disclosure refers to a device capable of acquiring text information corresponding to a voice signal. Specifically, the electronic device (100) can acquire text information corresponding to a voice signal using a neural network model for performing voice recognition, and in particular, can perform corrections, such as spelling correction, during the process of acquiring the text information.
[0045] The electronic device (100) according to the present disclosure may be implemented as a server including a neural network model for performing voice recognition, but is not limited thereto. The electronic device (100) may be implemented as a user terminal such as a smartphone, tablet PC, etc., or may be implemented as an edge computing device. That is, there are no particular limitations on the type of the electronic device (100) according to the present disclosure.
[0046] At least one instruction regarding the electronic device (100) may be stored in the memory (110). In addition, an O / S (Operating System) for driving the electronic device (100) may be stored in the memory (110). In addition, various software programs or applications for operating the electronic device (100) according to various embodiments of the present disclosure may be stored in the memory (110). In addition, the memory (110) may include a semiconductor memory such as a flash memory or a magnetic storage medium such as a hard disk.
[0047] Specifically, the memory (110) may store various software modules for operating the electronic device (100) according to various embodiments of the present disclosure, and the processor (120) may control the operation of the electronic device (100) by executing the various software modules stored in the memory (110). That is, the memory (110) is accessed by the processor (120), and data reading / recording / modifying / deleting / updating, etc., may be performed by the processor (120).
[0048] Meanwhile, in the present disclosure, the term memory (110) may be used to mean a memory (110), a ROM (not shown), a RAM (not shown) in a processor (120), or a memory card (not shown) (e.g., a micro SD card, a memory stick) mounted on an electronic device (100).
[0049] In particular, in various embodiments according to the present disclosure, the memory (110) may store a voice signal, a first vector, a second vector, a third vector, a phoneme sequence, a sub-word sequence, text information, etc. according to the present disclosure. In addition, the memory (110) may store data for implementing a plurality of modules according to the present disclosure and data for training a plurality of modules.
[0050] In addition, various information necessary within the scope of achieving the purpose of the present disclosure may be stored in the memory (110), and the information stored in the memory (110) may be updated as received from an external device or input by a user.
[0051] The processor (120) may include various processing circuits (as used herein, including in the claims, the term “processor” may include various processing circuits, including at least one processor, wherein one or more of the at least one processor may be configured to perform various functions described herein) and controls the overall operation of the electronic device (100). Specifically, the processor (120) is connected to a configuration of the electronic device (100) that includes a memory (110) and may control the overall operation of the electronic device (100) by executing at least one instruction stored in the memory (110) as described above.
[0052] The processor (120) may be implemented in various ways. For example, the processor (120) may be implemented as at least one of an application-specific integrated circuit (ASIC), an embedded processor, a microprocessor, hardware control logic, a hardware finite state machine (FSM), and a digital signal processor (DSP). Meanwhile, the term "processor (120)" in the present disclosure may be used to mean a central processing unit (CPU), a graphic processing unit (GPU), and a microprocessor unit (MPU).
[0053] In particular, in various embodiments according to the present disclosure, the processor (120) may obtain text information corresponding to a voice signal using a plurality of modules. The plurality of modules according to the present disclosure may include a common encoder (210), a first individual encoder (220), a second individual encoder (230), a first decoder (240), a second decoder (250), and a text information acquisition module (260), as illustrated in FIG. 2.
[0054] The plurality of modules may be implemented as software modules (e.g., including various executable program instructions) and / or hardware modules, and at least some of the plurality of modules may be integrated into one module, or other modules may be included in the plurality of modules in addition to the modules illustrated in FIG. 2. That is, the term 'module' used in this document may include various circuits (including processing circuits) and / or executable program instructions. In addition, the plurality of modules may include neural networks, and there is no particular limitation on the type of neural networks included in the plurality of modules. The overall configuration including the plurality of modules may be referred to as a neural network model or an automatic speech recognition model (ASR model), and it is also possible to implement a form in which some of the plurality of modules are included in one neural network model and other modules are included in another neural network model.
[0055] Meanwhile, the following description will assume that all of the multiple modules are included in the electronic device (100) as on-devices, but the present disclosure is not limited thereto. That is, it goes without saying that the embodiments according to the present disclosure can be implemented even when some of the multiple modules are included in the electronic device (100) and others are included in an external device. The following will describe in detail the process by which the processor (120) implements various embodiments according to the present disclosure using the multiple modules.
[0056] The processor (120) can acquire a voice signal. Specifically, the processor (120) may acquire a voice signal by receiving a voice signal through a microphone included in the electronic device (100), or may acquire a voice signal by receiving a voice signal from an external device through the communication unit (130). In the example of FIG. 2, it is assumed that the voice signal is a voice signal acquired when the user utters the word "Bhukansan."
[0057] When a speech signal is acquired, the processor (120) can input the speech signal to a common encoder (210) to acquire a first vector corresponding to each of a plurality of sections of the speech signal. In the present disclosure, the common encoder (210) refers to a module that can input and acquire a first vector corresponding to each of a plurality of sections of the speech signal. The common encoder (210) may be replaced with the term “shared encoder” in that it is for both the first individual encoder (220) and the second individual encoder (230). The common encoder (210) can be trained to acquire a first vector suitable for both the first individual encoder (220) and the second individual encoder (230) described below without any preset (e.g., specific) constraints.
[0058] That is, the common encoder (210) must be trained to obtain a first vector to be passed on to both the context-independent first individual encoder (220) and the context-dependent second individual encoder (230), and thus, it can be said to be a module that should not be limited by losses due to pre-learning. The first vector can be said to be a hidden representation obtained based on a limited context of a limited section of a speech signal.
[0059] The processor (120) can input the first vector into the first individual encoder (220) to obtain a second vector corresponding to each of the plurality of sections and independent of the context of the voice signal.
[0060] In the present disclosure, the first individual encoder (220) refers to a module capable of obtaining a second vector independent of the context of a speech signal. The first individual encoder (220) can be trained to obtain a second vector representing the characteristics of a phoneme sequence based on unlabeled training data.
[0061] Specifically, when a first vector corresponding to each of a plurality of sections is received from the common encoder (210), the first individual encoder (220) can obtain a second vector independent of the context based on the speech signal limited to each of the plurality of sections, and transmit the second vector to the first decoder (240). Therefore, the first individual encoder (220) can be referred to as a context-independent encoder. The second vector is not dependent on the context of the speech signal in that it is obtained based on a limited section of the speech signal, and can be referred to as hidden phoneme representations in that it is used to obtain a phoneme sequence as described below.
[0062] The processor (120) can input the second vector into the first decoder (240) to obtain a phoneme sequence corresponding to the second vector. In the present disclosure, the first decoder (240) refers to a module capable of obtaining a phoneme sequence corresponding to the second vector. The first decoder (240) can be trained to obtain a phoneme sequence based on labeled training data.
[0063] Specifically, when a second vector is received from the first individual encoder (220), the first decoder (240) can obtain a phoneme sequence corresponding to the second vector and transmit the phoneme sequence to the text information acquisition module (260). Since the phoneme sequence is obtained based on the second vector that is independent of the context of the speech signal, the phoneme sequence can also be said to be independent of the context of the speech signal. Therefore, the first decoder (240) can be referred to as a context-independent decoder. Meanwhile, since the first vector is obtained for each of a plurality of sections of the speech signal, the second vector is also obtained for each of a plurality of sections of the speech signal, but the phoneme sequence is a synthesis of the phonemes for each of the plurality of sections and corresponds to the entire section of the speech signal.
[0064] As in the example of Fig. 2, the phoneme sequence acquired by the first decoder (240) may be [ / bh / / u / / k / / a / / n / / s / / a / / n / ]. Each phoneme included in the phoneme sequence is expressed in a format such as / phoneme / . Since the phoneme sequence is independent of context, it can be said to be information that clearly expresses the pronunciation of a speech signal.
[0065] Meanwhile, the processor (120) can input the first vector into the second individual encoder (230) to obtain a third vector that corresponds to at least two sections among the plurality of sections and is dependent on the context of the speech signal. In the present disclosure, the second individual encoder (230) refers to a module that can obtain a second vector that is dependent on the context of the speech signal. The second individual encoder (230) can be trained to obtain a third vector that represents the characteristics of the subword sequence based on unlabeled training data.
[0066] Specifically, when a first vector corresponding to each of a plurality of sections is received from the common encoder (210), the second individual encoder (230) can obtain a context-dependent third vector based on unrestricted speech signals of at least two sections among the plurality of sections, and transmit the third vector to the second decoder (250).
[0067] At least two sections among the plurality of sections may be the entire section received before a specific point in time among the plurality of sections, or the entire section received before and after a specific point in time among the plurality of sections. That is, when the second individual encoder (230) inputs the first vector corresponding to each of the limited sections of the speech signal, it may obtain a third vector including a left-only context commonly used in a streaming speech recognition model (automatic speech recognition, ASR), or obtain a third vector including a full context commonly used in a full context ASR model (or causal ASR). Meanwhile, the length of the section or interval according to the present disclosure may be determined differently depending on the settings of the developer or user. The plurality of sections will be described in more detail with reference to FIG. 5.
[0068] The third vector may be context-dependent in that it is obtained based on the speech signal before or after a specific point in time, without being limited to each of the multiple segments of the speech signal. Therefore, the second individual encoder (230) may be referred to as a context-dependent encoder, and the third vector may be referred to as a hidden sub-word representation obtained based on an unrestricted segment of the speech signal.
[0069] The processor (120) can input the third vector into the second decoder (250) to obtain a subword sequence corresponding to the third vector. In the present disclosure, the second decoder (250) refers to a module that can obtain a subword sequence corresponding to the third vector. The second decoder (250) can be trained to obtain a subword sequence based on labeled training data. In the present disclosure, a subword sequence refers to a set of subwords that are sequentially connected, and a subword refers to a component that represents a smaller unit of meaning included in the meaning of a word. Depending on the embodiment, the second decoder may be implemented to obtain a word sequence in which words themselves are sequentially connected. However, using a subword sequence instead of a word sequence can alleviate problems with processing words that do not exist in the training data (e.g., Out-Of-Vocabulary (OoV) or Unknown Token (UNK)) or new words.
[0070] Specifically, when the third vector is received from the second individual encoder (230), the second decoder (250) can obtain a subword sequence corresponding to the third vector and transmit the subword sequence to the text information acquisition module (260). Since the subword sequence is obtained based on the third vector that is dependent on the context of the speech signal, the subword sequence can also be said to be dependent on the context of the speech signal. Therefore, the second decoder (250) can be referred to as a context-dependent decoder. Meanwhile, even if the third vector is not for the entirety of each of the plurality of sections of the speech signal, the subword sequence corresponds to the entire section of the speech signal as a composite of subwords for some sections.
[0071] As in the example of Fig. 2, the subword sequence obtained by the second decoder (250) may be [Book on sun]. That is, since the subword sequence obtained by the second decoder (250) is not limited to a specific section of the speech signal and is obtained by considering the context of the speech signal, unlike the phoneme sequence obtained by the first decoder (240), the linguistic information included in the speech signal can be clarified.
[0072] The processor (120) can acquire text information corresponding to multiple sections (e.g., multiple entire sections) by correcting subword sequences based on phoneme sequences through the text information acquisition module (260). In the present disclosure, the text information acquisition module (260) refers to a module capable of acquiring text information corresponding to the entire speech signal based on phoneme sequences and subword sequences.
[0073] Specifically, the text information acquisition module (260) identifies whether a subword sequence needs to be corrected, and if it is determined that the subword sequence needs to be corrected, it can acquire text information by correcting the subword sequence based on the phoneme sequence. On the other hand, if it is determined that the subword sequence does not need to be corrected, the text information acquisition module (260) can acquire text information based on the subword sequence without correcting the subword sequence. The process of identifying whether a subword sequence needs to be corrected will be described in more detail with reference to FIGS. 3 and 4.
[0074] As described above, the phoneme sequence acquired by the context-independent first decoder (240) can clearly express the pronunciation of the speech signal, and the subword sequence acquired by the context-dependent second decoder (250) can clearly express the linguistic information included in the speech signal. The text information acquisition module (260) can be trained to acquire text information that can clearly express the linguistic information included in the speech signal while clearly expressing the pronunciation of the speech signal. Meanwhile, since both the phoneme sequence and the subword sequence correspond to the entire section of the speech signal, the text information also corresponds to the entire section of the speech signal.
[0075] Specifically, the text information acquisition module (260) can acquire accurate text information corresponding to the user's speech by correcting the subword sequence based on the phoneme sequence. As in the example of Fig. 2, the text information acquired by the text information acquisition module (260) may be [Bhukansan]. That is, in the example of Fig. 2, the subword sequence is [Book on sun], but by correcting it using the phoneme sequence [ / bh / / u / / k / / a / / n / / s / / a / / n / ], the text information [Bhukansan] that matches the word spoken by the user can be acquired. The process of correcting the subword sequence through the text information acquisition module (260) will be described in more detail with reference to Figs. 3 and 4.
[0076] Meanwhile, when text information is acquired as described above, the acquired text information may be provided to the user as is, but may also be used to control the operation of the electronic device or external device through matching with a command for controlling the operation of the electronic device or external device.
[0077] Meanwhile, the above described embodiment of obtaining a phoneme sequence through the first decoder (240) may be used, but depending on the embodiment, a grapheme sequence may be obtained instead of a phoneme sequence. Here, a grapheme sequence refers to an individual letter or a group of letters representing one phoneme. In addition, the above described embodiment of obtaining a subword sequence through the second decoder (250) may be used, but it goes without saying that any lower-level concept constituting a word may be a subword according to the present disclosure.
[0078] For example, when acquiring text information for a speech signal, an ASR model may be trained not by using a limited section of the speech signal, but by using the entire section received before a specific point in time, or the entire section of the speech signal received before and after a specific point in time among multiple sections. In other words, since the ASR model according to the prior art utilizes a context-dependent encoder and decoder, it may be strongly biased by previous words, and thus, accurate recognition may not be achieved, especially in the case of foreign words.
[0079] On the other hand, according to the present disclosure, the electronic device (100) can train the context-independent encoder to acquire a phoneme sequence representing actual pronunciation by focusing on individual phonemes without biasing linguistic information using a speech signal of a limited section by dividing the encoder into a context-independent encoder and a context-dependent encoder. Accordingly, the electronic device (100) can perform spelling correction or named entity correction, etc. based on the phoneme sequence, thereby performing accurate speech recognition results for the speech signal.
[0080] FIG. 3 is a block diagram illustrating detailed modules of a text information acquisition module (260) according to an embodiment of the present disclosure. FIG. 4 is a flowchart illustrating a control method according to an embodiment of the present disclosure. Hereinafter, various embodiments of the present disclosure will be described with reference to FIGS. 3 and 4 together.
[0081] A text information acquisition module (260) according to an embodiment of the present disclosure may include at least one of a spelling correction module (261) and a named entity correction module (262). Each 'module' may include, for example, various circuits including various processing circuits and / or executable program instructions. That is, the text information acquisition module (260) may include only one module among the spelling correction module (261) and the named entity correction module (262), but the following description will be made on the premise that the text information acquisition module (260) includes both the spelling correction module (261) and the named entity correction module (262).
[0082] As described above with reference to FIGS. 1, 2, and 4, the processor (120) can obtain text information corresponding to multiple sections of a speech signal by correcting a subword sequence based on a phoneme sequence (S410).
[0083] The spelling correction module (261) refers to a module that can correct misspellings (or spellings) included in a text sequence to obtain a correct text sequence. Specifically, the spelling correction module (261) can identify whether a subword sequence violates a predefined spelling (S420). If the subword sequence is identified as violating the predefined spelling (S420-Y), the spelling correction module (261) can obtain text information by correcting the spelling of the subword sequence based on a phoneme sequence (S430). Conversely, if the subword sequence is identified as not violating the predefined spelling (S420-N), the spelling correction module (261) may not correct the spelling of the subword sequence, and the next step can be performed by the named entity correction module (262).
[0084] In one embodiment, the spelling correction module (261) may include a phoneme encoder for encoding a phoneme sequence, a sub-word encoder for encoding a sub-word sequence, and an attention decoder for outputting a correct sub-word sequence based on the output of the phoneme encoder and the sub-word sequence.
[0085] The entity name correction module (262) refers to a module that can correct incorrect entity names included in a text sequence to obtain a correct text sequence. Specifically, the entity name correction module (262) can identify an entity corresponding to a text sequence by using information on a plurality of entity names, which are words corresponding to pre-defined people, companies, places, times, units, etc. For example, referring to FIG. 4, the entity name correction module (262) can identify whether a subword sequence is included in a plurality of pre-defined entity names (S440). Then, if it is identified that the subword sequence is not included in a plurality of pre-defined entity names (S440-N), the entity name correction module (262) can correct the entity name of the subword sequence based on a phoneme sequence, thereby obtaining text information (S450). Conversely, if the subword sequence is identified as being included in a plurality of predefined entity names (S440-Y), the entity name correction module (262) may not correct the entity name of the subword sequence, and in this case, the text information acquisition module (260) may acquire text information corresponding to the subword sequence.
[0086] In particular, the spelling correction process by the spelling correction module (261) and the entity name correction process by the entity name correction module (262) can be performed using an edit distance. Here, the edit distance refers to the minimum number of character removals, insertions, and substitutions required to convert a subword sequence into text information with the correct spelling or entity name. For example, the spelling correction module (261) can obtain text information with the correct spelling by calculating the edit distance between a phoneme sequence and candidate text sequences with the correct spelling.
[0087] In the description of Figure 2, an example of obtaining text information called [Bhukansan] by correcting the subword sequence called [Book on sun] using the phoneme sequence [ / bh / / u / / k / / a / / n / / s / / a / / n / ] is described, but various other examples can be given.
[0088] For example, the text information acquisition module (260) can acquire text information called [Gyeongbokgung] by correcting the subword sequence called [Jumbo gum] or [Jungle gym] using the phoneme sequence called [ / g / / y / / eo / / ng / / b / / o / / k / / g / / u / / ng / ], and can also acquire text information called [Gyeonggi-do] by correcting the subword sequence called [Young dido] using the phoneme sequence called [ / g / / y / / eo / / n / / g / / i / / d / / o / ].
[0089] FIG. 5 is a diagram for explaining a plurality of sections of a voice signal according to one embodiment of the present disclosure.
[0090] The vectors representing the characteristics of the voice signal or the voice signal according to the present disclosure can be expressed as a continuous sequence as shown in Fig. 5. In Fig. 5, x trefers to a specific point in the sequence, y t is a specific point in time x t Indicates the unit of output by the encoder or decoder.
[0091] As described above, the common encoder (210) according to the present disclosure can output a first vector corresponding to each of a plurality of sections of a speech signal. In addition, the first individual encoder (220) can output a second vector corresponding to each of a plurality of sections and independent of the context of the speech signal based on the first vector. Specifically, as illustrated in the first image (410) of FIG. 5, the common encoder (210) and the first individual encoder (220) each output a first vector corresponding to each of a plurality of sections at a specific time point x t At a specific point in time x t From a specific point in time x to a previous point in time corresponding to one section based on t A limited section of a voice signal can be processed to output a first vector and a second vector. Here, it is of course possible to determine the length of one section differently depending on the developer or user's settings.
[0092] Meanwhile, the second individual encoder (230) according to the present disclosure may output a third vector corresponding to at least two sections among a plurality of sections of a voice signal based on the first vector. Here, the at least two sections among the plurality of sections may be all sections received before a specific point in time among the plurality of sections or all sections received before and after a specific point in time among the plurality of sections. For convenience of explanation, the expression "two or more sections" was used, but depending on the embodiment, any section that includes more time than one section may correspond to a section of the third vector according to the present disclosure.
[0093] Specifically, as shown in the second image (420) of FIG. 5, the second individual encoder (230) at a specific point in time x tFrom a specific point in time (x) to a specific point in time corresponding to two intervals based on t ) can process the first vector of a limited section to output the third vector. In other words, the second individual encoder (230) can output the third vector at a specific point in time x t The first vector corresponding to the entire previously received voice signal can be converted into a third vector and output. That is, when the second individual encoder (230) is input with the first vector corresponding to each of the limited sections of the voice signal, the second individual encoder (230) can output a third vector including a left-only context commonly used in a voice recognition model (e.g., automatic speech recognition, ASR).
[0094] Meanwhile, as shown in the third image (430) of FIG. 5, the second individual encoder (230) at a specific point in time x t Not only before but also at a specific point in time x t The first vector corresponding to the entire voice signal received thereafter may be converted into a third vector and output. That is, when the second individual encoder (230) receives the first vector corresponding to each of the limited sections of the voice signal, the second individual encoder (230) can obtain a third vector including the full context commonly used in the full context ASR model (or causal ASR). In other words, the second individual encoder (230) can obtain a third vector including the full context at a specific point in time x t In outputting the third vector corresponding to a specific point in time x t The third vector can be obtained using the first vector corresponding to the voice signal received thereafter.
[0095] According to the prior art, as illustrated in the second image (420) or the third image (430) of FIG. 5, a context-dependent encoder and decoder are included that utilize speech signals of the entire section received before a specific point in time or of the entire section received before and after a specific point in time among multiple sections. Therefore, it can be strongly biased by previous words.
[0096] On the other hand, the electronic device (100) according to the present disclosure may include a processing process of a context-independent path (e.g., a path according to the second individual encoder (230) and the second decoder (250)) that uses only a limited section of speech signal, as illustrated in the first image (410) of FIG. 5, as well as a context-dependent path like the prior art. Accordingly, the electronic device (100) can obtain a phoneme sequence representing an actual pronunciation without bias of language information, and correct a subword sequence using the phoneme sequence.
[0097] FIG. 6 is a block diagram showing in detail the configuration of an electronic device (100) according to an embodiment of the present disclosure.
[0098] As illustrated in FIG. 6, an electronic device (100) according to an embodiment of the present disclosure may further include a communication unit (130, e.g., including a communication circuit), an input unit (140, e.g., including an input circuit), and an output unit (150, e.g., including an output circuit). However, the configurations illustrated in FIGS. 1 and 6 are merely exemplary, and it is to be understood that new configurations may be added or some configurations may be omitted in addition to the configurations illustrated in FIGS. 1 and 6 when implementing the present disclosure.
[0099] The communication unit (130) includes a circuit and can perform communication with an external device. Specifically, the processor (120) can receive various data or information from an external device connected via the communication unit (130) and can also transmit various data or information to the external device.
[0100] The communication unit (130) may include at least one of a Wi-Fi module, a Bluetooth module, a wireless communication module, an NFC module, and a UWB (Ultra Wide Band) module. Specifically, the Wi-Fi module and the Bluetooth module may each perform communication in the Wi-Fi or Bluetooth manner. When using a Wi-Fi module or a Bluetooth module, various connection information, such as an SSID, may be first transmitted and received, and then communication may be established using this, after which various pieces of information may be transmitted and received.
[0101] In addition, the wireless communication module can perform communication according to various communication standards such as IEEE, Zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), 5G (5th Generation), etc. And, the NFC module can perform communication in the NFC (Near Field Communication) method using the 13.56MHz band among various RF-ID frequency bands such as 135kHz, 13.56MHz, 433MHz, 860~960MHz, 2.45GHz, etc. In addition, the UWB module can accurately measure the ToA (Time of Arrival), which is the time it takes for a pulse to reach a target, and the AoA (Angle of Arrival), which is the pulse arrival angle at the transmitting device, through communication between UWB antennas, and accordingly, precise distance and location recognition is possible within an error range of several tens of centimeters indoors.
[0102] In particular, in various embodiments according to the present disclosure, the processor (120) may receive a voice signal from an external device through the communication unit (130), and may also receive a user command to perform voice recognition on the voice signal. In addition, the processor (120) may control the communication unit (130) to transmit at least one of the first vector, the second vector, the third vector, the phoneme sequence, the subword sequence, and the text information according to the present disclosure to the external device.
[0103] The input unit (140) includes a circuit, and the processor (120) can receive a user command to control the operation of the electronic device (100) through the input unit (140). Specifically, the input unit (140) can be configured with components such as a microphone, a camera (not shown), and a remote control signal receiving unit (not shown). In addition, the input unit (140) can also be implemented in a form included in a display as a touch screen. In particular, the microphone can receive a voice signal and convert the received voice signal into an electrical signal.
[0104] In particular, in various embodiments according to the present disclosure, the processor (120) may receive a user command for performing voice recognition on a voice signal through the input unit (140). In addition, the processor (120) may also receive a user command for setting the sizes of a plurality of sections according to the present disclosure through the input unit (140).
[0105] The output unit (150) includes a circuit, and the processor (120) can output various functions that the electronic device (100) can perform through the output unit (150). In addition, the output unit (150) can include at least one of a display, a speaker, and an indicator.
[0106] The display can output image data under the control of the processor (120). Specifically, the display can output an image previously stored in the memory (110) under the control of the processor (120).
[0107] In particular, a display according to an embodiment of the present disclosure may display a user interface stored in a memory (110). The display may be implemented as a liquid crystal display panel (LCD), organic light emitting diodes (OLED), etc., and may also be implemented as a flexible display, a transparent display, etc., depending on the case. However, the display according to the present disclosure is not limited to a specific type. The speaker may output audio data under the control of the processor (120), and the indicator may be turned on under the control of the processor (120).
[0108] In particular, in various embodiments according to the present disclosure, the processor (120) may control the output unit (150) to output text information according to the present disclosure. In addition, the processor (120) may control the output unit (150) to provide a subword sequence before spelling correction or entity name correction according to the present disclosure is performed, together with text information on which spelling correction or entity name correction has been performed.
[0109] FIG. 7 is a flowchart illustrating a control method of an electronic device (100) according to one embodiment of the present disclosure.
[0110] As illustrated in FIG. 7, the electronic device (100) can acquire a voice signal (S710). Specifically, the electronic device (100) may acquire a voice signal by receiving a voice signal through a microphone included in the electronic device (100), or may acquire a voice signal by receiving a voice signal from an external device through the communication unit (130).
[0111] Once a voice signal is acquired, the electronic device (100) can input the voice signal into a common encoder (210) to acquire a first vector corresponding to each of a plurality of sections of the voice signal (S720). The common encoder (210) can be trained to acquire a first vector suitable for both the first individual encoder (220) and the second individual encoder (230) described below without any preset constraints.
[0112] Once the first vector is acquired, the electronic device (100) inputs the first vector into the first individual encoder (220) to acquire a second vector corresponding to each of the plurality of sections and independent of the context of the speech signal (S730). Specifically, once the first vector corresponding to each of the plurality of sections is acquired, the electronic device (100) can acquire a second vector independent of the context based on the speech signal limited to each of the plurality of sections.
[0113] When the second vector is acquired, the electronic device (100) can input the second vector into the first decoder (240) to acquire a phoneme sequence corresponding to the second vector (S740). Specifically, when the second vector is acquired, the electronic device (100) can acquire a phoneme sequence corresponding to the second vector.
[0114] Meanwhile, when the first vector is acquired, the electronic device (100) can input the first vector into the second individual encoder (230) to acquire a third vector that corresponds to at least two sections among the plurality of sections and is dependent on the context of the speech signal (S750). Specifically, when the first vector is acquired, the electronic device (100) can acquire a third vector that is dependent on the context based on the unrestricted speech signal of at least two sections among the plurality of sections.
[0115] When the third vector is acquired, the electronic device (100) can input the third vector into the second decoder (250) to acquire a subword sequence corresponding to the third vector (S760). Specifically, when the third vector is acquired, the electronic device (100) can acquire a subword sequence corresponding to the third vector.
[0116] Once the phoneme sequence and subword sequence are acquired, the electronic device (100) can acquire text information corresponding to multiple sections by correcting the subword sequence based on the phoneme sequence through the text information acquisition module (260) (S770). Specifically, the electronic device (100) can identify whether the subword sequence needs to be corrected, and if it is identified that the subword sequence needs to be corrected, correct the subword sequence using the phoneme sequence, thereby acquiring accurate text information corresponding to the user's voice.
[0117] Meanwhile, the control method of the electronic device (100) according to the above-described embodiment may be implemented as a program and provided to the electronic device (100). In particular, the program including the control method of the electronic device (100) may be stored and provided in a non-transitory computer readable medium.
[0118] Specifically, in a non-transitory computer-readable recording medium including a program for executing a control method of an electronic device, the control method of the electronic device includes the steps of: when a voice signal is acquired, inputting the voice signal into a common encoder to acquire a first vector corresponding to each of a plurality of sections of the voice signal; inputting the first vector into a first individual encoder to acquire a second vector corresponding to each of the plurality of sections and independent of the context of the voice signal; inputting the second vector into a first decoder to acquire a phoneme sequence corresponding to the second vector; inputting the first vector into a second individual encoder to acquire a third vector corresponding to at least two sections among the plurality of sections and dependent on the context of the voice signal; inputting the third vector into the second decoder to acquire a sub-word sequence corresponding to the third vector; and inputting the phoneme sequence and the sub-word sequence into a text information acquisition module to acquire text information corresponding to the plurality of sections.
[0119] In the above, a method for controlling an electronic device (100) and a computer-readable recording medium including a program for executing the method for controlling an electronic device (100) have been briefly described, but this is only to omit redundant descriptions, and it goes without saying that various embodiments of the electronic device (100) can also be applied to a method for controlling an electronic device (100) and a computer-readable recording medium including a program for executing the method for controlling an electronic device (100).
[0120] According to various embodiments of the present disclosure as described above, the electronic device (100) can train the context-independent encoder to acquire a phoneme sequence representing actual pronunciation by focusing on individual phonemes without biasing linguistic information using a speech signal of a limited section by dividing the encoder into a context-independent encoder and a context-dependent encoder. Accordingly, the electronic device (100) can perform spelling correction or named entity correction, etc. based on the phoneme sequence, thereby performing accurate speech recognition results for the speech signal.
[0121] The artificial intelligence-related function according to the present disclosure is operated through the processor (120) and memory (110) of the electronic device (100).
[0122] The processor (120) may be composed of one or more processors (120). At this time, the one or more processors (120) may include at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an NPU (Neural Processing Unit), but is not limited to the examples of the processors (120) described above. As described above, the processor (120) may include various processing circuits, and as used herein, including in the claims, the term "processor" may include various processing circuits including at least one processor, wherein at least one processor of the at least one processor may be configured to perform various functions according to the present disclosure.
[0123] The CPU is a general-purpose processor (120) capable of performing not only general calculations but also artificial intelligence calculations. Its multi-layer cache structure allows for the efficient execution of complex programs. The CPU is advantageous in a serial processing method, enabling organic linking of previous and subsequent calculation results through sequential calculations. The general-purpose processor (120) is not limited to the examples described above, except in cases where it is specifically referred to as a CPU.
[0124] A GPU is a processor (120) for large-scale operations such as floating point operations used in graphic processing, and can perform large-scale operations in parallel by integrating a large number of cores. In particular, a GPU may be advantageous compared to a CPU in parallel processing methods such as convolution operations. In addition, a GPU may be used as a co-processor (120) to supplement the functions of a CPU. The processor (120) for large-scale operations is not limited to the examples described above, except in cases where it is specified as a GPU as described above.
[0125] An NPU is a processor (120) specialized in artificial intelligence operations using an artificial neural network, and each layer constituting the artificial neural network can be implemented with hardware (e.g., silicon). At this time, since the NPU is designed specifically according to the required specifications of the company, it has a lower degree of freedom compared to a CPU or GPU, but it can efficiently process the artificial intelligence operations requested by the company. Meanwhile, as a processor (120) specialized in artificial intelligence operations, the NPU can be implemented in various forms such as a TPU (Tensor Processing Unit), an IPU (Intelligence Processing Unit), a VPU (Vision processing unit), etc. The artificial intelligence processor (120) is not limited to the above-described examples, except in cases where it is specified as the above-described NPU.
[0126] Additionally, one or more processors (120) may be implemented as a SoC (System on Chip). In this case, the SoC may further include, in addition to one or more processors (120), a memory (110), and a network interface such as a bus for data communication between the processor (120) and the memory (110).
[0127] When a plurality of processors (120) are included in a SoC (System on Chip) included in an electronic device (100), the electronic device (100) may perform operations related to artificial intelligence (e.g., operations related to learning or inference of an artificial intelligence model) by using some of the processors (120) among the plurality of processors (120). For example, the electronic device (100) may perform operations related to artificial intelligence by using at least one of a GPU, an NPU, a VPU, a TPU, and a hardware accelerator specialized in artificial intelligence operations such as convolution operations and matrix multiplication operations among the plurality of processors (120). However, this is merely an example, and it is of course possible to process operations related to artificial intelligence by using a CPU or a general-purpose processor (120).
[0128] In addition, the electronic device (100) can perform operations related to functions related to artificial intelligence by utilizing multiple cores (e.g., dual cores, quad cores, etc.) included in one processor (120). In particular, the electronic device (100) can perform artificial intelligence operations such as convolution operations, matrix multiplication operations, etc. in parallel by utilizing multiple cores included in the processor (120).
[0129] One or more processors (120) are controlled to process input data according to predefined operation rules or artificial intelligence models stored in the memory (110). The predefined operation rules or artificial intelligence models are characterized by being created through learning.
[0130] Here, "created through learning" means that a predefined set of behavioral rules or an AI model with desired characteristics is created by applying a learning algorithm to a large number of learning data. This learning may be performed on the device itself, where the AI according to the present disclosure is implemented, or through a separate server / system.
[0131] An artificial intelligence model may be composed of multiple neural network layers. At least one layer has at least one weight value and performs its operation through the operation result of the previous layer and at least one defined operation. Examples of neural networks include a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), deep Q-networks, and a transformer. The neural networks in the present disclosure are not limited to the above-described examples unless otherwise specified.
[0132] A learning algorithm is a method for training a target device (e.g., a robot) using a large amount of learning data, enabling the target device to make decisions or predictions on its own. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Unless otherwise specified, the learning algorithms in this disclosure are not limited to the aforementioned examples.
[0133] A device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, a "non-transitory storage medium" means a tangible device that does not contain signals (e.g., electromagnetic waves). For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.
[0134] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smartphones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product (e.g., downloadable app) may be at least temporarily stored or temporarily created in a device-readable storage medium, such as a memory (110) of a manufacturer's server, an application store's server, or an intermediary server.
[0135] Each of the components (e.g., modules or programs) according to the various embodiments of the present disclosure as described above may be composed of a single or multiple entities, and some of the sub-components described above may be omitted, or other sub-components may be further included in the various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into a single entity, which may perform the same or similar functions as those performed by each of the respective components prior to integration.
[0136] According to various embodiments, operations performed by a module, program or other component may be executed sequentially, in parallel, iteratively or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.
[0137] Meanwhile, the terms "part" or "module" used in the present disclosure include a unit composed of hardware, software, or firmware, or a combination thereof, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A "part" or "module" may be an integrally formed component, a minimum unit performing one or more functions, or a portion thereof. For example, a module may be composed of an application-specific integrated circuit (ASIC).
[0138] Various embodiments of the present disclosure may be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device is a device that can call instructions stored in the storage medium and operate according to the called instructions, and may include an electronic device (100) (e.g., electronic device (100)) according to the disclosed embodiments.
[0139] When the above command is executed by the processor (120), the processor (120) may perform a function corresponding to the command directly or by using other components under the control of the processor (120). The command may include code generated or executed by a compiler or interpreter.
[0140] While various embodiments according to the present disclosure have been illustrated and described above, the various embodiments are illustrative and not limiting. For example, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by those skilled in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims. Furthermore, such modifications should not be understood individually from the technical idea or prospect of the present disclosure. Furthermore, any embodiment described herein may be used in conjunction with any other embodiment described herein.
Claims
1. In electronic devices, a memory that stores at least one instruction; and At least one processor executing at least one instruction; At least one processor, When a voice signal is acquired, the voice signal is input to a common encoder to acquire a first vector corresponding to each of a plurality of sections of the voice signal, Inputting the first vector into a first individual encoder to obtain a second vector corresponding to each of the plurality of sections and independent of the context of the speech signal, By inputting the second vector into the first decoder, a phoneme sequence corresponding to the second vector is obtained, By inputting the first vector to a second individual encoder, a third vector corresponding to at least two sections among the plurality of sections and dependent on the context of the speech signal is obtained, By inputting the third vector into the second decoder, a sub-word sequence corresponding to the third vector is obtained, An electronic device that obtains text information corresponding to the plurality of sections by correcting the subword sequence based on the phoneme sequence through a text information acquisition module.
2. In paragraph 1, The above text information acquisition module includes a spell correction module, The above spelling correction module, An electronic device that obtains the text information by correcting the spelling of the subword sequence based on the phoneme sequence when the subword sequence is identified as being in violation of a specific spelling.
3. In paragraph 1, The above text information acquisition module includes a named entity correction module, The above entity name correction module is, An electronic device that obtains the text information by correcting the entity name of the subword sequence based on the phoneme sequence when the subword sequence is identified as not being included in a specified plurality of entity names.
4. In paragraph 1, The above common encoder is, An electronic device trained to obtain the first vector suitable for both the first individual encoder and the second individual encoder without any specific constraints.
5. In paragraph 1, The first individual encoder is trained to obtain the second vector representing the features of the phoneme sequence based on unlabeled training data, An electronic device wherein the first decoder is trained to obtain the phoneme sequence based on labeled training data.
6. In paragraph 1, The second individual encoder is trained to obtain the third vector representing the features of the subword sequence based on unlabeled training data, An electronic device wherein the second decoder is trained to obtain the subword sequence based on labeled training data.
7. In paragraph 1, An electronic device wherein at least two of the plurality of sections include the entire section received before a specific point in time among the plurality of sections or the entire section received before and after the specific point in time among the plurality of sections.
8. In a method for controlling an electronic device, When a voice signal is obtained, a step of inputting the voice signal into a common encoder to obtain a first vector corresponding to each of a plurality of sections of the voice signal; A step of inputting the first vector into a first individual encoder to obtain a second vector corresponding to each of the plurality of sections and independent of the context of the speech signal; A step of inputting the second vector into a first decoder to obtain a phoneme sequence corresponding to the second vector; A step of inputting the first vector into a second individual encoder to obtain a third vector corresponding to at least two sections among the plurality of sections and dependent on the context of the speech signal; A step of inputting the third vector into a second decoder to obtain a sub-word sequence corresponding to the third vector; and A control method of an electronic device, comprising: a step of obtaining text information corresponding to the plurality of sections by correcting the subword sequence based on the phoneme sequence through a text information obtaining module; 9. In paragraph 8, A control method of an electronic device further comprising: a step of obtaining the text information by correcting the spelling of the subword sequence based on the phoneme sequence, if the subword sequence is identified as being in violation of a specific spelling; 10. In paragraph 8, A control method of an electronic device further comprising: a step of obtaining the text information by correcting the entity name of the subword sequence based on the phoneme sequence, if the subword sequence is identified as not being included in a specified plurality of entity names; 11. In paragraph 8, A control method of an electronic device further comprising: a step of learning to obtain the first vector suitable for both the first individual encoder and the second individual encoder without any specific constraints; 12. In paragraph 8, A step of learning to obtain the second vector representing the characteristics of the phoneme sequence based on unlabeled training data; and A method for controlling an electronic device, further comprising: a step of learning to obtain the phoneme sequence based on labeled learning data; 13. In paragraph 8, A step of learning to obtain the third vector representing the characteristics of the subword sequence based on unlabeled learning data; and A control method of an electronic device further comprising: a step of learning to obtain the subword sequence based on labeled learning data; 14. In paragraph 8, A control method of an electronic device, wherein at least two of the plurality of sections include the entire section received before a specific point in time among the plurality of sections or the entire section received before and after the specific point in time among the plurality of sections.
15. In a non-transitory computer-readable recording medium including a program for executing a method of controlling an electronic device, The method of controlling the above electronic device is as follows: When a voice signal is obtained, a step of inputting the voice signal into a common encoder to obtain a first vector corresponding to each of a plurality of sections of the voice signal; A step of inputting the first vector into a first individual encoder to obtain a second vector corresponding to each of the plurality of sections and independent of the context of the speech signal; A step of inputting the second vector into a first decoder to obtain a phoneme sequence corresponding to the second vector; A step of inputting the first vector into a second individual encoder to obtain a third vector corresponding to at least two sections among the plurality of sections and dependent on the context of the speech signal; A step of inputting the third vector into a second decoder to obtain a sub-word sequence corresponding to the third vector; and A computer-readable recording medium comprising: a step of inputting the phoneme sequence and the subword sequence into a text information acquisition module to acquire text information corresponding to the plurality of sections;
Citation Information
Patent Citations
Voice recognizing device
JP1994175678A
Apparatus and method using two phase utterance verification architecture for computation speed improvement of n-best recognition word
KR1020110070688A
Utterance verification method for deep neural network speech recognition system
KR1020180117942A
Voice signal perturbation for speech recognition
WO2007117814A2
KR20220068679A