Electronic device and control method thereof

By performing acoustic enhancement on the original audio data, the problem of degradation of recognition performance when processing different users' speeches is solved, and higher speech recognition accuracy is achieved.

CN120112992APending Publication Date: 2025-06-06SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380078127.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-11
Filing Date
2023-10-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When existing speech recognition models process speeches of different users, it is difficult to accurately recognize speech, especially in the presence of various speech characteristics changes, resulting in a degradation of recognition performance.

Method used

By performing acoustic enhancement on the original audio data, enhancement audio data is obtained and both the original audio data and enhanced audio data are used in the decoder to improve the accuracy of speech recognition.

Benefits of technology

Through acoustic enhancement technology, the recognition performance of the speech recognition model can be improved and the accuracy of user speech recognition can be enhanced, especially in the case of changes in speech characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120112992A_ABST
    Figure CN120112992A_ABST
Patent Text Reader

Abstract

An electronic device is disclosed. The electronic device according to one embodiment of the present disclosure includes a memory storing instructions, and at least one processor executing the instructions, obtains first audio data including user voice, obtains second audio data by acoustically enhancing the first audio data, and transmits the second audio data to the memory. Obtaining a first interval corresponding to audio data between points including a user utterance among the first audio data and a second interval corresponding to the first interval of the second audio data, calculating a plurality of first scores corresponding to each of a plurality of estimation candidates based on the first audio data corresponding to the first interval, a plurality of second scores corresponding to each of the plurality of estimation candidates are calculated based on second audio data corresponding to the second interval, and one of the plurality of estimation candidates is determined as character data based on the plurality of first scores and the plurality of second scores.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an electronic device related to an artificial intelligence (AI) learning algorithm and a control method thereof, and more particularly, to an electronic device including a decoding process for enhancing the performance of a trained AI model and a control method thereof. Background Art

[0002] AI systems are computer systems that achieve human-level intelligence and are systems in which machines can learn, determine, and develop on their own, unlike existing rule-based intelligent systems.

[0003] The speech recognition technology applied by the AI ​​technology of the above-mentioned AI system is a technology for recognizing and applying / processing human language / characters. Here, the speech recognition model used for speech recognition technology can perform natural language processing, machine translation, conversation system, query response, speech recognition / synthesis, etc.

[0004] Specifically, for audio data different from the learning data, the trained speech recognition model may show incorrect speech recognition results. Therefore, acoustic enhancement can be used for more accurate speech recognition of a speech recognition model trained with limited audio data. Summary of the invention

[0005] Technical Solution

[0006] According to an embodiment of the present disclosure, an electronic device includes a memory storing instructions and at least one processor executing the instructions. The at least one processor obtains first audio data including user speech, performs acoustic enhancement on the first audio data to obtain second audio data, obtains a first interval corresponding to audio data between time points including user speech in the first audio data and a second interval corresponding to the first interval of the second audio data, calculates multiple first scores corresponding to each of multiple estimation candidates based on the first audio data corresponding to the first interval, calculates multiple second scores corresponding to each of the multiple estimation candidates based on the second audio data corresponding to the second interval, and determines one of the multiple estimation candidates as character data based on the multiple first scores and the multiple second scores.

[0007] According to an embodiment of the present disclosure, a method for controlling an electronic device includes: obtaining first audio data including user voice, performing acoustic enhancement on the first audio data to obtain second audio data, obtaining a first interval corresponding to audio data between time points including user speech in the first audio data and a second interval of the second audio data corresponding to the first interval, calculating multiple first scores corresponding to each of a plurality of estimation candidates based on the first audio data corresponding to the first interval, calculating multiple second scores corresponding to each of a plurality of estimation candidates based on the second audio data corresponding to the second interval, and determining one of the plurality of estimation candidates as character data based on the multiple first scores and the multiple second scores.

[0008] According to an embodiment for implementing one aspect of the present disclosure, a computer-readable recording medium including a program for executing a method for controlling an electronic device enables the electronic device to: obtain first audio data including user voice, perform acoustic enhancement on the first audio data to obtain second audio data, obtain a first interval corresponding to audio data between time points including user speech in the first audio data and a second interval of the second audio data corresponding to the first interval, calculate multiple first scores corresponding to each of a plurality of estimation candidates based on the first audio data corresponding to the first interval, calculate multiple second scores corresponding to each of a plurality of estimation candidates based on the second audio data corresponding to the second interval, and determine one of the plurality of estimation candidates as character data based on the multiple first scores and the multiple second scores. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 is a view showing an electronic device according to an embodiment;

[0010] Figure 2 is a block diagram showing a configuration of an electronic device according to an embodiment;

[0011] Figure 3 is a block diagram showing a detailed configuration of an electronic device according to an embodiment;

[0012] Figure 4 is a view showing a voice recognition process of an electronic device according to an embodiment;

[0013] Figure 5 is a view showing a process of determining character data of an electronic device according to an embodiment;

[0014] Figure 6 is a view showing a process of obtaining a plurality of first intervals and a plurality of second intervals according to an embodiment; and

[0015] Figure 7is a flowchart illustrating a method of controlling an electronic device according to an embodiment. DETAILED DESCRIPTION

[0016] The embodiments of the present disclosure may be modified in various forms and may have various embodiments, wherein specific embodiments will be illustrated in the drawings and specifically explained in the detailed description. However, it should be noted that the various embodiments are not intended to limit the scope of the present disclosure to specific embodiments, but they should be interpreted as including all modifications, equivalents and / or alternatives of the embodiments of the present disclosure. With respect to the description of the drawings, similar components may be represented by similar reference numerals.

[0017] In the case where it is determined that a detailed description of a related known function or configuration may unnecessarily obscure the gist of the present disclosure in describing the present disclosure, the detailed description thereof is omitted.

[0018] In addition, the following embodiments can be modified in various forms, and the scope of the technical ideas of the present disclosure is not limited to the following embodiments. On the contrary, these embodiments are provided to make the present disclosure more complete and complete, and to fully convey the technical ideas of the present disclosure to those skilled in the art.

[0019] The terms used in the present disclosure are only used to explain specific embodiments and are not intended to limit the scope of the present disclosure. Unless clearly defined differently in the context, a singular expression includes a plural expression.

[0020] In the present disclosure, expressions such as “have”, “may have”, “include” or “may include” indicate the presence of such characteristics (for example, numerical values, functions, operations, or components such as parts), and the expression does not exclude the presence of additional characteristics.

[0021] In the present disclosure, the expressions "A or B", "at least one of A and / or B", "one or more of A and / or B", etc. may include all possible combinations of the listed items. For example, "A or B", "at least one of A and B", or "at least one of A or B" may refer to all of the following: (1) including at least one A, (2) including at least one B, or (3) including all of at least one A and at least one B.

[0022] The expressions “first”, “second”, “first”, “second”, etc. used in the present disclosure may be used to describe various elements regardless of any order and / or importance, wherein the expressions are merely used to distinguish one element from another and are not intended to limit the elements.

[0023] Meanwhile, description that one element (e.g., a first element) is “(operably or communicatively) coupled” or “connected to” another element (e.g., a second element) should be interpreted such that one element is directly coupled to another element or one element is coupled to another element through another element (e.g., a third element).

[0024] In contrast, description that one element (eg, a first element) is “directly coupled” or “directly connected” to another element (eg, a second element) may be interpreted as indicating that another element (eg, a third element) does not exist between the one element and the other element.

[0025] The expression "configured to..." used in the present disclosure may be used interchangeably with other expressions, such as "suitable for...", "capable of...", "designed to...", "suitable for...", "made to..." or "capable of...", depending on the situation. The term "configured to..." may not necessarily mean that a device is "specifically designed to" do so in terms of hardware.

[0026] On the contrary, in some cases, the expression "a device configured to..." may mean that the device is "capable" of performing an operation together with another device or component. For example, the phrase "a processor configured to perform A, B, and C" may refer to a dedicated processor (e.g., an embedded processor) for performing the corresponding operations, or a general-purpose processor (e.g., a CPU or an application processor) that can perform the corresponding operations by executing one or more software programs stored in a memory device.

[0027] In the embodiments of the present disclosure, a "module" or "component" may perform at least one function or operation, and may be implemented as hardware or software, or as a combination of hardware and software. In addition, multiple "modules" or "components" may be integrated into at least one module and implemented as at least one processor, excluding "modules" or "components" that need to be implemented as specific hardware.

[0028] Meanwhile, various elements and regions in the drawings are schematically shown. Therefore, the technical concept of the present disclosure is not limited by the relative sizes or intervals shown in the drawings.

[0029] Hereinafter, with reference to the accompanying drawings, embodiments according to the present disclosure are described in detail to be easily implemented by those skilled in the art.

[0030] Figure 1 is a view showing an electronic device according to an embodiment.

[0031] refer to Figure 1, although the sentences are the same, user utterances may include different voice features depending on the user. Specifically, user utterances may include various voice features, such as speed, pitch, amplitude, etc., and various users may speak at various speeds, pitches, and amplitudes with respect to the same sentence. For example, user 1 may have a faster speech speed, a higher speech pitch, and a larger speech amplitude. In contrast, user 2 may have a slower speech speed, a lower speech pitch, and a smaller speech amplitude. The above-mentioned various voice features may differ according to the user's age, gender, health status, etc.

[0032] In addition, even in the case of the speech of the same user, although it is the same user, the speech may be speech including different speech features depending on where the user is or the distance from the user to the speech recognition model. For example, compared with the case where the user speaks in the living room or the room at home, the reverberation may be greater in the case where the user speaks in the bathroom. In addition, in the case of being outside compared to inside, noise caused from the surrounding environment may be included in the user speech, or in the case where the distance between the user and the speech recognition model is far, this situation may affect the speech features.

[0033] Therefore, regarding the same sentence, the speech recognition model is required to recognize the user utterance as the same character data while taking into account various speech features regarding the user utterance.

[0034] As described above, in the case of a speech recognition model for recognizing user utterances including various speech features, training it based on data including various speech features can be a factor that preferentially determines speech recognition performance. That is, since it learns the sound, voice, and language changes required for speech recognition based on the transcription data of the voice-character pair, a large amount of transcription data including various speech features is required for robust modeling. However, it may be difficult to collect transcription data including all speech features for training the speech recognition model, because collecting a large amount of data consumes a lot of cost and time. In addition, although the speech recognition model is trained by using limited transcription data and then collecting new transcription data, retraining the speech recognition model may consume a lot of cost and time.

[0035] Therefore, in order to enhance the recognition performance of the trained speech recognition model, enhanced audio data can be obtained by performing acoustic enhancement on the original audio data, and the original audio data and the enhanced audio data can be used simultaneously in the decoder.

[0036] Figure 2 is a block diagram showing a configuration of an electronic device according to an embodiment.

[0037] refer to Figure 2, the electronic device 100 may include a memory 110 and at least one processor 120. However, Figure 2 The illustrated configuration of the electronic device 100 is merely an example, wherein it is apparent that another configuration may be added or part of the configuration may be omitted.

[0038] The electronic device 100 may include a computer or a user terminal device such as, for example, a smart TV, a tablet PC, a monitor, a smart phone, a desktop computer, a laptop computer, a mobile device, or a wearable device.

[0039] The electronic device 100 may include a home appliance such as an air conditioner, a washing machine, a refrigerator, a speaker, an iron, a coffee maker, a vacuum cleaner, a dishwasher, an electric range, a gas range, an induction range, a fan, a cleaning robot, a service robot, or a medical robot.

[0040] Specifically, the electronic device 100 can train a speech recognition model through interaction between the memory 110 and the processor 120, and perform a speech recognition function through the speech recognition model.

[0041] According to an embodiment of the present disclosure, the memory 110 may store an operating system (OS) for controlling the overall operation of the components of the electronic device 100 and instructions or data related to the components of the electronic device 100. Specifically, the memory 110 may store an image obtained by photographing or capturing a display image by a camera. In addition, it may store an image obtained through a communication interface. In addition, in order to display a text image included in the obtained image as an alternative text image, the memory 110 may store instructions or data for generating an alternative text image. As described above, the memory 110 may include, for example, at least one of a main storage and an auxiliary storage. The main memory may be implemented by using a semiconductor storage medium such as a ROM and / or a RAM. The ROM may include, for example, a general-purpose ROM, an EPROM, an EEPROM, and / or a MASK-ROM. The RAM may include, for example, a DRAM and / or an SRAM. The auxiliary storage may be implemented by using at least one storage medium that can store data permanently or semi-permanently, such as a flash memory device, a secure digital (SD) card, a solid-state drive (SSD), a hard disk drive (HDD), a magnetic drum, an optical medium (such as a compact disk (CD), a DVD, or a laser disk), a magnetic tape, a magneto-optical disk, and / or a floppy disk.

[0042] Specifically, the memory 110 may store a speech recognition model trained by limited transcription data. That is, the electronic device 100 may determine character data corresponding to the audio data through the speech recognition model stored in the memory 110.

[0043] In addition, the memory 110 may store a decoder. Here, the decoder may output the probability of character data recognized from the initial time point to the time point before the recognition time point and the probability of character data of the estimated candidate characters including the recognition time point based on the probability value of the audio data about the recognition time point output from the speech recognition model. Here, the probability about the character data may be expressed as a score.

[0044] According to an embodiment, at least one processor 130 controls operations of the electronic device 100 as a whole.

[0045] According to an example of the present disclosure, at least one processor 130 may be implemented as a digital signal processor (DSP), a microprocessor, or a time controller (TCON) for processing digital signals. At the same time, the present disclosure is not limited thereto, and it may include one or more of a central processing unit (CPU), a microcontroller unit (MCU), a microprocessing unit (MPU), a controller, an application processor (AP), a communication processor (CP), an ARM processor, or an AI processor, or may be defined by related terms. In addition, at least one processor 130 may be implemented as a system on chip (SoC) or a large-scale integration (LSI) having a processing algorithm embedded thereon, and may be implemented as a field programmable gate array (FPGA). At least one processor 130 may perform various functions by executing computer executable instructions stored in a memory.

[0046] At least one processor 130 may include one or more of a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), an integrated many-core (MIC), a digital signal processor (DSP), a neural processing unit (NPU), a hardware accelerator, or a machine learning accelerator. At least one processor 130 may control one or any combination of other components of the electronic device and perform operations related to communication or data processing. At least one processor 130 may execute at least one program or instruction stored in a memory. For example, at least one processor 130 may execute a method according to an embodiment of the present disclosure by executing at least one instruction stored in a memory.

[0047] If the method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by one processor and may be performed by multiple processors. For example, when the first operation, the second operation, and the third operation are performed by the method according to the embodiment, the first operation, the second operation, and the third operation may all be performed by the first processor, and the first operation and the second operation are performed by the first processor (e.g., a general-purpose processor), and the third operation may be performed by the second processor (e.g., an artificial intelligence (AI) dedicated processor).

[0048] At least one processor 130 may be implemented as a single-core processor including one core, and may be implemented as at least one multi-core processor including multiple cores (e.g., homogeneous multi-cores or heterogeneous multi-cores). If at least one processor 130 is implemented as a multi-core processor, each of the multiple cores included in the multi-core processor may include a processor internal memory, such as a cache memory and an on-chip memory, wherein a common cache shared by the multiple cores may be included in the multi-core processor. In addition, each of the multiple cores included in the multi-core processor (or a portion of the multiple cores) may read and execute program instructions for independently implementing the method according to an embodiment of the present disclosure, and may also read and execute program instructions for implementing the method according to an embodiment of the present disclosure in conjunction with all (or a portion) of the multiple cores.

[0049] If the method according to an embodiment of the present disclosure includes multiple operations, the multiple operations can be performed by one of the multiple cores included in the multi-core processor, and can be performed by multiple cores. For example, when the first operation, the second operation, and the third operation are performed by the method according to the embodiment, the first operation, the second operation, and the third operation can all be performed by the first core included in the multi-core processor, and the first operation and the second operation can be performed by the first core included in the multi-core processor, and the third operation can be performed by the second core included in the multi-core processor.

[0050] In an embodiment of the present disclosure, at least one processor 130 may represent a system on chip (SoC) on which at least one processor and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, a GPU, an APU, a MIC, a DSP, an NPU, a hardware accelerator, or a machine learning accelerator, but the embodiments of the present disclosure are not limited thereto.

[0051] Specifically, the at least one processor 120 may obtain first audio data including a user voice.

[0052] In addition, the at least one processor 120 may perform acoustic enhancement on the first audio data to obtain the second audio data.

[0053] Furthermore, the at least one processor 120 may obtain a first interval corresponding to audio data between time points including the user's speech among the first audio data and a second interval of the second audio data corresponding to the first interval.

[0054] Next, the at least one processor 120 may calculate a plurality of first scores corresponding to each of the plurality of estimation candidates based on the first audio data corresponding to the first interval.

[0055] Furthermore, the at least one processor 120 may calculate a plurality of second scores corresponding to each of the plurality of estimation candidates based on the second audio data corresponding to the second interval.

[0056] Thereafter, the at least one processor 120 may determine one estimation candidate among the plurality of estimation candidates as character data based on the plurality of first scores and the plurality of second scores.

[0057] Meanwhile, if the length of the second interval is different from the length of the first interval, the at least one processor 120 may calculate a plurality of second scores for each of the plurality of estimation candidates based on the silent interval associated with the second interval.

[0058] In addition, if the length of the second interval is longer than the length of the first interval, the at least one processor 120 may calculate a plurality of second scores for each of the plurality of estimation candidates based on the silent interval included in the second interval.

[0059] In addition, if the length of the second interval is shorter than the length of the first interval, the at least one processor 120 may calculate a plurality of second scores for each of the plurality of estimation candidates based on the silent interval removed from the second interval.

[0060] Meanwhile, the at least one processor 120 may calculate a plurality of first scores associated with the plurality of estimation candidates based on character data determined at a time point before a time point corresponding to the first interval.

[0061] Furthermore, the at least one processor 120 may calculate a plurality of second scores associated with the plurality of estimation candidates based on character data determined at a time point before a time point corresponding to the second interval.

[0062] Meanwhile, the at least one processor 120 may add each of the plurality of first scores and the plurality of second scores to calculate a plurality of third scores corresponding to each of the plurality of estimation candidates.

[0063] Furthermore, the at least one processor 120 may determine, as character data, an estimation candidate corresponding to a highest score among a plurality of third scores among a plurality of estimation candidates.

[0064] Meanwhile, the at least one processor 120 may determine, as character data, an estimation candidate among the plurality of estimation candidates corresponding to a highest score among the plurality of first scores and the plurality of second scores.

[0065] Meanwhile, the at least one processor 120 may perform acoustic enhancement on the first audio data based on at least one of velocity disturbance, amplitude disturbance, vocal tract length disturbance (VTLP), and pitch disturbance with respect to the first audio data to obtain second audio data.

[0066] Figure 3 is a block diagram showing a detailed configuration of an electronic device according to an embodiment.

[0067] refer to Figure 3 , the electronic device 100 may include a memory 110, at least one processor 120, a microphone 130, a display 140, a communication interface 150, an input interface 160, and a speaker 170. Figure 2 A detailed description of the overlapping parts of the description.

[0068] The microphone 130 may refer to a module that obtains sound and converts the sound into an electric signal, and may be a condenser microphone, a ribbon microphone, a dynamic microphone, a piezoelectric element microphone, a carbon microphone, or a micro-electromechanical system (MEMS) microphone. In addition, the microphone may be implemented in an omnidirectional method, a bidirectional method, a unidirectional method, a sub-cardioid method, a super-cardioid method, or a hyper-cardioid method.

[0069] Specifically, the microphone 130 may receive audio data including user speech. Here, the audio data may include noise caused by the user speech and the surrounding environment.

[0070] The display 140 may include various types of display panels such as a liquid crystal display (LCD) panel, an organic light emitting diode (OLED) panel, an active matrix organic light emitting diode (AM-OLED) panel, a liquid crystal on silicon (LcoS) panel, a quantum dot light emitting diode (QLED) panel, and a digital light processing (DLP) panel, a plasma display panel (PDP), an organic LED panel, and a micro LED panel, but is not limited thereto. Meanwhile, the display 140 may be configured with a touch panel together with a touch screen, and may be configured by a flexible panel.

[0071] Specifically, the display 140 may display character data determined by the electronic device 100 recognizing a voice with respect to audio data.

[0072] The communication interface 150 may include a wireless communication interface, a wired communication interface, or an input interface. The wireless communication interface may perform communication with various external devices by using wireless communication technology or mobile communication technology. The wireless communication technology may include, for example, Bluetooth, Bluetooth low energy, CAN communication, Wi-Fi, Wi-Fi direct, ultra-wideband (UWB), Zigbee, infrared data association (IrDA), or near field communication (NFC), and the mobile communication technology may include 3GPP, Wi-Max, long-term evolution (LTE), or 5G. The wireless communication interface may be implemented by using an antenna, a communication chip, a substrate, etc. that can send electromagnetic waves to the outside or can receive electromagnetic waves sent from the outside. Specifically, the communication interface 150 may obtain an image or receive a mobile state specified by a user. Specifically, the communication interface 150 may receive audio data. In addition, the electronic device 100 may send audio data to an external electronic device or server through the communication interface 150, and receive character data determined by recognizing the voice of the sent audio data from the external electronic device or server.

[0073] The input interface 160 may include a circuit and receive a user command for setting or selecting various functions supported by the electronic device 100. For the above, the input interface 160 may include a plurality of buttons and be implemented as a touch screen capable of simultaneously performing the function of a display.

[0074] In this case, the at least one processor 130 may control the operation of the electronic device 100 based on a user command input through the input interface 160. For example, the at least one processor 130 may control the electronic device 100 based on an on / off command of the electronic device 100, an on / off command of a function of the electronic device 100, etc. input through the input interface 160.

[0075] Specifically, the input interface 160 may select the operation of the decoder. That is, when the estimation candidate is determined as character data according to the user input, the input interface 160 may change the method of calculating the score.

[0076] The speaker 170 may output audio sounds. Specifically, the at least one processor 130 may output various alarm sounds or voice guide messages regarding the operation of the electronic device 100 through the speaker 170.

[0077] Specifically, if character data is determined as a result of voice recognition, the speaker 170 may output the determined character data.

[0078] Figure 4 is a view illustrating a voice recognition process of an electronic device according to an embodiment.

[0079] refer to Figure 4, at least one processor 120 may obtain audio data including user voice (S401). Here, at least one processor 120 may obtain audio data received through a microphone, or may receive audio data from a server or an external electronic device through a communication interface. Here, the audio data may include user voice and noise.

[0080] In addition, at least one processor 120 may obtain a plurality of audio data on which acoustic enhancement is performed (S402). Here, at least one processor 130 may apply at least one of velocity disturbance, amplitude disturbance, VTLP, and pitch disturbance to the audio data to obtain a plurality of audio data on which acoustic enhancement is performed. That is, each of the plurality of audio data may be audio data to which a separate acoustic enhancement method is applied. As described above, the method of performing acoustic enhancement on audio data is an example, and the present disclosure is not limited thereto.

[0081] Then, at least one processor 120 may input the plurality of audio data into a speech recognition model (S403). Here, the speech recognition model may be an end-to-end speech recognition model. Here, the end-to-end speech recognition model may express combined information between speech / languages ​​while reducing system complexity by using a single deep neural network. Examples of end-to-end speech recognition models may include Connectionist Temporal Classification (CTC), an attention model, an RNN converter, and the like.

[0082] The speech recognition model may include an encoder. Here, the encoder may obtain information in which speech features included in the input data for speech recognition are converted into vectors on a latent space suitable for speech recognition. Here, the vector on the latent space may be a set of feature values ​​of audio data for calculating scores by an AI neural network.

[0083] In addition, the speech recognition model can calculate the score of the character data corresponding to the audio data to be recognized at the current time point. Here, the score of the character data corresponding to the audio data to be recognized at the current time point can be calculated taking into account the score of the character data determined until the previous time point. For example, the score of the character data corresponding to the audio data to be recognized at the current time point can be expressed based on conditional probability.

[0084] In addition, at least one processor 120 may input the output of the speech recognition model to the decoder (S404). That is, the decoder may receive the score of the character data corresponding to the audio data to be currently recognized. Here, the decoder may calculate the score of the character data, wherein the character data determined at the previous time point and the character data corresponding to the audio data to be recognized at the current time point are combined. Then, at least one processor 120 may determine the character data based on the score calculated by the decoder. As described above, the decoder may correspond to a beam search process.

[0085] In the following, reference Figure 5 and Figure 6 , describes a process of performing acoustic enhancement on audio data to determine character data and determining the character data based on a calculated score.

[0086] Figure 5 is a view illustrating a process of determining character data of an electronic device according to an embodiment.

[0087] refer to Figure 5 , at least one processor 120 may obtain first audio data including a user voice (S501). Here, at least one processor 120 may obtain the first audio data through a microphone or obtain the first audio data through a communication interface.

[0088] In addition, at least one processor 120 may obtain the second audio data by acoustically enhancing the first audio data. Here, the acoustic enhancement may correspond to at least one of a speed disturbance, an amplitude disturbance, a VTLP, and a pitch disturbance. This is an example of acoustic enhancement and is not limited thereto.

[0089] Here, the length of the second audio data on which acoustic enhancement is performed on the first audio data may be different from the length of the first audio data. As described above, if the length of the first audio data is different from the length of the second audio data, the processor may create an interval to determine the character data corresponding thereto and calculate a score about the relevant interval.

[0090] Therefore, at least one processor 120 can obtain a first interval corresponding to the audio data between the time points including the user's speech in the first audio data and a second interval corresponding to the first interval of the second audio data. Here, the time point including the user's speech in the first audio data can be the next time point immediately following the time point at which the character data corresponding to the first audio data is determined in the entire first audio data. That is, the time point including the user's speech in the first audio data can correspond to the time point at which speech recognition is performed on the first audio data. In the following, reference is made to Figure 6 , describing a second interval of the second audio data corresponding to the first interval of the first audio data.

[0091] Figure 6 is a view illustrating a process of obtaining a plurality of first intervals and a plurality of second intervals according to an embodiment.

[0092] refer to Figure 6 , the first audio data 610 may be audio data including the user's voice, and may be original audio data without any data transformation. In addition, the time point for determining the character data about the first audio data 610 may be t. Therefore, the character data corresponding to the first audio data 610 before the time point t may be determined. Here, the fact of determining the character data may mean that the speech recognition of all and / or part of the audio data has been completed. Here, the character data corresponding to the first audio data 610 before the time point t may be determined as "Hi".

[0093] Therefore, at least one processor 120 can then determine the character data about the specific interval from the time point t to determine the character data about the first audio data 610. Here, the specific interval can be composed of a plurality of frames. Here, a frame can represent audio data divided by a specific time interval. A frame can be composed of audio data in units of 10ms to 20ms. As described above, a specific interval can be composed of a plurality of frames divided in units of 10ms to 20ms. Here, if the time point at which a frame starts is t, the time point at which a frame ends can be represented by t+1.

[0094] That is, the first interval corresponding to the audio data between the time points including the user's speech in the first audio data 610 may be an interval of time points configured with a plurality of frames. That is, if the first interval consists of two frames, the first interval may correspond to a time frame from time point t to time point t+2, and if the first interval consists of three frames, the first interval may correspond to a time frame from time point t to time point t+3. Hereinafter, for ease of description, the first interval consists of two frames and is an interval of time points from time point t to time point t+2.

[0095] The second audio data is audio data on which acoustic enhancement is performed on the first audio data, and therefore, the length of the second audio data may be different from the length of the first audio data. For example, if acoustic enhancement is performed on the first audio data by speed perturbation to reduce the speed of the first audio data by two times, the data length of the second audio data 620 may be twice the data length of the first audio data 610.

[0096] If the data length of the second audio data 620 can be twice the data length of the first audio data 610, the second interval corresponding to the first interval can be an interval ranging from time point t0 to a time point of four frames. That is, the second interval can correspond to an interval from time point t0 to a time point of t0+4.

[0097] In addition to this, if acoustic enhancement is performed on the first audio data through speed perturbation to increase the speed of the first audio data by two times, the data length of the second audio data 640 may be two times shorter than the data length of the first audio data.

[0098] If the data length of the second audio data 640 is twice shorter than the data length of the first audio data 610, the second interval corresponding to the first interval may be an interval ranging from time point t0 to a time point of four frames. That is, the second interval may be an interval from time point t0 to a time point of t0+4.

[0099] The at least one processor 120 may determine character data “B” corresponding to the first interval based on the first interval corresponding to the first audio data from time point t to time point t+2 (ie, after “Hi” is determined before time point t).

[0100] In the following, reference Figure 6 The process of calculating the score for determining the character data "B" is described.

[0101] In addition, at least one processor 120 may calculate a plurality of first scores corresponding to each of a plurality of estimation candidates based on the first audio data 610 corresponding to the first interval (S504). Here, the estimation candidate may correspond to character data, which may be included in the first audio data 610 corresponding to the first interval. That is, the plurality of estimation candidates may be a plurality of character data having a possibility of being included in the first audio data 610 corresponding to the first interval. For example, the character data that may be included in the first audio data 610 corresponding to the first interval may include "A", "B", "C", ... "Z".

[0102] Here, the at least one processor 120 may calculate a plurality of first scores corresponding to each of the plurality of estimation candidates based on a score output by inputting the first audio data 610 corresponding to the first interval into the trained speech recognition model. Here, the at least one processor 120 may calculate a plurality of first scores associated with the plurality of estimation candidates based on character data determined at a time point prior to the time point corresponding to the first interval.

[0103] Specifically, the first audio data 610 corresponding to the first interval is the audio data to be currently recognized, wherein at least one processor 120 may input the first audio data 610 corresponding to the first interval into the trained speech recognition model to obtain the score of the character data corresponding to the audio data to be currently recognized, that is, the conditional probability of the audio data to be currently recognized under the condition of the character data determined until the previous time point, as the output score of the speech recognition model. For example, the output score of the speech recognition model may be expressed as P ("B" | "Hi", x_t), which is the probability of determining "B" at time point t under the condition that "Hi" is determined before time point t. Here, x_t is the first audio data 610 corresponding to the first interval, and may be the audio data from t to t+2 among the first audio data 610.

[0104] In addition, at least one processor 120 may calculate a plurality of first scores corresponding to each of the plurality of estimation candidates based on the output score of the speech recognition model. Here, the plurality of first scores corresponding to each of the plurality of estimation candidates may be scores not only about the plurality of estimation candidates themselves, but also about the form in which the character data determined until the previous time point and the estimation candidate are combined. For example, if "Hi" is determined before time point t and the estimation candidate is "B", the score corresponding to the estimation candidate "B" may be calculated as P ("Hi B").

[0105] Here, in the case of calculating the score corresponding to the estimated candidate, the silent interval may be further considered. Here, the silent interval may be an interval in the audio data that does not include the user's speech. Here, the character data corresponding to the silent interval may be represented as blank data. As described above, the blank data may be included at the end of the determined character data in the form of Φ. For example, if "Hi" is determined before time point t and the estimated candidate is "B", the score corresponding to the estimated candidate "B" may be calculated as P ("HiΦBΦ").

[0106] In addition, at least one processor 120 can calculate multiple second scores corresponding to each of the multiple estimation candidates based on the second audio data corresponding to the second interval. Here, at least one processor 120 can calculate multiple second scores related to the multiple estimation candidates based on character data determined at a time point before the time point corresponding to the second interval.

[0107] Specifically, the second audio data corresponding to the first interval is the audio data to be currently recognized, wherein at least one processor 120 may input the second audio data corresponding to the second interval into the trained speech recognition model to obtain the score of the character data corresponding to the audio data to be currently recognized, that is, the conditional probability of the audio data to be currently recognized under the condition of the character data determined until the previous time point, as the output score of the speech recognition model. For example, the output score of the speech recognition model may be expressed as P ("B"|"Hi", x_t0), which is the probability of determining "B" at time point t0 under the condition that "Hi" is determined before time point t0. Here, x_t0 is the second audio data 620 corresponding to the second interval, and may be the audio data from t0 to t0+4 in the second audio data 620.

[0108] In addition, at least one processor 120 may calculate a plurality of second scores corresponding to each of the plurality of estimation candidates based on the output score of the speech recognition model. Here, the plurality of second scores corresponding to each of the plurality of estimation candidates may be scores not only about the plurality of estimation candidates themselves, but also about the form in which the character data determined until the previous time point and the estimation candidate are combined. For example, if "Hi" is determined before time point t0 and the estimation candidate is "B", the score corresponding to the estimation candidate "B" may be calculated as P ("Hi B").

[0109] Here, in the case of calculating the score corresponding to the estimated candidate, the silent interval can be further considered. Here, the silent interval can be a part of the audio data that does not include the user's speech. Here, the character data corresponding to the silent interval can be represented as blank data. As described above, the blank data can be included in the form of Φ at the end of the determined character data. At the same time, the second audio data is audio data that performs acoustic enhancement on the first audio data, wherein the data length may be different. In this case, even if the problematic data is determined at the same point, the length of the first interval and the length of the second segment may be different.

[0110] Therefore, if the length of the second interval is different from the length of the first interval, the at least one processor 120 may calculate multiple second scores for each of the multiple estimation candidates based on the silent interval associated with the second interval. That is, if the length of the second interval is longer than the length of the first interval, the at least one processor 120 needs to further consider the blank data, and if the length of the second interval is shorter than the length of the first interval, the at least one processor 120 may further consider the removed blank data to calculate the multiple second scores.

[0111] As described above, at least one processor may determine whether the length of the second interval is longer than the length of the first interval ( S505 ).

[0112] If the length of the second interval is longer than the length of the first interval, at least one processor 120 may calculate each of the multiple second scores based on the silence interval included in the second interval and the multiple estimation candidates (S506). That is, if the speed of the second audio data is reduced by acoustic enhancement, the second interval corresponding to the first interval may be configured to be longer than the length of the first interval. Therefore, the second interval includes a silence interval longer than the silence interval included in the first interval, wherein if at least one processor 120 calculates the score while considering the silence interval longer than the first interval, the estimation candidate may be determined as character data while considering all of the multiple first scores for the first interval and the multiple second scores for the second interval, wherein the data lengths of the first interval and the second interval are different.

[0113] For example, in Figure 6 , if the second audio data 620 whose speed is reduced by performing acoustic enhancement on the first audio data 610 is obtained, the at least one processor 120 may determine that the length of the second interval becomes longer than the length of the first interval. Here, if the length of the second interval is longer than the length of the first interval by adding an interval corresponding to the silence interval, the at least one processor 120 may calculate a second score (P(BΦΦ)+P(ΦBΦ)) by adding the probability (P(BΦΦ)) that "B" as one of the plurality of estimation candidates appears at t0 and the probability (P(ΦBΦ)) that "B" appears at t0+1.

[0114] At the same time, if the length of the second interval is shorter than the length of the first interval, at least one processor 120 may calculate multiple second scores for each of the multiple estimation candidates based on the silence interval removed from the second interval (S507). That is, if the speed of the second audio data is increased by acoustic enhancement, the second interval corresponding to the first interval may be configured to be shorter than the length of the first interval. Therefore, the second interval includes a silence interval shorter than the silence interval included in the first interval, or the silence interval may be removed, wherein, if at least one processor 120 calculates the score while considering the silence interval shorter than the silence interval of the first interval or the removed silence interval, the processor may determine the estimation candidate as character data while considering all of the multiple first scores about the first interval and the multiple second scores about the second interval, wherein the data lengths of the first interval and the second interval are different.

[0115] For example, in Figure 6, if the second audio data 640 whose speed is increased by performing acoustic enhancement on the first audio data 610 is obtained, the at least one processor 120 may determine that the length of the second interval becomes shorter than the length of the first interval. Here, if the interval corresponding to the silent interval is removed, and thus the length of the first interval is shorter than the length of the second interval, the at least one processor 120 may determine that the silent interval of the previously determined interval "Hi" overlaps with the silent interval of "B" which is one of the multiple estimation candidates. Therefore, the at least one processor 120 should correct the score about the overlapped silent interval as described above to determine the estimation candidate as character data in consideration of all of the multiple first scores about the first interval and the multiple second scores about the second interval, where the data lengths of the first interval and the second interval are different. That is, if at t2, the score about "HiΦ" is calculated, but at t2+1 ​​where the silent interval is removed, the probability (P(BΦ)) of the occurrence of "BΦ" is calculated, the score is calculated by overlapping Φ corresponding to the silent interval. Therefore, the at least one processor 120 may calculate the second score as (1 / P(Φ))*P(BΦ)).

[0116] Meanwhile, the at least one processor 120 may determine one of the plurality of estimation candidates as character data based on the plurality of first scores and the plurality of second scores. That is, the at least one processor 120 may determine the character data as the final result of the speech recognition up to the current time point.

[0117] Here, the at least one processor 120 may add each of the plurality of first scores and the plurality of second scores to calculate a plurality of third scores corresponding to each of the plurality of estimation candidates (S508). That is, the at least one processor 120 may add a plurality of first scores corresponding to each of the plurality of estimation candidates calculated based on the first audio data as the original data and a plurality of second scores corresponding to each of the plurality of estimation candidates calculated based on the second audio data as the data on which acoustic enhancement is performed to calculate a plurality of third scores.

[0118] Here, as a method for calculating the plurality of third scores, not only the accumulation method but also any operation including at least one of addition, multiplication, division and subtraction can be considered. In addition, the plurality of third scores can be calculated by further considering the weight values ​​of the plurality of first scores and the plurality of second scores.

[0119] In addition, the at least one processor 120 may determine the estimation candidate corresponding to the highest score among the plurality of third scores among the plurality of estimation candidates as character data (S509). As described above, if the character data is determined by calculating the plurality of third scores, the at least one processor 120 may calculate the plurality of third scores including all the first scores and the second scores calculated for one estimation candidate, and may determine the estimation candidate corresponding to the highest score among the plurality of third scores as character data.

[0120] At the same time, at least one processor 120 may determine the estimation candidate corresponding to the highest score among the plurality of first scores and the plurality of second scores among the plurality of estimation candidates as character data. If the character data is determined as described above, at least one processor 120 may determine the estimation candidate corresponding to the highest score among the first score and the second score calculated for one estimation candidate as character data.

[0121] Figure 7 is a flowchart illustrating a method of controlling an electronic device according to an embodiment.

[0122] refer to Figure 7 , the method may include obtaining first audio data including a user voice (S701).

[0123] Furthermore, the method may include performing acoustic enhancement on the first audio data to obtain second audio data ( S702 ).

[0124] Then, the method may include obtaining a first section corresponding to audio data between time points including the user's speech among the first audio data and a second section of the second audio data corresponding to the first section ( S703 ).

[0125] Furthermore, the method may include calculating a plurality of first scores corresponding to each of the plurality of estimation candidates based on the first audio data corresponding to the first section ( S704 ).

[0126] Furthermore, the method may include calculating a plurality of second scores corresponding to each of the plurality of estimation candidates based on the second audio data corresponding to the second section ( S705 ).

[0127] Furthermore, the method may include determining one estimation candidate among the plurality of estimation candidates as character data based on the plurality of first scores and the plurality of second scores ( S706 ).

[0128] Step S704 may include calculating a plurality of first scores associated with the plurality of estimation candidates based on character data determined at a time point before the time point corresponding to the first interval.

[0129] Meanwhile, step S704 may include calculating a plurality of first scores associated with the plurality of estimation candidates based on character data determined at a time point before the time point corresponding to the first interval.

[0130] Step S705 may include calculating a plurality of second scores for each of the plurality of estimation candidates based on a silent interval associated with the second interval if the length of the second interval is different from the length of the first interval.

[0131] Here, if the length of the second interval is longer than the length of the first interval, the step may include calculating a plurality of second scores based on a silent interval included in the second interval and a plurality of estimation candidates.

[0132] Alternatively, if the length of the second interval is shorter than the length of the first interval, the step may include calculating a plurality of second scores for each of the plurality of estimation candidates based on the silent interval removed from the second interval.

[0133] Step S705 may include calculating a plurality of second scores associated with the plurality of estimation candidates based on character data determined at a time point before a time point corresponding to the second interval.

[0134] Meanwhile, step S705 may include calculating a plurality of second scores associated with the plurality of estimation candidates based on character data determined at a time point before the time point corresponding to the second interval.

[0135] Furthermore, step S706 may include adding each of the plurality of first scores and the plurality of second scores to calculate a plurality of third scores corresponding to each of the plurality of estimation candidates.

[0136] Then, the step may include determining, as character data, an estimation candidate corresponding to a highest score among the plurality of third scores among the plurality of estimation candidates.

[0137] Meanwhile, step S706 may include determining, as character data, an estimation candidate among the plurality of estimation candidates corresponding to a highest score among the plurality of first scores and the plurality of second scores.

[0138] Meanwhile, step S702 may apply at least one of velocity disturbance, amplitude disturbance, vocal tract length disturbance (VTLP) and pitch disturbance to the first audio data, and perform acoustic enhancement on the first audio data to obtain second audio data.

[0139] The AI-related functions according to the present disclosure are operated by the processor and memory of the electronic device.

[0140] The processor may be configured by one or more processors. Here, the one or more processors may include at least one of a central processing unit (CPU), a graphics processing unit (GPU), and a neural processing unit (NPU), but is not limited to the above examples of the processor.

[0141] The CPU is a general-purpose processor that can perform not only general operations but also AI operations, and can efficiently execute complex programs through a multi-layer cache structure. The CPU is advantageous for a serial processing method by which an organic connection between a previous calculation result and a next calculation result is possible through sequential calculations. The general-purpose processor is not limited to the above-mentioned examples, and does not include the case where the present disclosure specifies it as the above-mentioned CPU.

[0142] A GPU is a processor for a large number of operations (such as floating point operations for graphics processing), and can integrate cores on a large scale to perform a large number of operations in parallel. Specifically, compared with a CPU, a GPU may be advantageous for parallel processing methods such as convolution operations. In addition, a GPU may be used as a coprocessor to supplement the functions of a CPU. The processor for a large number of operations is not limited to the above examples, and does not include the case where the present disclosure specifies it as the above-mentioned GPU.

[0143] NPU is a processor specific to AI operations using artificial neural networks, where each layer configuring the artificial neural network can be implemented as hardware (e.g., silicon). Here, the NPU is designed to be specific to the specifications required by the manufacturer, and therefore its degree of freedom is lower than that of the CPU or GPU, but it can effectively perform the AI ​​operations required by the manufacturer. At the same time, as a processor specific to AI operations, the NPU can be implemented in various forms such as a tensor processing unit (TPU), an intelligent processing unit (IPU), or a visual processing unit (VPU). The AI ​​processor is not limited to the above examples, and does not include the case where the present disclosure specifies it as the above-mentioned NPU.

[0144] In addition, one or more processors may be implemented as a system on chip (SoC). Here, the SoC may further include a memory and a network interface such as a bus for data communication between the processor other than the one or more processors and the memory.

[0145] If the SoC included in the electronic device includes a plurality of processors, the electronic device may perform AI-related operations (e.g., operations related to learning or reasoning of an AI model) by using some of the processors among the plurality of processors. For example, the electronic device may perform AI-related operations by using at least one of a GPU, an NPU, a VPU, a TPU, and a hardware accelerator specific to AI operations (such as convolution operations or matrix product calculations) among the plurality of processors. Meanwhile, this is merely an example, and it is apparent that AI-related operations may be processed by using a general-purpose processor such as a CPU.

[0146] In addition, the electronic device can perform operations related to AI-related functions by using multiple cores (e.g., dual cores, quad cores) included in one processor. Specifically, the electronic device can perform AI operations such as convolution operations and matrix product calculations in parallel by using multiple cores included in the processor.

[0147] One or more processors can control the electronic device to process input data according to predefined operating rules or AI models stored in the memory. The predefined operating rules or AI models are constructed through learning.

[0148] Here, building by learning means building a predefined operating rule or an AI model with desired characteristics by applying a learning algorithm to various learning data. The learning can be performed in the device itself that executes the AI ​​according to the present disclosure, and can also be performed by a separate server / system.

[0149] The AI ​​model can be composed of multiple neural network layers. At least one layer has at least one weight value, and the operation of the layer is performed by the operation result of the previous layer and at least one defined operation. Examples of neural networks are convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), deep Q networks, or transformers, wherein the neural networks of the present disclosure are not limited to the above examples, excluding the case where the neural network is specified as the above examples.

[0150] A learning algorithm is a method for training a given target device (e.g., a robot) by using a plurality of learning data so that the given target device can make or predict decisions by itself. Examples of learning algorithms are supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, wherein the learning algorithm of the present disclosure is not limited to the above examples, excluding the case where the learning algorithm is specified as the above examples.

[0151] Meanwhile, according to an embodiment of the present disclosure, the above-mentioned various examples may be implemented as software including instructions stored in a machine (e.g., computer) readable storage medium. A machine may refer to a device that calls instructions stored in a storage medium and can operate according to the called instructions, wherein the machine may include a device according to the disclosed embodiment. If the instructions are executed by a processor, the processor may perform functions corresponding to the instructions directly or by using other components under the control of the processor. The instructions may include code generated or executed by a compiler or an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" only means that the storage medium is a tangible device and does not include signals (e.g., electromagnetic waves), wherein the term does not distinguish between a case where data is semi-permanently stored in a storage medium and a case where data is temporarily stored in a storage medium. For example, a "non-transitory storage medium" may include a buffer that temporarily stores data.

[0152] According to an embodiment, the method according to various examples disclosed in the present disclosure may be provided to be included in a computer program product. The computer program product may be traded between a seller and a buyer as a commodity. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or via an application store (e.g., a play store). TM ) online distribution (e.g., download or upload), or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be at least temporarily stored in a machine-readable storage medium, or may be temporarily generated in a machine-readable storage medium, such as a memory of a manufacturer's server, an application store's server, or a relay server.

[0153] Above, preferred examples of the present disclosure are shown and described. However, it is obvious that the present disclosure is not limited to the aforementioned specific examples, and those skilled in the art can implement various modifications without departing from the main idea of ​​the present disclosure claimed in the scope of the claims, wherein these modifications should not be understood independently from the technical spirit or prospect of the present disclosure.

Claims

1. An electronic device, include: a memory for storing instructions; and at least one processor for executing instructions, Wherein, the at least one processor is configured to: Obtaining first audio data including user voice; Obtaining second audio data by acoustically enhancing the first audio data; Obtain a first interval corresponding to audio data between time points including the user's speech in the first audio data and a second interval of the second audio data corresponding to the first interval; calculating a plurality of first scores corresponding to each of the plurality of estimation candidates based on the first audio data corresponding to the first interval; calculating a plurality of second scores corresponding to each of the plurality of estimation candidates based on second audio data corresponding to a second interval; and One estimation candidate among the plurality of estimation candidates is determined as character data based on the plurality of first scores and the plurality of second scores.

2. The electronic device according to claim 1, in, The at least one processor is configured to: Based on the difference between the length of the second interval and the length of the first interval, a plurality of second scores for each of the plurality of estimation candidates are calculated based on a silent interval associated with the second interval.

3. The electronic device according to claim 2, in, The at least one processor is configured to: Based on the length of the second interval being longer than the length of the first interval, each of the plurality of second scores is calculated based on the silent interval included in the second interval and the plurality of estimation candidates.

4. The electronic device according to claim 2, in, The at least one processor is configured to: Based on the length of the second interval being shorter than the length of the first interval, a plurality of second scores for each of the plurality of estimation candidates are calculated based on the silent interval removed from the second interval.

5. The electronic device according to claim 1, in, The at least one processor is configured to: A plurality of first scores associated with the plurality of estimation candidates are calculated based on character data determined at a time point before a time point corresponding to a first interval.

6. The electronic device according to claim 1, in, The at least one processor is configured to: A plurality of second scores associated with the plurality of estimation candidates are calculated based on character data determined at a time point before a time point corresponding to a second interval.

7. The electronic device according to claim 1, in, The at least one processor is configured to: adding the plurality of first scores and each of the plurality of second scores to calculate a plurality of third scores corresponding to each of the plurality of estimation candidates; as well as An estimation candidate corresponding to a highest score among the plurality of third scores among the plurality of estimation candidates is determined as character data.

8. The electronic device according to claim 1, in, The at least one processor is configured to: An estimation candidate among the plurality of estimation candidates corresponding to a highest score among the plurality of first scores and the plurality of second scores is determined as character data.

9. The electronic device according to claim 1, in, The at least one processor is configured to: At least one of velocity perturbation, amplitude perturbation, vocal tract length perturbation (VTLP) or pitch perturbation is applied to the first audio data, and acoustic enhancement is performed on the first audio data to obtain second audio data.

10. A method of controlling an electronic device, include: Obtaining first audio data including user voice; Obtaining second audio data by acoustically enhancing the first audio data; Obtain a first interval corresponding to audio data between time points including the user's speech in the first audio data and a second interval of the second audio data corresponding to the first interval; calculating a plurality of first scores corresponding to each of the plurality of estimation candidates based on the first audio data corresponding to the first interval; calculating a plurality of second scores corresponding to each of the plurality of estimation candidates based on second audio data corresponding to a second interval; as well as One estimation candidate among the plurality of estimation candidates is determined as character data based on the plurality of first scores and the plurality of second scores.

11. The method according to claim 10, in, Calculating the plurality of second scores comprises: Based on the difference between the length of the second interval and the length of the first interval, a plurality of second scores for each of the plurality of estimation candidates are calculated based on a silent interval associated with the second interval.

12. The method according to claim 11, in, Calculating the plurality of second scores comprises: Based on the length of the second interval being longer than the length of the first interval, each of the plurality of second scores is calculated based on the silent interval included in the second interval and the plurality of estimation candidates.

13. The method according to claim 11, in, Calculating the plurality of second scores comprises: Based on the length of the second interval being shorter than the length of the first interval, a plurality of second scores for each of the plurality of estimation candidates are calculated based on the silent interval removed from the second interval.

14. The method according to claim 10, in, Calculating the plurality of first scores comprises: A plurality of first scores associated with the plurality of estimation candidates are calculated based on character data determined at a time point before a time point corresponding to a first interval.

15. A non-transitory computer-readable recording medium storing computer instructions, which, when executed by a processor of an electronic device, cause the electronic device to: Obtaining first audio data including user voice; Performing acoustic enhancement on the first audio data to obtain second audio data; Obtain a first interval corresponding to audio data between time points including the user's speech in the first audio data and a second interval of the second audio data corresponding to the first interval; calculating a plurality of first scores corresponding to each of the plurality of estimation candidates based on the first audio data corresponding to the first interval; calculating a plurality of second scores corresponding to each of the plurality of estimation candidates based on second audio data corresponding to a second interval; as well as One estimation candidate among the plurality of estimation candidates is determined as character data based on the plurality of first scores and the plurality of second scores.