Electronic device and operation method thereof

The electronic device uses multiple microphones and advanced signal processing to enhance voice recognition accuracy by identifying call words and adjusting orientation, addressing performance degradation from user movement.

JP2025131477APending Publication Date: 2025-09-09HYUNDAI MOTOR CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024106626
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-28
Filing Date
2024-07-02
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing voice recognition systems in electronic devices face degradation in performance when users move or change orientation, necessitating improved methods to track user location and orientation for accurate voice recognition.

Method used

An electronic device equipped with multiple microphones, a processor, and a memory that performs voice recognition on a first audio signal, identifies a designated call word, tracks user position using phase differences, enhances target voice signals, and reduces noise, allowing for accurate voice recognition and device orientation adjustment.

Benefits of technology

The system effectively identifies call words and adjusts device orientation, enhancing voice recognition accuracy and reducing computational load by processing audio signals efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025131477000001_ABST
    Figure 2025131477000001_ABST
Patent Text Reader

Abstract

To provide an electronic device for identifying a call word and an operation method thereof.SOLUTION: An electronic device includes a plurality of microphones, a speaker, a processor and a memory. The processor is configured to acquire a plurality of voice signals by using the plurality of microphones, to identify a designated call word by performing voice recognition on a first voice signal acquired from a first microphone in the plurality of voice signals in accordance with a flow of time in which the plurality of voice signals is acquired, to acquire a second time point which is before utterance time corresponding to the designated call word from a first time point in which identification of the designated call word is completed and to voice-recognize a partial portion acquired from the second time point in the plurality of voice signals.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to electronic devices and methods of operation thereof, and more particularly to techniques for identifying audio signals. [Background technology]

[0002] A voice recognition service is a service based on technology that converts voice signals into text or commands that an electronic device can process. Voice recognition services typically involve an electronic device processing what a user says and converting the content into text, or recognizing voice commands to perform a specific task. For an electronic device to provide a voice recognition service, sound source localization or beamforming technology is required to track the user's location. However, if a user speaks while moving, a fixed beamforming direction can cause a rapid degradation in voice recognition performance. Therefore, there is a need to research ways to change the location and / or orientation of an electronic device depending on the user's location. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2022-128579 Summary of the Invention [Problem to be solved by the invention]

[0004] SUMMARY OF THE INVENTION The present invention has been made in view of the above-mentioned prior art, and an object of the present invention is to provide an electronic device for identifying an invocation word and a method of operating the same. [Means for solving the problem]

[0005] In order to achieve the above object, one aspect of the present invention provides an electronic device comprising a plurality of microphones, a speaker, a processor, and a memory, wherein the processor is configured to acquire a plurality of audio signals using the plurality of microphones, identify a designated call word by performing voice recognition on a first audio signal acquired from a first microphone among the plurality of audio signals in accordance with the time flow of acquiring the plurality of audio signals, acquire a second time point that is before an utterance time corresponding to the designated call word from a first time point at which identification of the designated call word is completed, and perform voice recognition on a portion of the plurality of audio signals acquired from the second time point.

[0006] In one embodiment, the processor is configured to, based on performing voice recognition on the portion, track a user's position associated with the plurality of voice signals using a phase difference between the plurality of voice signals, and identify a second voice signal from the acquired plurality of voice signals, in which a target voice corresponding to the user's speech is enhanced, by beamforming using the plurality of microphones toward the user's position. In one embodiment, the processor is configured to identify a noise signal included in the second audio signal based on identifying the second audio signal, and reduce the amplitude of the noise signal. In one embodiment, the processor is configured to identify a target audio signal indicative of the target audio based on identifying the second audio signal, increase an amplitude of the target audio signal included in the second audio signal, and perform speech recognition using the second audio signal with the increased amplitude of the target audio signal. In one embodiment, the processor is configured to reduce the amplitude of a reference signal output from the speaker in each of the plurality of audio signals, and track the position of the user using a phase difference between the plurality of audio signals with the reference signal having the reduced amplitude. In one embodiment, the processor is configured to identify, within the first audio signal, a reference signal indicative of designated text and a noise signal distinct from the reference signal, reduce an amplitude of at least one of the reference signal, the noise signal, or any combination thereof, and identify the designated invocation word using the first audio signal with the at least one amplitude reduced. In one embodiment, the processor is configured to bypass voice recognition for other voice signals among the plurality of voice signals that are distinct from the first voice signal while performing voice recognition for a first voice signal acquired from the first microphone. In one embodiment, the processor is configured to perform speech recognition on the first speech signal using a first data set corresponding to the first speech signal and temporarily suspend deletion of other data sets corresponding to the other speech signals based on a specified data size while bypassing speech recognition on the other speech signals. In one embodiment, the processor is configured to identify text data corresponding to the specified invocation word from the first dataset, obtain the second time point by performing rollback on the first dataset and rollback on the other dataset based on the identified text data, and perform speech recognition using the first dataset and the other dataset as a whole from the second time point. In one embodiment, the processor is configured to identify a first microphone of the plurality of microphones based on a distance from a user associated with the plurality of audio signals, and to acquire the first audio signal using the first microphone.

[0007] In order to achieve the above object, according to one aspect of the present invention, a method for operating an electronic device includes the steps of: acquiring a plurality of audio signals using a plurality of microphones; identifying a designated call word by performing voice recognition on a first audio signal acquired from a first microphone among the plurality of audio signals according to a time flow of acquiring the plurality of audio signals; acquiring a second time point that is before an utterance corresponding to the designated call word from a first time point at which identification of the designated call word is completed; and performing voice recognition on a portion of the plurality of audio signals acquired from the second time point.

[0008] In one embodiment, the step of performing voice recognition on the portion includes: tracking a user's position associated with the plurality of voice signals using a phase difference of the plurality of voice signals based on performing voice recognition on the portion; and identifying a second voice signal having an enhanced target voice corresponding to the user's speech from the acquired plurality of voice signals by beamforming toward the user's position using the plurality of microphones. In one embodiment, identifying the second audio signal includes identifying a noise signal included in the second audio signal and reducing an amplitude of the noise signal. In one embodiment, identifying the second audio signal includes identifying a target audio signal indicative of the target audio, increasing the amplitude of the target audio signal included in the second audio signal, and performing audio recognition using the second audio signal with the increased amplitude of the target audio signal. In one embodiment, the step of tracking the user's position includes the steps of: reducing the amplitude of a reference signal output from a speaker in each of the plurality of audio signals, the reference signal being included; and tracking the user's position using a phase difference between the plurality of audio signals in which the amplitude of the reference signal has been reduced. In one embodiment, identifying the designated invocation word includes identifying, within the first audio signal, a reference signal indicative of designated text and a noise signal distinct from the reference signal; reducing the amplitude of at least one of the reference signal, the noise signal, or any combination thereof; and identifying the designated invocation word using the first audio signal with the at least one amplitude reduced. In one embodiment, the method further includes bypassing voice recognition for other voice signals among the plurality of voice signals that are distinct from the first voice signal while performing voice recognition on the first voice signal acquired from the first microphone. In one embodiment, bypassing speech recognition for the other speech signal includes performing speech recognition for the first speech signal using a first data set corresponding to the first speech signal, and temporarily suspending deletion of the other data set corresponding to the other speech signal based on a specified data size while bypassing speech recognition for the other speech signal. In one embodiment, the method further includes identifying text data corresponding to the specified invocation term from the first dataset; obtaining the second time point by performing rollback on the first dataset and rollback on the other dataset based on the identified text data; and performing speech recognition from the second time point using the first dataset and the other dataset as a whole. In one embodiment, the method further includes identifying a first microphone of the plurality of microphones based on a distance to a user associated with the plurality of audio signals, and acquiring the first audio signal using the first microphone. [Effects of the Invention]

[0009] According to the present invention, a call word can be identified using at least one of the microphones. Furthermore, when a call word is identified, voice recognition can be performed using all of the microphones. Furthermore, when a call word is identified, the location of the electronic device can be changed based on the location where the call word was generated.

[0010] In addition, various other effects can be provided that are understood directly or indirectly throughout this specification. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a block diagram of an example of an electronic device according to an embodiment of the present invention. [Figure 2] 10A and 10B are diagrams illustrating an example of an operation of an electronic device performing voice recognition according to an embodiment of the present invention. [Figure 3] FIG. 2 illustrates an example of a call word preprocessor included in an electronic device according to an embodiment of the present invention. [Figure 4] FIG. 2 illustrates an example of an invocation word recognizer included in an electronic device according to an embodiment of the present invention. [Figure 5] 10A and 10B are diagrams illustrating an example of an operation of an electronic device according to an embodiment of the present invention to identify connected words; [Figure 6] FIG. 2 illustrates an example of a connected word preprocessor included in an electronic device according to an embodiment of the present invention. [Figure 7] 1 is a flowchart illustrating an example of an operation of an electronic device according to an embodiment of the present invention. [Figure 8a] 4A and 4B are diagrams illustrating an example of an operation of an electronic device receiving a voice signal from a user according to an exemplary embodiment of the present invention; [Figure 8b] 4A and 4B are diagrams illustrating an example of an operation of an electronic device receiving a voice signal from a user according to an exemplary embodiment of the present invention; [Figure 9] 4 is a flowchart illustrating an example of a method performed by an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, specific examples of embodiments of the present invention will be described in detail with reference to the drawings.

[0013] When adding reference numerals to components in each drawing, care is taken to use the same numerals as much as possible for the same components even if they are displayed in different drawings. Furthermore, when describing embodiments of the present invention, if a detailed description of related known structures or functions is deemed to obscure understanding of the embodiments of the present invention, the detailed description will be omitted.

[0014] When describing components of an embodiment of the present invention, terms such as "first," "second," "A," "B," "(a)," and "(b)" are used. These terms are used to distinguish a component from other components and do not limit the nature, order, or sequence of the components. Furthermore, unless otherwise defined, all terms used herein, including technical and scientific terms, have the same meaning as commonly understood by those skilled in the art to which the present invention pertains. Terms similar to those defined in commonly used dictionaries should be interpreted as meanings consistent with the meanings they have in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense unless expressly defined in this specification.

[0015] The term "module" used in various embodiments of the present invention includes units embodied as hardware, software, or firmware, and is used interchangeably with terms such as logic, logic block, component, or circuit. A module is an integrated component, or the smallest unit or portion of a component that performs one or more functions. For example, in one embodiment, a module is embodied in the form of an ASIC (application-specific integrated circuit). According to various embodiments, operations performed by a module, program, or other component may be performed sequentially, in parallel, or iteratively, or one or more of the operations may be performed in a different sequence, omitted, or one or more other operations may be added.

[0016] Various embodiments of the present invention may be embodied as software (e.g., a program) including one or more instructions stored in a storage medium (e.g., internal memory or external memory) readable by a machine (e.g., electronic device 100). For example, a processor (e.g., processor 110) of the machine (e.g., electronic device 100) retrieves and executes at least one instruction from the one or more instructions stored in the storage medium. This enables the machine to be operated to perform at least one function according to the retrieved at least one instruction. The one or more instructions may include code generated by a compiler or code executed by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves). This term does not distinguish between data being stored semi-permanently or temporarily on the storage medium.

[0017] Hereinafter, an embodiment of the present invention will be described in detail with reference to FIGS.

[0018] FIG. 1 is a block diagram of an example of an electronic device according to one embodiment of the present invention.

[0019] The electronic device 100 according to this embodiment includes at least one of a processor 110, a memory 120, a speaker 140, and a plurality of microphones 150. The processor 110, the memory 120, the speaker 140, and the plurality of microphones 150 are electrically and / or operably coupled to each other by electronic components, including a communication bus. Hereinafter, "operably coupled" means that a direct or indirect connection between the hardware is established, either wired or wireless, such that a first piece of hardware controls a second piece of hardware. Although illustrated based on separate blocks, the embodiment is not limited thereto. Part of the hardware in FIG. 1 (e.g., at least a portion of the processor 110, the memory 120, and a communication circuit (not shown)) may be included in a single integrated circuit, such as a system on a chip (SoC). The electronic device 100 may further include components not shown in FIG. 1.

[0020] The processor 110 of the electronic device 100 according to this embodiment includes hardware components for processing data based on one or more instructions. The hardware components for processing data include, for example, an arithmetic and logic unit (ALU), a floating point unit (FPU), a field programmable gate array (FPGA), a central processing unit (CPU), a microcontroller unit (MCU), and / or an application processor (AP). The number of processors 110 may be one or more. For example, the processor 110 may have a multi-core processor architecture, including a dual-core, quad-core, hexa-core, or octa-core architecture.

[0021] The memory 120 of the electronic device 100 according to this embodiment includes hardware components for storing data and / or instructions input and / or output to and from the processor 110. The memory 120 includes volatile memory, such as random-access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM). For example, the volatile memory includes at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). For example, the non-volatile memory includes at least one of programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, a hard disk, a compact disk, and an embedded multimedia card (eMMC). For example, the following components contained within memory 120 (eg, speech recognition module 130, invocation word preprocessor 131, invocation word recognizer 132, connected word preprocessor 133, and / or TTS module 135) are logically partitioned.

[0022] The speaker 140 of the electronic device 100 according to the present embodiment outputs an audio signal. For example, the electronic device 100 acquires an electrical signal. For example, the electronic device 100 converts the electrical signal into an acoustic signal using the speaker 140. For example, the electronic device 100 outputs an audio signal including the converted acoustic signal using the speaker 140.

[0023] The plurality of microphones 150 of the electronic device 100 according to this embodiment receive audio (or voice) signals from the outside. For example, the plurality of microphones 150 are disposed in a portion of the housing of the electronic device 100. As an example, at least one of the plurality of microphones 150 is referred to as a feedforward microphone on a side of the electronic device 100 facing the outside. The plurality of voice signals acquired by the electronic device 100 using the plurality of microphones 150 are divided into channels.

[0024] According to some embodiments, memory 120 of electronic device 100 stores one or more instructions (or command words) that indicate operations and / or actions to be performed on data by processor 110. A set of one or more instructions may be referred to as firmware, an operating system (OS), a process, a routine, a subroutine, and / or an application. For example, electronic device 100 and / or processor 110 may perform at least one of the operations described below in FIGS. 7 and 9 when a set of a plurality of instructions distributed in the form of an operating system (OS), firmware, a driver, and / or an application is executed.

[0025] The electronic device 100 according to the present embodiment uses the voice recognition module 130 to perform operations (or functions) corresponding to a plurality of voice signals acquired through a plurality of microphones 150. The operations corresponding to the plurality of voice signals include functions (or services) performed by the electronic device 100 in response to a user's speech.

[0026] For example, when the electronic device 100 identifies a designated call word (or a call command word) using the voice recognition module 130 (or after identifying the designated call word), the electronic device 100 activates (or executes) a voice recognition service (or a voice recognition function). The voice recognition module 130 receives (or inputs) a plurality of voice signals together with the designated call word or after detecting the designated call word, and recognizes the received voice (e.g., automatic speech recognition (ASR)). For example, using the voice recognition module 130, the electronic device 100 performs language processing (e.g., natural language understanding (NLU)) and / or dialogue management (DM)) on the voice signals.

[0027] For example, the specified invocation word includes text information (or audio information) that causes the electronic device 100 to initiate the performance of a voice recognition function.

[0028] The electronic device 100 according to the present embodiment uses a TTS (text to speech) module 135 to acquire a reference signal to be output from the speaker 140. The reference signal includes an audio signal for guiding a user using a voice recognition service. The reference signal is generated using data acquired from the memory 120 or an external server. The electronic device 100 outputs the generated reference signal through the speaker 140. When the reference signal output from the speaker 140 is acquired through multiple microphones 150, the reference signal is referred to as echo (or noise). The TTS module 135 is used to perform a text to speech function (or service).

[0029] The electronic device 100 according to the present embodiment uses the call word preprocessor 131 to modify a first voice signal among a plurality of voice signals acquired through a plurality of microphones 150. For example, the electronic device 100 removes a noise signal from the first voice signal. The electronic device 100 removes a reference signal from the first voice signal. The operation of the electronic device 100 to modify the first voice signal using the call word preprocessor 131 will be described in more detail below with reference to FIG. 3.

[0030] The electronic device 100 according to the present embodiment uses the call word recognizer 132 to identify a designated call word from the first audio signal modified by the call word preprocessor 131. The call word recognizer 132 is used to perform a KWS (keyword spotting) function (or service) for identifying a designated call word (or keyword). The call word recognizer 132 includes a model that is set (or trained) to identify a predefined (or set) call word. The operation of the electronic device 100 to recognize a designated call word using the call word recognizer 132 will be described in more detail below with reference to FIG. 4.

[0031] The electronic device 100 according to the present embodiment uses the connected word preprocessor 133 to tune (or change) a data set corresponding to a plurality of voice signals identified through the voice recognition module 130. The plurality of voice signals correspond to an invocation word and / or connected words spoken by a user. For example, the connected words indicate a user's command word that is spoken (or acquired) following the invocation word. For example, the electronic device 100 performs a function corresponding to the connected words by identifying the connected words.

[0032] For example, the electronic device 100 enhances a target voice corresponding to a user's voice in the plurality of voice signals using the continuous word preprocessor 133. For example, the electronic device 100 removes noise contained in the plurality of voice signals using the continuous word preprocessor 133. However, the present invention is not limited to this. The operation of the electronic device 100 tuning the plurality of voice signals using the continuous word preprocessor 133 will be described in more detail below with reference to FIG. 6.

[0033] The electronic device 100 according to this embodiment uses multiple microphones 150 to acquire multiple audio signals.

[0034] In this embodiment, the electronic device 100 performs voice recognition on a first voice signal acquired from a first microphone among the plurality of voice signals according to the time course of acquiring the plurality of voice signals.

[0035] For example, the electronic device 100 may identify a first microphone among the plurality of microphones 150 based on the distance from the user associated with the plurality of audio signals. The electronic device 100 may identify a first audio signal acquired using the first microphone. However, the present invention is not limited thereto. For example, the electronic device 100 may set the first microphone to acquire a first audio signal for identifying a specified invocation phrase from the plurality of audio signals.

[0036] For example, the electronic device 100 performs voice recognition on the first voice signal to identify the specified invocation word.

[0037] The electronic device 100 according to the present embodiment identifies a first time point at which the identification of the designated invocation word is completed according to the time flow of acquiring a plurality of voice signals.

[0038] For example, the electronic device 100 identifies a speech time corresponding to a specified call word. The speech time corresponding to the specified call word varies depending on the length of the specified call word. As an example, the electronic device 100 sets a speech time corresponding to the specified call word.

[0039] The electronic device 100 according to the present embodiment acquires a second time point that is before the utterance time corresponding to the specified invocation word from the first time point at which the identification of the specified invocation word is completed.

[0040] For example, the electronic device 100 acquires a plurality of data sets corresponding to the plurality of audio signals, and performs speech recognition on the first audio signal using a first data set corresponding to a first audio signal among the plurality of audio signals.

[0041] For example, the electronic device 100 identifies text data from the first data set that corresponds to the specified invocation term.

[0042] For example, the electronic device 100 performs a rollback on a first data set based on the text data. For example, the electronic device 100 performs a rollback on another data set that is separate from the first data set based on the text data. For example, the electronic device 100 performs a rollback on all of the multiple data sets to obtain a second point in time.

[0043] The electronic device 100 according to the present embodiment performs speech recognition on all of the plurality of voice signals from the second time point onward, for example, the electronic device 100 performs speech recognition on a plurality of data sets rolled back based on text data.

[0044] For example, the electronic device 100 generates the reference signal output from the TTS module 135 based on performing speech recognition on all of the plurality of speech signals from the second time point.

[0045] For example, but not limited to, the reference signal may include information indicative of a response to a user query corresponding to a plurality of audio signals.

[0046] The electronic device 100 according to the present embodiment as described above acquires multiple audio signals from multiple microphones 150. For example, the electronic device 100 identifies a designated call word using a first audio signal of the multiple audio signals with a smaller amount of computation than identifying a designated call word using all of the multiple audio signals. For example, the electronic device 100 performs rollback on multiple data sets corresponding to the multiple audio signals based on identifying the designated call word. For example, the electronic device 100 processes multiple data sets corresponding to a point in time before identifying the designated call word by performing rollback on the multiple data sets. The electronic device 100 processes the rolled-back multiple data sets, thereby improving the accuracy of voice recognition.

[0047] 2 is a diagram illustrating an example of an operation of an electronic device performing speech recognition according to an embodiment of the present invention. Referring to FIG. 2, the electronic device 100 is referred to as the electronic device 100 of FIG. 2. Referring to FIG. 2, one or more filters for processing (or modifying) a plurality of speech signals are included in the call word preprocessor 131, the call word recognizer 132, and / or the connected word preprocessor 133.

[0048] The electronic device 100 according to this embodiment acquires multiple audio signals using multiple microphones (e.g., the multiple microphones 150 in FIG. 1). For example, the first microphone 151 and / or the other microphone 152 are included in the multiple microphones 150 in FIG. 1. For example, the electronic device 100 tunes (or modifies) the first audio signal 210 acquired using the first microphone 151 of the multiple microphones 150 via the call word preprocessor 131.

[0049] The electronic device 100 according to this embodiment performs acoustic echo cancellation (AEC) 201 on the first audio signal 210. The acoustic echo cancellation 201 is used to remove acoustic feedback between the speaker 140 and the plurality of microphones 150.

[0050] For example, when the reference signal 208 output from the speaker 140 is included in the first audio signal 210, the electronic device 100 performs acoustic echo cancellation 201 on the first audio signal 210 to reduce (or remove) the amplitude of the portion of the first audio signal 210 that corresponds to the reference signal 208.

[0051] For example, the reference signal 208 is obtained using at least one data contained in a memory (eg, memory 120 of FIG. 1).

[0052] For example, when the reference signal 208 output from the speaker 140 is included in multiple audio signals acquired using multiple microphones 150, the reference signal 208 is referred to as an echo signal 208-1. The electronic device 100 performs acoustic echo cancellation 201 to reduce echoes included in the audio signals for speech recognition.

[0053] The electronic device 100 according to this embodiment performs noise reduction and residual echo suppression 202 on the first audio signal 210. By performing noise reduction and residual echo suppression 202, the electronic device 100 removes (or reduces) noise and / or residual echo contained in the first audio signal 210. The residual echo is identified based on noise, time, and / or frequency. For example, the noise is distinguished from the reference signal 208 and includes acoustic information generated from outside the electronic device 100.

[0054] The electronic device 100 according to this embodiment identifies a designated call word via the first voice signal 210 using the call word recognizer 132. For example, the electronic device 100 identifies the designated call word by performing voice recognition on the first voice signal 210 acquired using the first microphone 151 of the multiple microphones 150, before identifying the designated call word using the call word recognizer 132. For example, the electronic device 100 bypasses voice recognition on other voice signals acquired using the other microphones 152 of the multiple microphones 150, before identifying the designated call word using the call word recognizer 132.

[0055] For example, before identifying the specified invocation word, the electronic device 100 performs voice recognition on the first voice signal 210 while bypassing voice recognition on other voice signals among the plurality of voice signals that are distinct from the first voice signal 210.

[0056] For example, the electronic device 100 may temporarily store in memory a data set corresponding to the other audio signals while bypassing speech recognition for the other audio signals.

[0057] For example, the electronic device 100 initiates speech recognition using the plurality of audio signals 220 as a whole based on identifying the specified call word using the call word recognizer 132 .

[0058] For example, the electronic device 100 may identify a specified invocation word and roll back a data set corresponding to the plurality of audio signals 220 based on a specified time point. The specified time point may range from the time at which the identification of the specified invocation word is completed to a time before the time corresponding to the specified invocation word. For example, the electronic device 100 may perform speech recognition on the entire plurality of audio signals 220 from the specified time point.

[0059] The electronic device 100 according to this embodiment uses the continuous word preprocessor 133 to perform acoustic echo cancellation 203 on the entire plurality of audio signals 220. The acoustic echo cancellation 203 is referred to as acoustic echo cancellation 201. For example, the acoustic echo cancellation 201 is used to process a data set corresponding to a single channel (or one audio signal). The acoustic echo cancellation 203 is used to process a data set corresponding to multiple channels (or multiple audio signals). The echo signal 208-2 is referred to as echo signal 208-1.

[0060] The electronic device 100 according to this embodiment performs sound source localization and beamforming tracking 204 using the multiple audio signals 220 from which echoes have been removed. For example, the electronic device 100 performs beamforming by tracking the position of a user associated with the multiple audio signals 220. For example, the electronic device 100 identifies the source of the multiple audio signals 220 (e.g., the user's position) using the direction in which each of the multiple microphones 150 faces and / or the position in which each of the multiple microphones 150 is disposed. For example, the electronic device 100 acquires the source of the multiple audio signals 220 using a filter for identifying the source (e.g., a Kalman / particle filter). For example, the electronic device 100 acquires multiple audio signals in which the target voice spoken by the user is enhanced based on identifying the source of the multiple audio signals 220 and performing beamforming.

[0061] For example, electronic device 100 synchronizes multiple audio signals 220 based on a phase difference between multiple audio signals 220 and / or a time difference at which multiple audio signals 220 were acquired. For example, using synchronized multiple audio signals 230, electronic device 100 performs noise reduction and residual echo suppression 205. For example, noise reduction and residual echo suppression 205 is referred to as noise reduction and residual echo suppression 202.

[0062] For example, the electronic device 100 uses an active gain control 206 to enhance the amplitude (or gain) of a region corresponding to a target voice contained in the plurality of audio signals 220, and processes the data set corresponding to the plurality of audio signals 220 in which the target voice has been enhanced using the voice recognition module 130.

[0063] For example, the electronic device 100 uses the TTS module 135 to obtain a reference signal 208 indicative of a response corresponding to speech recognition based on processing a dataset using the speech recognition module 130. The electronic device 100 modifies at least one region of the reference signal 208 obtained using the TTS module 135 by performing an active gain control 207 to enhance the amplitude (or gain) of at least one region of the reference signal 208. The modified reference signal 208 is output via the speaker 140, but is not limited thereto.

[0064] The electronic device 100 according to the present embodiment as described above can reduce the amount of calculation required for data processing by identifying a designated call word based on a single channel (a) compared to identifying a designated call word based on multiple channels. The electronic device 100 can track the location of a user who uttered the designated call word based on identifying the designated call word. The electronic device 100 can obtain a voice signal in which the target voice uttered by the user is enhanced by performing beamforming toward the user's location. The electronic device 100 can perform more accurate voice recognition of the target voice by obtaining a voice signal in which the target voice is enhanced.

[0065] 3 is a diagram illustrating an example of a call word preprocessor included in an electronic device according to one embodiment of the present invention. The electronic device 100 of FIG. 3 is referred to as the electronic device 100 of FIG.

[0066] Referring to FIG. 3, the electronic device 100 according to this embodiment acquires a plurality of audio signals (210, 310) using a plurality of microphones (151, 152).

[0067] For example, the electronic device 100 selects the first microphone 151 among the multiple microphones (151, 152) to recognize the specified invocation word.

[0068] For example, the electronic device 100 performs voice recognition on the first voice signal 210 acquired from the first microphone 151. The electronic device 100 performs call word pre-processing using the first voice signal 210.

[0069] While performing voice recognition on a first voice signal 210 acquired from a first microphone 151, the electronic device 100 according to the present embodiment bypasses voice recognition on another voice signal 310, which is distinct from the first voice signal 210, among a plurality of voice signals (210, 310). The other voice signal 310 is acquired through another microphone 152.

[0070] The electronic device 100 according to the present embodiment performs speech recognition on the first speech signal 210 using a first data set corresponding to the first speech signal.

[0071] The electronic device 100 according to this embodiment temporarily suspends deletion of the other data set corresponding to the other audio signal 310 based on a specified data size (or length) while bypassing voice recognition for the other audio signal 310. The electronic device 100 temporarily stores the other data set. When the electronic device 100 identifies the specified invocation word, it at least temporarily suspends deletion of the other data set in order to use the other data set.

[0072] The electronic device 100 according to this embodiment distinguishes between a reference signal representing specified text (e.g., the reference signal 208 in FIG. 2) and a noise signal distinct from the reference signal within the first audio signal 210. For example, if the reference signal is identified via a microphone, the reference signal is referred to as an echo signal 208-1.

[0073] For example, the reference signal is obtained based on a TTS module (eg, TTS module 135 in FIG. 1).

[0074] The electronic device 100 according to this embodiment reduces the amplitude of at least one of the reference signal (eg, echo signal 208-1), the noise signal, or any combination thereof.

[0075] For example, the electronic device 100 uses a reference signal detector (Ref. signal detector) 320 to identify whether the first audio signal 210 contains a reference signal (eg, echo signal 208-1) output from a speaker.

[0076] For example, if the reference signal is included in the first audio signal 210, the electronic device 100 performs acoustic echo cancellation 201 on the reference signal.

[0077] For example, if the electronic device 100 uses the reference signal detector 320 to identify a reference signal (e.g., echo signal 208-1) having an amplitude greater than or equal to a critical amplitude, the electronic device 100 modifies the first audio signal 210 by applying acoustic echo cancellation 201 to the reference signal.

[0078] Based on identifying the noise signal, the electronic device 100 according to the present embodiment applies noise reduction and residual echo suppression 202 to the first audio signal 210 to remove the noise or reduce the residual echo.

[0079] The electronic device 100 according to the present embodiment identifies a designated call word using a first voice signal in which the amplitude of at least one of a noise signal and an echo signal is reduced. For example, the operation of the electronic device 100 to identify a designated call word by executing a call word recognizer using the first voice signal will be described with reference to FIG. 4.

[0080] The electronic device 100 according to the present embodiment described above can reduce the amount of calculation required to recognize a call word by performing call word pre-processing on one of the multiple voice signals (210, 310) (e.g., the first voice signal 210). The electronic device 100 starts performing acoustic echo cancellation 201 to remove echoes in one voice signal depending on whether the echo signal 208-1 is identified, thereby reducing the amount of calculation (or calculation time) required to perform call word pre-processing.

[0081] 4 is a diagram illustrating an example of an invocation word recognizer included in an electronic device according to an embodiment of the present invention. The electronic device 100 of FIG. 4 is referred to as the electronic device 100 of FIG.

[0082] The electronic device 100 according to the present embodiment performs voice recognition on the first voice signal 210 while bypassing voice recognition on the other voice signals 310. By performing voice recognition on the first voice signal 210, the electronic device 100 identifies the specified call word using the call word recognizer 132.

[0083] For example, the electronic device 100 may use the invocation word recognizer 132 to perform keyword spotting (KWS) in a data set (or text stream) corresponding to the first audio signal 210 to identify text data corresponding to the specified invocation word.

[0084] For example, the electronic device 100 temporarily maintains a saved data set corresponding to the other audio signal 310 while bypassing speech recognition for the other audio signal 310. The electronic device 100 buffers the data set corresponding to the other audio signal 310.

[0085] For example, the electronic device 100 may initiate speech recognition on the entire plurality of audio signals 220 based on identifying a specified invocation word. The electronic device 100 may then perform a rollback on the entire plurality of data sets corresponding to the plurality of audio signals 220, and then perform speech recognition using the rolled-back plurality of data sets. The plurality of audio signals 220 may include the first audio signal 210 and / or the other audio signals 310.

[0086] For example, the electronic device 100 identifies a second time point that is before the utterance time corresponding to the specified invocation word from the first time point when the identification of the specified invocation word is completed, and rolls back the multiple data sets corresponding to the multiple audio signals 220 based on information (or versions) corresponding to the second time point.

[0087] The electronic device 100 according to the present embodiment as described above can improve the accuracy of voice recognition by recognizing a designated invocation phrase and then performing voice recognition using all of a plurality of voice signals corresponding to multi-channels.

[0088] 5 is a diagram illustrating an example of an operation of an electronic device according to an embodiment of the present invention to identify connected words. The electronic device 100 in FIG. 5 is referred to as the electronic device 100 in FIG.

[0089] In state 510, the electronic device 100 according to this embodiment acquires multiple audio signals using multiple microphones.

[0090] For example, electronic device 100 identifies designated call word 515 using a first audio signal of the plurality of audio signals.

[0091] For example, electronic device 100 may acquire a plurality of audio signals over time and perform speech recognition on a first one of the plurality of audio signals to identify designated invocation word 515 .

[0092] For example, the electronic device 100 acquires multiple audio signals over time and acquires multiple data sets 511 (or text streams) corresponding to the multiple audio signals, such that the multiple audio signals are synchronized so that each of the multiple data sets 511 contains substantially the same information at a given point in time.

[0093] For example, electronic device 100 uses a first data set of multiple data sets 511 corresponding to a first audio signal to identify a specified invocation word 515 (eg, halodal in FIG. 5).

[0094] The electronic device 100 according to this embodiment identifies text data corresponding to the specified invocation term 515 from the first data set.

[0095] In this embodiment, at state 520, electronic device 100 identifies a first point in time 521 at which it has completed identifying the specified invocation word 515.

[0096] For example, the electronic device 100 identifies the speech time corresponding to the specified invocation word 515 .

[0097] For example, the electronic device 100 acquires a second time point 531 that is before the utterance time corresponding to the specified invocation word 515 from the first time point 521 .

[0098] The electronic device 100 according to the present embodiment performs rollback on the first data set and rollback on the other data set based on the text data corresponding to the specified invocation word 515. The other data set corresponds to the other audio signals excluding the first audio signal among the plurality of audio signals.

[0099] For example, but not limited to, the electronic device 100 may at least temporarily suspend processing of at least one data set of the plurality of data sets 511 until the first time point 521 when the specified invocation term 515 is identified. In other words, processing (e.g., deletion) of at least one data set is put on hold.

[0100] For example, after identifying the specified invocation word 515, the electronic device 100 performs a rollback on the entire plurality of data sets 511. By performing a rollback on the entire plurality of data sets 511, the electronic device 100 restores the entire plurality of data sets 511 to a version corresponding to before the second point in time 531. The electronic device 100 performs speech recognition from the second point in time 531 using the restored entire plurality of data sets 511.

[0101] In a state 530, the electronic device 100 according to this embodiment performs speech recognition using the first data set and the other data sets as a whole from a second time point 531.

[0102] According to this embodiment, the electronic device 100 performs speech recognition on a portion 535 of the plurality of audio signals acquired from a second time point. For example, the electronic device 100 uses a plurality of data sets 511 corresponding to the plurality of audio signals to perform speech recognition on the portion 535 acquired after the second time point 531. The portion 535 includes text data corresponding to a specified invocation word 515 and / or a sequence of words (or command words) for invoking at least one function (or task) performed by the electronic device 100.

[0103] For example, if portion 535 does not include a sequence of words for invoking at least one function, electronic device 100 enters state 510 from state 530. As an example, if portion 535 does not include a sequence of words for invoking at least one function, electronic device 100 resumes identifying designated invocation word 515 using at least one audio signal of the plurality of audio signals.

[0104] 6 is a diagram illustrating an example of a connected word preprocessor included in an electronic device according to an embodiment of the present invention. The electronic device 100 of FIG. 6 is referred to as the electronic device 100 of FIG.

[0105] The electronic device 100 according to the present embodiment performs a preprocessing operation using the connected word preprocessor 153 to perform speech recognition on all of the plurality of speech signals from a second time point (for example, the second time point 531 in FIG. 5).

[0106] For example, the electronic device 100 removes the echo signal 208-2 included in the rolled-back audio signals at the second time point. For example, the electronic device 100 applies acoustic echo cancellation 203 to the rolled-back audio signals as a whole to identify the audio signals 610 from which the echo signal 208-2 has been removed.

[0107] For example, the echo signal 208 - 2 is included in the reference signal 208 obtained based on the TTS module 135 .

[0108] The electronic device 100 according to the present embodiment performs voice recognition on a portion of the plurality of voice signals (e.g., portion 535 in FIG. 5 ) and tracks the position of the user associated with the plurality of voice signals 610 using a phase difference between the plurality of voice signals 610. The electronic device 100 tracks the position of the user associated with the plurality of voice signals 610 by performing sound source localization and beamforming tracking 204.

[0109] For example, the electronic device 100 reduces the amplitude of the reference signal 208 (or the echo signal 208-2) in each of the plurality of audio signals 610 including the reference signal 208 output from the speaker 140. The electronic device 100 tracks the user's position using the phase difference of the plurality of audio signals 610 with the reduced amplitude of the reference signal.

[0110] For example, the electronic device 100 uses multiple microphones and performs beamforming toward the user's position to identify a second audio signal 615 from multiple audio signals that has an enhanced target audio corresponding to the user's speech.

[0111] For example, the second audio signal 615 is obtained based on the synchronized plurality of audio signals 610. For example, the second audio signal 615 includes at least one audio signal of the synchronized plurality of audio signals 610.

[0112] The electronic device 100 according to the present embodiment identifies a noise signal included in the second audio signal 615 based on the acquired second audio signal 615 .

[0113] For example, the electronic device 100 removes the noise signal using the noise reduction and residual echo suppression 205. For example, the electronic device 100 reduces the amplitude of the noise signal.

[0114] The electronic device 100 according to the present embodiment identifies a target voice signal indicating a target voice corresponding to the user's speech based on the acquired second voice signal 615. For example, the electronic device 100 identifies the target voice signal after removing a noise signal included in the second voice signal 615.

[0115] For example, the electronic device 100 may enhance the target voice signal to improve the accuracy of voice recognition, for example, by increasing the amplitude of the target voice signal included in the second voice signal 615.

[0116] For example, the electronic device 100 performs voice recognition via the voice recognition module 130 using a second voice signal that is an increased amplitude version of the target voice signal.

[0117] After performing speech recognition, the electronic device 100 according to this embodiment generates a reference signal 208 for guiding the user in responding to the speech recognition result via the TTS module 135. The electronic device 100 applies active gain control 207 to at least a portion of the reference signal 208, and then outputs the reference signal 208 using the speaker 140.

[0118] The electronic device 100 according to the present embodiment as described above can improve the accuracy of voice recognition for consecutive words spoken by a user after the call word by identifying the call word and then performing voice recognition on all of the multiple voice signals acquired using multiple microphones.

[0119] FIG. 7 is a flowchart illustrating an example of the operation of the electronic device according to an embodiment of the present invention.

[0120] Hereinafter, it will be assumed that the electronic device 100 of Figure 1 performs the process of Figure 7. Also, in the description of Figure 7, each step described as being performed by the electronic device will be understood to be controlled by the processor 110 of the electronic device 100. Although the steps of Figure 7 are performed sequentially, they are not necessarily performed sequentially. For example, the order of each step may be changed, or at least two steps may be performed in parallel.

[0121] 7, an electronic device according to this embodiment removes echoes included in a first audio signal among a plurality of audio signals in step S710. The first audio signal is acquired through a first microphone among a plurality of microphones for identifying a designated call phrase. The electronic device acquires the plurality of audio signals over time using the plurality of microphones and removes echoes included in the first audio signal among the acquired plurality of audio signals.

[0122] 7, the electronic device according to the present embodiment removes residual echo or noise from the echo-removed first voice signal in step S720. Steps S710 and S720 are included in pre-processing steps for identifying an invocation word.

[0123] 7, the electronic device according to the present embodiment checks whether the calling word is recognized in step S730. For example, if the electronic device does not recognize the calling word (step S730-No), the electronic device performs step S710.

[0124] Referring to FIG. 7, if the electronic device according to this embodiment recognizes the call word using the first voice signal (step S730-Yes), in step S740, it rolls back multi-channel based microphone data (e.g., the multiple data sets 511 in FIG. 5).

[0125] For example, the electronic device rolls back the microphone data from the point at which identification of the call word is completed to the point before the time corresponding to the call word. The electronic device performs voice recognition on the microphone data using the first voice signal based on one channel from step S710 to the point at which identification of the call word is completed.

[0126] For example, after identifying the invocation word, the electronic device performs voice recognition using microphone data corresponding to a plurality of multi-channel based voice signals, and processes the entire plurality of multi-channel based voice signals to perform steps S750 to S790.

[0127] Referring to FIG. 7, the electronic device according to the present embodiment removes echoes contained in a plurality of audio signals in step S750.

[0128] 7, the electronic device according to this embodiment infers and / or tracks a voice direction using a plurality of voice signals from which echoes have been removed in step S760. For example, the electronic device estimates or tracks the voice direction based on a phase difference between the plurality of voice signals. The voice direction includes a direction from the electronic device toward a user's position for voice recognition.

[0129] 7, the electronic device according to the present embodiment performs beamforming based on the voice direction in step S770. Based on the beamforming, the electronic device acquires a plurality of voice signals in which the target voice uttered by the user is enhanced for voice recognition.

[0130] Referring to FIG. 7, the electronic device according to the present embodiment removes residual echo or noise contained in the plurality of audio signals in which the target audio has been enhanced in step S780.

[0131] 7, the electronic device according to the present embodiment recognizes connected words using a plurality of voice signals from which residual echo or noise has been removed in step S790. The connected words indicate voice command words uttered by the user after the invocation word has been identified, and performs a function corresponding to the connected words based on the recognition of the connected words.

[0132] 8a and 8b are diagrams illustrating an example of an operation of an electronic device according to an embodiment of the present invention to receive a voice signal from a user. The electronic device of FIG. 8a and 8b is referred to as the electronic device 100 of FIG. 1.

[0133] Referring to FIG. 8a, in a state 800, the electronic device 100 according to this embodiment is positioned such that one side (for example, the front side) of the electronic device 100 faces a direction 811 in which a user 810 faces.

[0134] For example, the electronic device 100 uses multiple microphones to acquire multiple audio signals indicative of a designated call phrase 820 from the user 810. The electronic device 100 identifies the designated call phrase 820 using a first audio signal acquired using a first microphone of the multiple microphones.

[0135] For example, multiple microphones (e.g., multiple microphones 150 in FIG. 1) are positioned within electronic device 100 facing different directions, allowing electronic device 100 to acquire multiple audio signals generated around electronic device 100.

[0136] For example, the electronic device 100 tracks the location of the user 810 based on identifying a designated call word 820 using the phase difference of multiple audio signals.

[0137] In state 805, the electronic device 100 according to this embodiment changes its orientation toward a direction 812 in which one side (e.g., the front) of the electronic device 100 faces the user 810. For example, the electronic device 100 changes its orientation toward a direction set to acquire an audio signal via a first microphone among multiple microphones arranged facing different directions. For example, the operation of the electronic device 100 changing its orientation is included in the beamforming operation, but is not limited thereto.

[0138] The electronic device 100 according to the present embodiment as described above can improve the accuracy of speech recognition for successive words uttered by the user 810 after the call word 820 by performing beamforming toward the location where the call word 820 is generated based on identifying the call word 820.

[0139] Referring to FIG. 8b, in state 850, when a user 810 who has uttered a call word 820 moves along a movement direction 851, the electronic device 100 according to this embodiment uses all of the multiple microphones to acquire a series of words 855 (or multiple audio signals indicating the series of words) while tracking the movement path of the user 810.

[0140] The electronic device 100 according to the present embodiment as described above acquires the target voice-enhanced consecutive words 855 by changing the direction and / or position of the electronic device 100 based on the identification of the invocation word. The electronic device 100 can improve the accuracy of the voice recognition service by changing the position (or direction) of the electronic device 100 according to the user 810 while the user 810 is moving.

[0141] FIG. 9 is a flowchart illustrating an example of a method performed by an electronic device according to an embodiment of the present invention. Hereinafter, it is assumed that the electronic device 100 of FIG. 1 performs the process of FIG. 9. In the description of FIG. 9, each step described as being performed by the electronic device is understood to be controlled by the processor 110 of the electronic device 100. The steps of FIG. 9 are performed sequentially, but not necessarily sequentially. For example, the order of each step may be changed, or at least two steps may be performed in parallel.

[0142] Referring to FIG. 9, in step S910, the method according to this embodiment includes obtaining a plurality of audio signals using a plurality of microphones.

[0143] For example, the plurality of audio signals are represented as channels based on the number of the plurality of microphones. For example, if the number of the plurality of microphones is four, the method includes acquiring the plurality of audio signals based on multi-channels.

[0144] Referring to FIG. 9, in step S920, the method according to this embodiment includes a step of identifying a designated invocation word by performing voice recognition on a first voice signal acquired from a first microphone among the plurality of voice signals according to a time flow of acquiring the plurality of voice signals.

[0145] For example, the method includes performing speech recognition on a first speech signal while bypassing speech recognition on other speech signals of the plurality of speech signals that are distinct from the first speech signal.

[0146] For example, identifying the designated call word includes removing an echo signal contained in the first audio signal.

[0147] For example, an echo signal indicates acoustic feedback between a speaker and a microphone.

[0148] For example, identifying the designated call phrase may include applying at least one filter (eg, acoustic echo cancellation 201 of FIG. 2) to the first audio signal to remove the echo signal.

[0149] For example, identifying the designated call word may include removing residual echo and / or noise after removing the echo signal from the first audio signal.

[0150] 9, in step S930, the method according to this embodiment includes acquiring a second time point that is before the speech time corresponding to the specified invocation word from the first time point at which the identification of the specified invocation word is completed. The speech time corresponding to the specified invocation word is changed depending on the length of the text included in the specified invocation word.

[0151] 9, in step S940, the method according to this embodiment includes performing speech recognition on a portion of the plurality of speech signals acquired from a second time point. For example, the method may include performing rollback on a plurality of data sets corresponding to the plurality of speech signals after identifying a specified invocation term.

[0152] For example, the method includes performing rollback for multiple data sets based on information obtained prior to identifying the specified invocation word.

[0153] For example, the method may include temporarily buffering other data sets corresponding to other audio signals distinct from the first audio signal among the plurality of audio signals while identifying the invocation word using a first data set corresponding to the first audio signal.

[0154] For example, the method may include performing a rollback on the buffered other data sets and the entire first data set to obtain a rolled back plurality of data sets.

[0155] For example, the portion acquired from the second time point includes text data corresponding to an input (e.g., an invocation word) for activating the voice recognition function and text data corresponding to an input (e.g., a connected word) for the user to receive the voice recognition function.

[0156] For example, the step of performing voice recognition on the portion acquired from the second time point may include, if the source locations of the plurality of voice signals are changed, changing the position (or direction) of the electronic device according to the changed source locations.

[0157] The above description is merely an illustrative example of the technical concept of the present invention, and various modifications and variations can be made by a person having ordinary knowledge in the technical field to which the present invention pertains without departing from the essential characteristics of the present invention.

[0158] Therefore, the embodiments disclosed in this specification are intended to illustrate, not limit, the technical idea of ​​the present invention, and the scope of the technical idea of ​​the present invention should not be limited by such embodiments. The scope of protection of the present invention should be interpreted by the claims, and all technical ideas within the equivalent range should be interpreted as being included in the scope of the present invention. [Explanation of symbols]

[0159] 100 Electronic equipment 110 processors 120 Memory 130 Voice Recognition Module 131 Call word preprocessor 132 Call word recognizer 133, 153 Continuous word preprocessor 135 TTS (text to speech) module 140 speakers 150 microphone 151 First Microphone 152 Other Microphones 201, 203 Acoustic echo cancellation (AEC) 202, 205 Noise reduction and residual echo suppression 204 Sound Source Localization and Beamforming Tracking 206, 207 Active gain control 208 Reference Signal 208-1, 208-2 Echo signal 210, 615 1st and 2nd audio signals 220, 610 audio signals 310 Other Audio Signals 320 Reference signal detector (Ref. signal detector) 510, 520, 530, 800, 805, 850 status 511 datasets 515 specified invocation words 521, 531 1st and 2nd time points 535 Portion acquired from the second point in time 810 User 811 User's direction 812 Direction towards user 820 specified invocation words 851 Direction of movement 855 consecutive words

Claims

1. Multiple microphones and With a speaker, a processor; Memory and preparation, The processor: acquiring a plurality of audio signals using the plurality of microphones; performing voice recognition on a first voice signal acquired from a first microphone among the plurality of voice signals according to a time flow of acquiring the plurality of voice signals, thereby identifying a designated call word; obtaining a second time point that is before an utterance time corresponding to the specified invocation word from a first time point at which the identification of the specified invocation word is completed; an electronic device configured to perform speech recognition on a portion of the plurality of speech signals obtained from the second time point;

2. The processor: tracking a user's position associated with the plurality of voice signals using a phase difference of the plurality of voice signals based on performing voice recognition on the portion; 2. The electronic device of claim 1, further comprising: a second audio signal having an enhanced target audio corresponding to the user's speech from the acquired audio signals by beamforming the plurality of microphones toward the user's position.

3. The processor: Identifying a noise signal included in the second audio signal based on the identification of the second audio signal; 3. The electronic device of claim 2, configured to reduce the amplitude of the noise signal.

4. The processor: identifying a target audio signal indicative of the target audio based on identifying the second audio signal; Increasing the amplitude of the target audio signal included in the second audio signal; 3. The electronic device of claim 2, configured to perform speech recognition using the second audio signal with the target audio signal having an increased amplitude.

5. The processor: reducing an amplitude of a reference signal in each of the plurality of audio signals including the reference signal output from the speaker; 3. The electronic device of claim 2, configured to track the user's position using a phase difference between the plurality of audio signals with the reference signal amplitude reduced.

6. The processor: identifying, within the first audio signal, a reference signal indicative of designated text and a noise signal distinct from the reference signal; reducing the amplitude of at least one of the reference signal, the noise signal, or any combination thereof; 2. The electronic device of claim 1, configured to identify the designated call word using the at least one amplitude-reduced first audio signal.

7. 2. The electronic device of claim 1, wherein the processor is configured to bypass voice recognition for other voice signals distinct from the first voice signal among the plurality of voice signals while performing voice recognition for the first voice signal acquired from the first microphone.

8. The processor: performing speech recognition on the first speech signal using a first data set corresponding to the first speech signal; 8. The electronic device of claim 7, wherein the electronic device is configured to temporarily suspend deletion of other data sets corresponding to the other audio signals based on a specified data size while bypassing speech recognition for the other audio signals.

9. The processor: identifying text data from the first data set corresponding to the specified invocation term; Based on the identified text data, perform a rollback on the first data set and a rollback on the other data set to obtain the second point in time; The electronic device of claim 8 , configured to perform speech recognition using the first data set and the other data set as a whole from the second point in time.

10. The processor: identifying the first microphone of the plurality of microphones based on a distance to a user associated with the plurality of audio signals; 2. The electronic device of claim 1, configured to capture the first audio signal using the first microphone.

11. 1. A method of operating an electronic device, comprising: acquiring a plurality of audio signals using a plurality of microphones; performing voice recognition on a first voice signal acquired from a first microphone among the plurality of voice signals according to a time flow of acquiring the plurality of voice signals, thereby identifying a designated call word; acquiring a second time point that is before an utterance time corresponding to the designated invocation word from a first time point at which the identification of the designated invocation word is completed; and performing speech recognition on a portion of the plurality of speech signals obtained from the second time point.

12. The step of performing speech recognition on the portion includes: tracking a user's location associated with the plurality of voice signals using a phase difference of the plurality of voice signals based on performing voice recognition on the portion; and identifying a second audio signal from the plurality of audio signals acquired by beamforming the plurality of microphones toward the user's position, the second audio signal having an enhanced target audio corresponding to the user's speech.

13. The step of identifying the second audio signal comprises: identifying a noise signal included in the second audio signal; and reducing the amplitude of the noise signal.

14. The step of identifying the second audio signal comprises: identifying a target speech signal indicative of the target speech; increasing the amplitude of the target audio signal included in the second audio signal; and performing speech recognition using the second speech signal with the target speech signal having an increased amplitude.

15. The step of tracking the user's location includes: reducing an amplitude of a reference signal in each of the plurality of audio signals including the reference signal output from a speaker; and tracking the user's position using a phase difference between the plurality of audio signals with the amplitude of the reference signal reduced.

16. The step of identifying the designated invocation word comprises: identifying, within the first audio signal, a reference signal indicative of designated text and a noise signal distinct from the reference signal; reducing the amplitude of at least one of the reference signal, the noise signal, or any combination thereof; and identifying the designated call word using the at least one amplitude-reduced first audio signal.

17. 12. The method of claim 11, further comprising: bypassing voice recognition for other voice signals distinct from the first voice signal among the plurality of voice signals while performing voice recognition for the first voice signal acquired from the first microphone.

18. The step of bypassing speech recognition for the other speech signals comprises: performing speech recognition on the first speech signal using a first data set corresponding to the first speech signal; 18. The method of claim 17, further comprising: temporarily suspending deletion of other data sets corresponding to the other audio signals based on a specified data size while bypassing speech recognition for the other audio signals.

19. The method comprises: identifying text data from the first data set corresponding to the specified invocation term; obtaining the second point in time by performing a rollback on the first data set and a rollback on the other data set based on the identified text data; 20. The method of claim 18, further comprising: performing speech recognition using the first data set and the other data set as a whole from the second time point.

20. The method comprises: identifying the first microphone of the plurality of microphones based on a distance to a user associated with the plurality of audio signals; 12. The method of claim 11, further comprising: acquiring the first audio signal using the first microphone.

Citation Information

Patent Citations

  • Position estimation device, robot system including the same, and position estimation method thereof

    JP2022128579A