Method for providing call translation service and electronic device thereof
By overlapping translation data voice with the speaker's voice based on frequency analysis, the method addresses delays in real-time translation, enhancing the naturalness and emotional conveyance in multilingual calls.
Patent Information
- Application Number
- PCT/KR2025/010894
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-11
- Filing Date
- 2025-07-23
- Publication Date
- 2026-02-05
AI Technical Summary
Existing real-time translation during calls on electronic devices often causes delays and interrupts the conversation flow, making it difficult to convey real-time reactions and emotions due to the separation of speaker's voice and translation data output.
The method involves overlapping the translation data voice with the speaker's voice by analyzing voice characteristics through frequency analysis to reduce delay and improve recognition, allowing simultaneous transmission of both signals.
This approach enhances the naturalness of conversations by reducing delays and enabling users to convey emotions more effectively during calls by overlapping translated voice with the original voice, thus improving the real-time translation experience.
Smart Images

Figure KR2025010894_05022026_PF_FP_ABST
Abstract
Description
Method for providing currency translation services and electronic device thereof
[0001] The present disclosure relates to a method for providing a currency translation service and an electronic device thereof.
[0002] With the recent increase in the prevalence of various electronic devices, various services that operate through networked devices are being developed. Recently, translation functions have been introduced to overcome language barriers in various situations, such as overseas travel, business, and study. For example, electronic devices can provide real-time call translation services to ensure seamless communication between users speaking different languages on the phone.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.
[0004] According to an embodiment of the present disclosure, an electronic device may include one or more microphones, a communication circuit, at least one processor, and a memory storing instructions. The instructions, when executed by the at least one processor, may cause the electronic device to obtain a first audio signal including voice data in a first language from the one or more microphones while a call is established with an external electronic device. The instructions, when executed by the at least one processor, may cause the electronic device to obtain a second audio signal in which the voice data in the first language is translated into a second language specified by a user. The instructions, when executed by the at least one processor, may cause the electronic device to identify a voice characteristic of the first audio signal based on a frequency analysis of the voice data. The instructions, when executed by the at least one processor, may cause the electronic device to determine setting information related to the output of the first audio signal or the second audio signal based on the voice characteristic. The instructions may be executed by the at least one processor to cause the electronic device to transmit the first audio signal and the second audio signal to the external electronic device using the communication circuit based on the determined setting information.
[0005] A method according to one embodiment of the present disclosure may include an operation of obtaining a first audio signal including voice data in a first language from one or more microphones of an electronic device while a call is established with an external electronic device. The method may include an operation of obtaining a second audio signal in which the voice data in the first language is translated into a second language specified by a user. The method may include an operation of identifying a voice characteristic of the first audio signal based on a frequency analysis of the voice data. The method may include an operation of determining configuration information related to the output of the first audio signal or the second audio signal based on the voice characteristic. The method may include an operation of transmitting the first audio signal and the second audio signal to the external electronic device based on the determined configuration information using a communication circuit of the electronic device.
[0006] A computer-readable recording medium according to one embodiment of the present disclosure may include programs executable on a computer. The programs may perform an operation of acquiring a first audio signal including voice data in a first language from one or more microphones of an electronic device while a call is established with an external electronic device. The programs may perform an operation of acquiring a second audio signal in which the voice data in the first language is translated into a second language specified by a user. The programs may perform an operation of identifying a voice characteristic of the first audio signal based on a frequency analysis of the voice data. The programs may perform an operation of determining setting information related to the output of the first audio signal or the second audio signal based on the voice characteristic. The programs may perform an operation of transmitting the first audio signal and the second audio signal to the external electronic device based on the determined setting information using a communication circuit of the electronic device.
[0007] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.
[0008] FIG. 2 is a drawing illustrating the configuration of an electronic device according to one embodiment.
[0009] FIGS. 3A and 3B are diagrams illustrating a method of transmitting a first audio signal and a second audio signal according to various embodiments.
[0010] FIG. 4 is a diagram illustrating a method for transmitting a first audio signal and a second audio signal from an electronic device to an external electronic device, according to one embodiment.
[0011] FIG. 5 is a diagram illustrating a method for providing a translation function while a call is established between a first electronic device and a second electronic device.
[0012] FIG. 6 is a flowchart illustrating a method of operating an electronic device according to one embodiment.
[0013] FIG. 7 is a flowchart illustrating a method for determining a voice of a second audio signal based on a voice characteristic of a first audio signal in an electronic device, according to one embodiment.
[0014] FIG. 8A is a flowchart illustrating a method for determining a voice of a second audio signal based on frequency component similarity of a first audio signal and a second audio signal in an electronic device, according to one embodiment.
[0015] FIG. 8b is a diagram illustrating a method for determining frequency component similarity between a first audio signal and a second audio signal, according to one embodiment.
[0016] FIG. 9 is a flowchart illustrating a method of applying acoustic distortion to a first audio signal in an electronic device, according to one embodiment.
[0017] FIG. 10 is a flowchart illustrating a method for adjusting the volume of a first audio signal in an electronic device, according to one embodiment.
[0018] FIG. 11 is a diagram illustrating a user interface provided while using a translation function during a call, according to one embodiment.
[0019] FIG. 12 is a flowchart illustrating a method for transmitting a first audio signal and a second audio signal using two channels in an electronic device, according to one embodiment.
[0020] FIG. 13 is a flowchart illustrating a method for receiving a first audio signal and a second audio signal transmitted using two channels in an electronic device, according to one embodiment.
[0021] In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0022] Hereinafter, various embodiments disclosed in this document will be described with reference to the attached drawings. It should be understood that this is not intended to limit the various embodiments of the present disclosure to a specific form, but rather to encompass various modifications, equivalents, and / or alternatives of the present disclosure.
[0023] When using real-time translation during a call on an electronic device, the translation data (e.g., translated text-to-speech (TTS)) may be delivered to the listener after the speaker has finished speaking. For example, the listener may receive the speaker's voice, receive the translation data voice after a certain period of time, and then speak the desired voice to the other party. This can cause significant delays during the call, frequently interrupting the flow of conversation, and making it difficult to convey real-time reactions, which can be uncomfortable for both parties. To reduce the inconvenience caused by delays when providing translation during a call on an electronic device, only the translation data voice can be transmitted without the speaker's voice output. However, this has limitations in conveying the speaker's voice and the emotions and surrounding sounds contained within it.
[0024] In various embodiments of this document, when using a real-time translation function during a call, various embodiments can be provided to improve the recognition rate while reducing the delay time for the voice signal of the translation data in the process of transmitting the voice signal of the speaker and the voice signal of the translation data by overlapping at least some of them.
[0025] The technical tasks to be achieved in this document are not limited to the technical tasks mentioned above, and other technical tasks not mentioned will be clearly understood by those with ordinary knowledge in the technical field to which this document pertains.
[0026] FIG. 1 is a diagram illustrating an electronic device within a network environment (100) according to one embodiment.
[0027] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0028] The processor (120) may control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing, for example, software (e.g., a program (140)), and may perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculation, the processor (120) may store a command or data received from another component (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the command or data stored in the volatile memory (132), and store the resulting data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0029] The auxiliary processor (123) may control at least a part of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0030] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0031] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0032] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0033] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. According to one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0034] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0035] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0036] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0037] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0038] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0039] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0040] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0041] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0042] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0043] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, Wi-Fi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0044] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0045] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. According to some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0046] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., a bottom surface) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent to a second surface (e.g., a top surface or a side surface) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0047] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0048] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In one embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0049] FIG. 2 is a drawing illustrating the configuration of an electronic device (200) according to one embodiment.
[0050] Referring to FIG. 2, the electronic device (200) is a device that transmits a voice signal of a speaker during a call and a voice signal of translation data to an external electronic device set for a call, and may include one or more microphones (210), a communication circuit (220), a display (230), at least one processor (240), and a memory (250). In FIG. 2, the electronic device (200) may correspond to the electronic device (101) illustrated in FIG. 1.
[0051] In one embodiment, one or more microphones (210) (e.g., input module (150) of FIG. 1) can acquire voice data spoken by a user.
[0052] In one embodiment, the communication circuit (220) (e.g., the communication module (190) of FIG. 1) may support a communication connection with an external electronic device. For example, the communication circuit (220) may establish a call connection with the external electronic device. According to various embodiments, the communication circuit (220) may establish a communication connection with an electronic device located within a predetermined distance using short-range communication such as Bluetooth, Wi-Fi direct, or BLE (Bluetooth low energy), or may establish a communication connection with an external electronic device located outside a predetermined distance using long-range communication such as cellular or Wi-Fi to exchange data.
[0053] In one embodiment, the display (230) (e.g., the display module (160) of FIG. 1) may display information desired to be provided to the user. For example, if a translation function is set during a call, the display (230) may display text corresponding to voice data spoken by the user and / or a translated text of the text.
[0054] In one embodiment, the display (230) may be configured with at least one of a liquid crystal display (LCD), a thin film transistor LCD (TFT-LCD), an organic light emitting diode (OLED), a light emitting diode (LED), an active matrix organic LED (AMOLED), a flexible display, and a 3-dimensional display. In addition, some of these displays may be configured as transparent or light-transmitting so that the outside can be viewed through them. This may be configured in the form of a transparent display including a TOLED (transparent OLED).
[0055] In addition, the components described with reference to FIG. 1 may be appropriately applied to the electronic device (200). For example, the electronic device (200) may further include a speaker (e.g., the audio output module (155) of FIG. 1) for outputting the other party's voice data and translated voice data transmitted from an external electronic device (not shown in FIG. 2).
[0056] In one embodiment, the memory (250) (e.g., the memory (130) of FIG. 1) may store instructions that, when executed, cause the electronic device (200) to perform various operations by at least one processor (240) (e.g., the processor (120) of FIG. 1). For example, the at least one processor (240) may control operations to be performed to increase the recognition rate while reducing the delay time for a translated voice signal when providing a translation function during a call.
[0057] In one embodiment, at least one processor (240) may obtain a first audio signal comprising voice data of a first language from one or more microphones (210) while a call connection is established with an external electronic device via the communication circuit (220).
[0058] In one embodiment, at least one processor (240) may perform speech recognition on speech data in the first language. For example, at least one processor (240) may identify a first text corresponding to the speech data in the first language as a result of the speech recognition. At least one processor (240) may obtain a second text in which the first text is translated into a second language. The second language is a language different from the first language and may correspond to a language used by a user of the external electronic device. According to various embodiments, when the second language is specified by the user, at least one processor (240) may obtain the second text by translating the first text into the second language based on a generative artificial intelligence (AI) model or obtain the second text from an external server that provides a real-time translation function.
[0059] In one embodiment, at least one processor (240) may convert the second text into speech data to obtain a second audio signal. The second audio signal may be delivered to the user using the speech of the translation data.
[0060] In one embodiment, at least one processor (240) may identify a voice characteristic of the first audio signal based on frequency analysis of voice data included in the first audio signal, and may determine setting information related to the output of the first audio signal or the second audio signal based on the identified voice characteristic of the first audio signal. According to various embodiments, at least one processor (240) may determine a first voice type corresponding to the first audio signal based on the voice characteristic of the first audio signal, and may set a voice of a type different from the first voice type as the voice of the second audio signal. For example, if at least one processor (240) determines that the voice of the first audio signal is classified as a female voice based on the frequency analysis, it may set the voice of the second audio signal as a male voice. As another example, if at least one processor (240) determines that the voice of the first audio signal is classified as a voice of an age group of 60 or older based on the frequency analysis, it may set the voice of the second audio signal as a voice of an age group of 30.
[0061] In one embodiment, at least one processor (240) may verify the frequency components of voice data included in the first audio signal, and set a voice in a frequency band that does not overlap with the verified frequency components as the voice of the second audio signal. For example, if the at least one processor (240) analyzes the frequency components of voice data included in the first audio signal and verifies that sound energy is concentrated in a band of 1k to 2kHz, the at least one processor (240) may determine a voice in which energy is concentrated in a frequency band that does not overlap with 1k to 2kHz as the voice of the second audio signal. As another example, the at least one processor (240) may select, from among a plurality of voice types that can be set as the voice of the second audio signal, a voice type that has the lowest similarity with the frequency components verified for the first audio signal as the voice of the second audio signal. In this case, the at least one processor (240) may verify a first frequency spectrum for voice data of the first audio signal and a plurality of second frequency spectra corresponding to each of the plurality of voice types. At least one processor (240) may compare the first frequency spectrum with the plurality of second frequency spectra, and select a voice type corresponding to a frequency spectrum having the lowest similarity with the first frequency spectrum among the plurality of second frequency spectra as the voice of the second audio signal. As another example, at least one processor (240) may generate a voice in which a ratio of a frequency band where energy is concentrated in the first audio signal is lower than a specified ratio, and set the generated voice as the voice of the second audio signal.
[0062] In one embodiment, at least one processor (240) may process the first audio signal to have an artificial sense of space through acoustic distortion of the first audio signal. For example, at least one processor (240) may adjust the gain of a designated frequency band (e.g., 6 to 16 kHz) of the first audio signal to be lowered based on multi-band dynamic range control (MBDRC). In this case, the first audio signal may be perceived as being transmitted from a distance away from the user and / or from behind the user.
[0063] In one embodiment, at least one processor (240) may set the volume level of the first audio signal to a specified level or lower. For example, at least one processor (240) may adjust the volume level of the first audio signal to the specified level (e.g., about 9 dB level) and maintain the volume level of the second audio signal or adjust it to output at a volume level greater than that of the first audio signal, thereby enabling a user to better recognize the second audio signal.
[0064] In one embodiment, at least one processor (240) may transmit the first audio signal and the second audio signal to the external electronic device using the communication circuit (220) based on configuration information determined for the first audio signal and / or the second audio signal. For example, at least one processor (240) may transmit the first audio signal and the second audio signal to the external electronic device through one channel. In this case, at least one processor (240) may mix the first audio signal and the second audio signal based on the configuration information and transmit the mixed signal to the external electronic device through the first channel. As another example, at least one processor (240) may transmit the first audio signal and the second audio signal to the external electronic device through two channels. At least one processor (240) may transmit the first audio signal to the external electronic device through the first channel and the second audio signal to the external electronic device through the second channel, which is distinct from the first channel, based on the configuration information. In this case, the first audio signal and the second audio signal may be output through one channel or two channels depending on whether the external electronic device supports stereo output. According to various embodiments, if the external electronic device does not support stereo output, the first audio signal and the second audio signal may be mixed into one signal and output by the external electronic device. If the external electronic device supports stereo output, the first audio signal and the second audio signal may be output through two channels, respectively.For example, the first audio signal may be output through the left speaker and the second audio signal may be output through the right speaker, or only one audio signal selected by the external electronic device among the two audio signals may be output.
[0065] According to various embodiments, at least one processor (240) may display texts corresponding to audio signals transmitted and received between the electronic device (200) and the external electronic device through the display (230) while a call with the external electronic device is established. For example, at least one processor (240) may provide a user interface that displays a first text in a first language identified through voice recognition for the first audio signal while the call is established, and a second text translated into a second language of the first text. As another example, at least one processor (240) may additionally display a third text in addition to the first text and the second text through the user interface so that the user can check the accuracy of the translation transmitted to the external electronic device. The third text may include at least one of a text translated from the second text in the second language back into the first language, or a text translated from the first text in the first language or the second text in the second language into a third language.
[0066] FIGS. 3A and 3B are diagrams illustrating a method of transmitting a first audio signal and a second audio signal according to various embodiments.
[0067] Referring to FIG. 3A, the electronic device (200) may transmit a first audio signal (310) including voice data spoken by a user in a first language to an external electronic device (350), and transmit a second audio signal (315) including voice data translated into a second language to the external electronic device (350) after a specified time (e.g., 1 to 2 seconds). Thereafter, the electronic device (200) may receive a third audio signal (320) spoken by a user of the external electronic device in a second language from the external electronic device (350) in response to the first audio signal (310), and receive a fourth audio signal (325) including voice data translated into the first language from the external electronic device after a specified time (e.g., 1 to 2 seconds). In this case, since the translated audio signal is transmitted to the other device after the user's speech is completed and the above-mentioned time has elapsed, the transmission of the translated audio signal is delayed, which may cause the flow of conversation between users during the call to be unnatural and make it difficult to transmit real-time reactions.
[0068] Referring to FIG. 3B, the electronic device (200) may transmit a first audio signal (330) including voice data uttered by a user in a first language to an external electronic device (350), while transmitting a second audio signal (335) including voice data translated into a second language to the external electronic device (350). For example, the electronic device (200) may transmit the second audio signal (335) to the external electronic device (350) by at least partially overlapping the first audio signal (330) even when the user's speech is not completed. In this case, the electronic device (200) may transmit a translated sentence to the external electronic device (350) in succession at each sentence conclusion during the user's speech. Thereafter, the electronic device (200) may receive a third audio signal (340) spoken in a second language by a user of the external electronic device (350) in response to the first audio signal from the external electronic device (350), and may receive a fourth audio signal (345) from the external electronic device (350) that includes voice data translated into the first language and at least partially overlapping with the third signal (340). In this case, since the translated audio signal is transmitted to the other device together with the user's speech, a more natural conversation can proceed between users who speak different languages, and it may be easier for the users to convey their emotions during a call.
[0069] FIG. 4 is a diagram illustrating a method of transmitting a first audio signal and a second audio signal from an electronic device (200) to an external electronic device (450), according to one embodiment.
[0070] Referring to FIG. 4, the electronic device (200) can obtain a first audio signal including voice data spoken by a user in a first language through a microphone (410), perform processing on the first audio signal using a Tx call processing module (420), and then transmit the first audio signal to an external electronic device (450).
[0071] In one embodiment, the Tx call processing module (420) may perform operations of echo canceling / noise suppression (EC / NS) (421), multi-band dynamic range control (MBDRC) (422), and voice processing (423). EC / NS (421) may be an operation of removing echo components and noise components included in the first audio signal, and MBDRC (422) may be an operation of applying an artificial spatial effect by adjusting the gain so that a specified frequency band (e.g., 6 to 16 kHz) in the first audio signal is attenuated. Voice processing (423) may be an operation of applying various filters (e.g., high pass filter (HPF), finite impulse response (FIR) filter) to the first audio signal and performing tuning such as gain adjustment or volume adjustment by considering distance (or space) in order to effectively transmit the user's spoken voice to an external electronic device (450).
[0072] In one embodiment, the electronic device (200) may obtain a second audio signal by translating voice data spoken by the user in a first language into a second language using the translation data processing module (430), and transmit the obtained second audio signal together with the first audio signal to the external electronic device (450). According to various embodiments, the second audio signal may be transmitted with at least a portion of the first audio signal overlapping the first audio signal while the first audio signal is transmitted to the external electronic device (450).
[0073] In one embodiment, the translation data processing module (430) may perform operations of speech recognition (431), speech to text (STT) (432), translation (433), text to speech (TTS) (434), frequency analysis (435), and translation data voice determination (436). Speech recognition (431) may be an operation of performing speech recognition on speech data of a first language spoken by a user, and STT (432) may be an operation of obtaining a first text corresponding to the speech data based on the speech recognition on the speech data. Translation (433) may be an operation of obtaining a second text translated from the first text into a second language, and TTS (434) may be an operation of converting the second text into speech data to obtain a second audio signal. According to various embodiments, the translation data processing module (430) may perform a frequency analysis (435) on the voice-recognized voice data, and determine (436) a voice of the translation data to be applied to the second audio signal based on the result of the frequency analysis. For example, the translation data processing module (430) may classify a voice characteristic of the voice data uttered by the user based on the frequency analysis (435) on the voice data, identify a first voice type corresponding to the voice data, and determine to apply a translation data voice of a type different from the identified first voice type to the second audio signal. For example, the translation data processing module (430) may estimate at least one of a gender or an age group of the voice data uttered by the user based on the frequency analysis (435) to identify a type of the voice data, and determine a translation data voice classified as a type different from the identified type as the voice of the second audio signal.For another example, the translation data processing module (430) may determine to apply the translation data voice having a low similarity to the identified frequency component based on the frequency analysis (435) of the voice data spoken by the user to the second audio signal. In this case, the translation data processing module (430) may determine a frequency band where energy is concentrated in the voice data spoken by the user based on the frequency analysis (435), and determine the translation data voice having energy concentrated in a frequency band that does not overlap with the identified frequency band as the voice of the second audio signal. According to various embodiments, the translation data processing module (430) may also generate a voice having a ratio of a frequency band where energy is concentrated in the voice data spoken by the user lower than a specified ratio and apply the generated voice to the second audio signal.
[0074] In one embodiment, the electronic device (200) may perform processing on the first audio signal to increase the recognition rate of the second audio signal during the process of transmitting the second audio signal to the external electronic device (450) by at least partially overlapping the first audio signal. For example, the electronic device (200) may perform an operation of MBDRC (422) on the first audio signal using the Tx call processing module (420) to process the first audio signal so that the first audio signal has an artificial sense of space. As another example, the electronic device (200) may adjust the volume level of the first audio signal to a specified level or lower while performing an operation of voice processing (523) using the Tx call processing module (420) so that the second audio signal is more emphasized.
[0075] In one embodiment, the electronic device (200) can transmit a first audio signal processed by the Tx call processing module (420) and a second audio signal processed by the translation data processing module (430) to an external electronic device (450). When using one channel, the electronic device (200) can mix the first audio signal and the second audio signal and transmit the mixed signal to the external electronic device (450). When using two channels, the electronic device (200) can transmit the first audio signal and the second audio signal to the external electronic device (450) through each channel without mixing the audio signals. When the external electronic device (450) receives the first audio signal and the second audio signal from the electronic device (200), the external electronic device (450) can adjust at least one of volume, sound quality, and tone through the Rx call processing module and then output the first and second audio signals.
[0076] FIG. 5 is a diagram illustrating a method for providing a translation function while a call is established between a first electronic device (510) and a second electronic device (520), according to one embodiment. In FIG. 5 , a first user of the first electronic device (510) and a second user of the second electronic device (520) may be understood as users who speak different languages.
[0077] Referring to FIG. 5, a first electronic device (510) can obtain a first audio signal (502) including first voice data (501) spoken in a first language by a first user through a microphone (511). The first electronic device (510) can perform various processing on the first audio signal (502) through a Tx call processing module (512). For example, the Tx call processing module (512) can apply an artificial sense of space to the first audio signal (502) through gain suppression for a specified frequency band (e.g., 6 to 16 kHz) of the first audio signal (502). As another example, the Tx call processing module (512) can adjust the volume level of the first audio signal (502) to a specified level or lower.
[0078] In one embodiment, the first electronic device (510) may obtain a second audio signal (503) by translating first voice data (501) of a first language into a second language using the translation data processing module (530). The translation data processing module (530) may identify a second text that is a translation of a first text corresponding to the first voice data (501) into a second language, and may convert the second text into voice data to obtain a second audio signal (503). The second audio signal (503) may be output using the translation data voice determined by the translation data processing module (530). According to various embodiments, the translation data processing module (530) may determine the translation data voice to output the second audio signal (503) based on frequency analysis of the first audio signal (502). For example, the translation data processing module (530) can identify voice characteristics based on frequency analysis of the first audio signal (502), and determine a translation data voice of a different type from the voice type corresponding to the identified voice characteristics as the voice of the second audio signal (503). The voice type can be classified according to gender and / or age group. As another example, the translation data processing module (530) can identify a frequency band in which energy is concentrated in the first audio signal (502), and select a voice type in which energy is concentrated in a frequency band that does not overlap with the identified frequency band as the voice of the second audio signal (503).
[0079] In one embodiment, a first electronic device (510) may transmit a first audio signal (502) and a second audio signal (503) to a second electronic device (520). At this time, while transmitting the first audio signal (502), the first electronic device (510) may transmit the second audio signal (503) to the second electronic device (520) by at least partially overlapping the first audio signal (502). According to various embodiments, the first electronic device (510) may transmit the first audio signal (502) and the second audio signal (503) to the second electronic device (520) using one or two channels. When using one channel, the first electronic device (510) can mix the first audio signal (502) and the second audio signal (503) into one signal and transmit the mixed signal to the second electronic device (520) through one channel. When using two channels, the first electronic device (510) can transmit the first audio signal (502) and the second audio signal (503) to the second electronic device (520) through two different channels, respectively.
[0080] In one embodiment, the first audio signal (502) and the second audio signal (503) transmitted from the first electronic device (510) may pass through the Rx call processing module (522) of the second electronic device (520) and be output through the speaker (521). The Rx call processing module (522) may apply gain control and various filters (e.g., a high pass filter (HPF), a finite impulse response (FIR) filter) to the first audio signal (502) and the second audio signal (503) so that the first audio signal (502) and the second audio signal (503) received from the first electronic device (510) may be efficiently transmitted to the user of the second electronic device (520), and then may transmit the first audio signal (502) and the second audio signal (503) to the speaker (521).
[0081] In one embodiment, the second electronic device (520) may acquire a third audio signal (506) including second voice data (505) spoken in a second language by a second user through a microphone (523). The second electronic device (520) may perform various processing on the third audio signal (506) through a Tx call processing module (524). For example, the Tx call processing module (524) may apply an artificial sense of space to the third audio signal (506) through gain suppression for a specified frequency band (e.g., 6 to 16 kHz) of the third audio signal (506). As another example, the Tx call processing module (524) may adjust the volume level of the third audio signal (506) to a specified level or lower.
[0082] In one embodiment, the second electronic device (520) may obtain a fourth audio signal (507) by translating second voice data (506) of a second language into a first language using the translation data processing module (530). The translation data processing module (530) may identify a fourth text that is a third text corresponding to the third voice data (506) translated into the first language, and may convert the fourth text into voice data to obtain a fourth audio signal (507). The fourth audio signal (507) may be output using the translation data voice determined by the translation data processing module (530). According to various embodiments, the translation data processing module (530) may determine the translation data voice to output the fourth audio signal (507) based on frequency analysis of the third audio signal (506). For example, the translation data processing module (530) can identify voice characteristics based on frequency analysis of the third audio signal (506) and determine a translation data voice of a different type from the voice type corresponding to the identified voice characteristics as the voice of the fourth audio signal (507). The voice type can be classified according to gender and / or age group. As another example, the translation data processing module (530) can identify a frequency band in which energy is concentrated in the third audio signal (506) and select a translation data voice in which energy is concentrated in a frequency band that does not overlap with the identified frequency band as the voice of the fourth audio signal (507).
[0083] In one embodiment, the second electronic device (520) may transmit a third audio signal (506) and a fourth audio signal (507) to the first electronic device (510). At this time, the second electronic device (520) may transmit a fourth audio signal (507) to the first electronic device (510) while transmitting the third audio signal (506) by at least partially overlapping the third audio signal (506). According to various embodiments, the second electronic device (520) may transmit the third audio signal (506) and the fourth audio signal (507) to the first electronic device (510) using one or two channels. When using one channel, the second electronic device (520) can mix the third audio signal (506) and the fourth audio signal (507) into one signal and transmit the mixed signal to the first electronic device (510) through one channel. When using two channels, the second electronic device (520) can transmit the third audio signal (506) and the fourth audio signal (507) to the first electronic device (510) through two different channels, respectively.
[0084] In one embodiment, the third audio signal (506) and the fourth audio signal (507) transmitted from the second electronic device (520) may pass through the Rx call processing module (514) of the first electronic device (510) and be output through the speaker (513). The Rx call processing module (514) may apply gain control and various filters (e.g., a high pass filter (HPF), a finite impulse response (FIR) filter) to the third audio signal (506) and the fourth audio signal (507) so that the third audio signal (506) and the fourth audio signal (507) received from the second electronic device (520) may be efficiently transmitted to the user of the first electronic device (510), and then may transmit the same to the speaker (513).
[0085] FIG. 6 is a flowchart illustrating an operating method of an electronic device (200) according to an embodiment. According to an embodiment, the electronic device (200) is a device that transmits a voice signal of a speaker during a call and a voice signal of translation data to an external electronic device set for a call, and may correspond to the electronic device (101) illustrated in FIG. 1. The operations of FIG. 6 may be performed by at least one processor included in the electronic device (200) (e.g., the processor (120) of FIG. 1 or the processor (240) of FIG. 2).
[0086] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0087] Referring to FIG. 6, in operation 610, the electronic device (200) may obtain a first audio signal including voice data of a first language from one or more microphones (e.g., the input module (150) of FIG. 1 or the microphone (210) of FIG. 2) while a call connection is established with an external electronic device.
[0088] According to one embodiment, in operation 620, the electronic device (200) may obtain a second audio signal in which the voice data of the first language is translated into a second language specified by the user. The electronic device (200) may identify a first text corresponding to the voice data of the first language based on voice recognition of the voice data of the first language, and may obtain a second text in which the first text is translated into a second language. The second language may be a language different from the first language and may correspond to a language used by the user of the external electronic device. The electronic device (200) may convert the second text into voice data to obtain the second audio signal. The second audio signal may be transmitted to the user of the electronic device (200) and / or the external electronic device using the translated data voice.
[0089] According to one embodiment, in operation 630, the electronic device (200) can identify the voice characteristics of the first audio signal based on frequency analysis of the voice data included in the first audio signal. For example, the electronic device (200) can identify a voice type corresponding to the first audio signal based on the frequency analysis. The voice type can be classified based on gender and / or age group. As another example, the electronic device (200) can identify the frequency components of the voice data included in the first audio signal based on the frequency analysis. The electronic device (200) can identify a frequency band where sound energy is concentrated in the first audio signal through frequency analysis of the first audio signal.
[0090] According to one embodiment, in operation 640, the electronic device (200) may determine setting information related to the output of the first audio signal or the second audio signal based on the voice characteristics of the identified first audio signal. According to various embodiments, in operation 640, the electronic device (200) may set a voice of a different type from the voice type identified for the first audio signal as the voice of the second audio signal. For example, if the electronic device (200) determines that the voice of the first audio signal is classified as a female voice based on the frequency analysis, the electronic device (200) may set the voice of the second audio signal as a male voice. As another example, if the electronic device (200) determines that the voice of the first audio signal is classified as a voice of an age group of 60 or older based on the frequency analysis, the electronic device (200) may set the voice of the second audio signal as a voice of an age group of 30.
[0091] According to various embodiments, the electronic device (200) may, in operation 640, set a voice in a frequency band that does not overlap with a frequency component identified for the first audio signal as the voice of the second audio signal. For example, if the electronic device (200) determines that sound energy is concentrated in a 1k to 2kHz band based on a frequency analysis of the first audio signal, the electronic device (200) may determine a voice in which energy is concentrated in a frequency band that does not overlap with 1k to 2kHz as the voice of the second audio signal. As another example, the electronic device (200) may select a voice type having the lowest similarity with a frequency component identified for the first audio signal among a plurality of voice types that can be set as the voice of the second audio signal as the voice of the second audio signal. In this case, the electronic device (200) may check a first frequency spectrum for voice data of the first audio signal and a plurality of second frequency spectra corresponding to each of the plurality of voice types. At least one processor (240) may compare the first frequency spectrum with the plurality of second frequency spectra, and select a voice type corresponding to a frequency spectrum having the lowest similarity with the first frequency spectrum among the plurality of second frequency spectra as the voice of the second audio signal. As another example, the electronic device (200) may generate a voice in which a ratio of a frequency band where energy is concentrated in the first audio signal is lower than a specified ratio, and set the generated voice as the voice of the second audio signal.
[0092] According to various embodiments, the electronic device (200) may process the first audio signal to have an artificial sense of space through acoustic distortion of the first audio signal in operation 640. For example, the electronic device (200) may adjust the gain of a designated frequency band (e.g., 6 to 16 kHz) of the first audio signal to be lowered based on multi-band dynamic range control (MBDRC). In this case, the first audio signal may be perceived as being transmitted from a distance away from the user and / or from behind the user.
[0093] According to various embodiments, the electronic device (200) may, in operation 640, set the volume level of the first audio signal to a specified level or lower. For example, the electronic device (200) may adjust the volume level of the first audio signal to the specified level (e.g., about 9 dB level) and maintain the volume level of the second audio signal or adjust it to output at a volume level greater than that of the first audio signal, thereby enabling the user to better recognize the second audio signal.
[0094] According to one embodiment, in operation 650, the electronic device (200) may transmit the first audio signal and the second audio signal to the external electronic device using a communication circuit (e.g., the communication module (190) of FIG. 1 or the communication circuit (220) of FIG. 2) based on the configuration information determined for the first audio signal and / or the second audio signal. For example, in operation 650, the electronic device (200) may transmit the first audio signal and the second audio signal to the external electronic device through one channel. In this case, the electronic device (200) may mix the first audio signal and the second audio signal based on the configuration information, and transmit the mixed signal to the external electronic device through the first channel. As another example, in operation 650, the electronic device (200) may transmit the first audio signal and the second audio signal to the external electronic device through two channels. The electronic device (200) may transmit the first audio signal to the external electronic device through a first channel based on the setting information, and the second audio signal to the external electronic device through a second channel that is distinct from the first channel. In this case, the first audio signal and the second audio signal may be output through one channel or two channels depending on whether the external electronic device supports stereo output. According to various embodiments, when the external electronic device does not support stereo output, the first audio signal and the second audio signal may be mixed into one signal by the external electronic device and output. When the external electronic device supports stereo output, the first audio signal and the second audio signal may be output through two channels, respectively.For example, the first audio signal may be output through the left speaker and the second audio signal may be output through the right speaker, or only one audio signal selected by the external electronic device among the two audio signals may be output.
[0095] According to various embodiments, the electronic device (200) may display texts corresponding to audio signals transmitted and received between the electronic device (200) and the external electronic device while a call is established with the external electronic device through a display (e.g., the display module (160) of FIG. 1 or the display (230) of FIG. 2). For example, the electronic device (200) may provide a user interface that displays a first text in a first language identified through voice recognition for the first audio signal while the call is established, and a second text translated into a second language of the first text. As another example, the electronic device (200) may additionally display a third text in addition to the first text and the second text through the user interface so that the user can check the accuracy of the translation transmitted to the external electronic device. The third text may be a text that is a translation of the second text in the second language back into the first language, or may include at least one of a second translation of the first text in the first language or the second text in the second language into the third language.
[0096] FIG. 7 is a flowchart illustrating a method for determining a voice of a second audio signal based on a voice characteristic of a first audio signal in an electronic device (200), according to an embodiment. According to various embodiments, the electronic device (200) may obtain a first audio signal including voice data of a first language spoken by a user while a call connection is established with an external electronic device through a microphone (e.g., the input module (150) of FIG. 1 or the microphone (210) of FIG. 2). The electronic device (200) may obtain a second audio signal obtained by translating the voice data of the first language into a second language, and transmit the obtained second audio signal together with the first audio signal to the external electronic device. The operations of FIG. 7 may be understood as functions performed by at least one processor included in the electronic device (200) (e.g., the processor (120) of FIG. 1 or the processor (240) of FIG. 2).
[0097] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0098] Referring to FIG. 7, in operation 710, the electronic device (200) can identify the voice characteristics of the first audio signal based on frequency analysis. When the electronic device (200) obtains a first audio signal including voice data spoken by a user in a first language through a microphone (210), the electronic device (200) can perform frequency analysis on the first audio signal to identify the voice characteristics of the first audio signal.
[0099] According to one embodiment, in operation 720, the electronic device (200) may determine the voice type of the first audio signal based on the voice characteristics of the first audio signal identified above. For example, in operation 720, the electronic device (200) may identify the voice of the first audio signal as a female voice based on the voice characteristics identified by the frequency analysis. As another example, in operation 720, the electronic device (200) may estimate the voice of the first audio signal as a voice of an age group of 60 or older based on the voice characteristics identified by the frequency analysis.
[0100] According to one embodiment, in operation 730, the electronic device (200) may determine a voice different from the voice type of the first audio signal as the voice of the second audio signal. For example, if the electronic device (200) identifies the voice type of the first audio signal as a female voice, the electronic device (200) may determine the voice to be applied to the second audio signal as a male voice. As another example, if the electronic device (200) identifies the voice type of the first audio signal as a voice of an age group of 60 or older, the electronic device (200) may determine the voice to be applied to the second audio signal as a voice of an age group of 20.
[0101] According to one embodiment, the electronic device (200) may perform mixing of the first audio signal and the second audio signal in operation 740, and transmit the mixed audio signal to the external electronic device in operation 750. The electronic device (200) may perform mixing of the first audio signal and the second audio signal so that the second audio signal may be output by at least partially overlapping with the first audio signal, and transmit the mixed audio signal to the external electronic device through one channel.
[0102] According to various embodiments, when two channels are used, the electronic device (200) may not perform operation 740. In this case, the electronic device (200) may transmit the first audio signal and the second audio signal to the external electronic device through two different channels, respectively. For example, in operation 750, the electronic device (200) may transmit the second audio signal to the external electronic device through a second channel while transmitting the first audio signal to the external electronic device through a first channel.
[0103] FIG. 8A is a flowchart illustrating a method for determining a voice of a second audio signal based on frequency component similarity of a first audio signal and a second audio signal in an electronic device (200) according to an embodiment, and FIG. 8B is a diagram explaining a method for determining frequency component similarity of a first audio signal and a second audio signal according to an embodiment. According to various embodiments, the electronic device (200) may obtain a first audio signal including voice data of a first language spoken by a user while a call connection is established with an external electronic device through a microphone (e.g., the input module (150) of FIG. 1 or the microphone (210) of FIG. 2). The electronic device (200) may obtain a second audio signal obtained by translating the voice data of the first language into a second language, and transmit the obtained second audio signal together with the first audio signal to the external electronic device. The operations of FIG. 8a may be understood as functions performed by at least one processor (e.g., processor (120) of FIG. 1 or processor (240) of FIG. 2) included in the electronic device (200).
[0104] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0105] Referring to FIG. 8A, in operation 810, the electronic device (200) can identify the frequency components of voice data included in the first audio signal based on frequency analysis. When the electronic device (200) obtains the first audio signal including voice data spoken by the user in a first language through the microphone (210), the electronic device (200) can perform frequency analysis on the first audio signal to identify the frequency components of the first audio signal. For example, in operation 810, the electronic device (200) can identify a frequency band where sound energy is concentrated in the first audio signal based on the frequency analysis.
[0106] According to one embodiment, in operation 820, the electronic device (200) can determine whether there is a translation data voice among a plurality of translation data voices that can be set for the second audio signal, the similarity between the analyzed frequency components of the first audio signal and the translation data voice is less than a threshold value. The determination of the similarity with respect to the frequency components of the first audio signal is described with reference to FIG. 8B.
[0107] Referring to FIG. 8B, the electronic device (200) can confirm that sound energy is concentrated in a 1k to 2kHz band (801) based on the frequency spectrum for the first audio signal (800). The electronic device (200) can also confirm a frequency band in which sound energy is concentrated in each of the plurality of translation data voices (870, 880, 890) based on frequency spectrum analysis. For example, the electronic device (200) can confirm that energy is concentrated in a 10k to 1kHz band (871) for the first translation data voice (870). The electronic device (200) can confirm that energy is concentrated in a 1.5k to 4kHz band (881) for the second translation data voice (880). The electronic device (200) can confirm that energy is concentrated in a 4k to 6kHz band (891) for the third translation data voice (890). The electronic device (200) can select, among the plurality of translation data voices (870, 880, 890), a third translation data voice (890) in which energy is concentrated in a frequency band (891) of 4k to 6kHz that does not overlap with a frequency band (801) of 1k to 2kHz identified from the first audio signal (800), as the translation data voice with the lowest similarity to the first audio signal (800).
[0108] Referring back to FIG. 8A, if, as a result of the determination in operation 820, there exists a translation data voice whose similarity is less than the threshold (operation 820-Yes), the electronic device (200) may select the translation data voice whose similarity to the frequency component of the first audio signal is the lowest in operation 830. For example, the electronic device (200) may select, from among the plurality of translation data voices, a translation data voice whose energy is concentrated in a frequency band that does not overlap with the frequency band where the energy of the first audio signal is concentrated.
[0109] As a result of the determination in operation 820, if there is no translation data voice with the similarity less than the threshold value (operation 820-No), the electronic device (200) may select the translation data voice with the lowest similarity to the frequency component of the first audio signal from among the plurality of translation data voices in operation 840. In this case, the electronic device (200) may select the translation data voice with the least overlapping frequency band even if it overlaps with the frequency band where the energy of the first audio signal is concentrated from among the plurality of translation data voices. The electronic device (200) may apply a filter that reduces a frequency band specified in the selected translation data voice in operation 845. For example, the electronic device (200) may apply a filter for reducing a frequency band with the largest distribution in the first audio signal to the selected translation data voice. The electronic device (200) may generate a new translation data voice based on the translation data voice to which the filter is applied in operation 847, and select the generated new translation data voice.
[0110] According to one embodiment, in operation 850, the electronic device (200) may determine the selected translation data voice as the voice of the second audio signal. The second audio signal may be output using the determined translation data voice.
[0111] According to one embodiment, in operation 860, the electronic device (200) may transmit the first audio signal and the second audio signal to the external electronic device. In operation 860, the electronic device (200) may transmit the first audio signal and the second audio signal to the external electronic device using one or two channels. For example, when using one channel, the electronic device (200) may perform mixing of the first audio signal and the second audio signal so that the second audio signal is output by at least partially overlapping with the first audio signal, and may transmit the mixed audio signal to the external electronic device through one channel. As another example, when using two channels, the electronic device (200) may transmit the second audio signal to the external electronic device through a second channel that is distinct from the first channel while transmitting the first audio signal to the external electronic device through the first channel.
[0112] FIG. 9 is a flowchart illustrating a method of applying acoustic distortion to a first audio signal in an electronic device (200), according to an embodiment. According to various embodiments, the electronic device (200) may obtain a first audio signal including voice data of a first language spoken by a user while a call connection is established with an external electronic device through a microphone (e.g., the input module (150) of FIG. 1 or the microphone (210) of FIG. 2). The electronic device (200) may obtain a second audio signal obtained by translating the voice data of the first language into a second language, and transmit the obtained second audio signal together with the first audio signal to the external electronic device. The operations of FIG. 9 may be understood as functions performed by at least one processor included in the electronic device (200) (e.g., the processor (120) of FIG. 1 or the processor (240) of FIG. 2).
[0113] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0114] Referring to FIG. 9, in operation 910, the electronic device (200) may adjust the gain of a designated frequency band of the first audio signal. In operation 910, the electronic device (200) may adjust the gain so that a designated frequency band (e.g., 6 to 16 kHz) of the first audio signal is attenuated based on multi-band dynamic range control (MBDRC). In this case, the first audio signal may be processed to have an artificial sense of space, so that it may be perceived as being transmitted from a distance away from the user and / or from behind the user.
[0115] According to one embodiment, the electronic device (200) may perform mixing of the adjusted first audio signal and the second audio signal in operation 920, and transmit the mixed audio signal to the external electronic device in operation 930. The electronic device (200) may perform mixing of the first audio signal and the second audio signal so that the second audio signal may be output by at least partially overlapping with the first audio signal, and transmit the mixed audio signal to the external electronic device through one channel.
[0116] According to various embodiments, when two channels are used, the electronic device (200) may not perform operation 920. In this case, the electronic device (200) may transmit the first audio signal and the second audio signal to the external electronic device through two different channels, respectively. For example, in operation 930, the electronic device (200) may transmit the second audio signal to the external electronic device through a second channel while transmitting the first audio signal to the external electronic device through a first channel.
[0117] FIG. 10 is a flowchart illustrating a method for adjusting the volume of a first audio signal in an electronic device (200), according to an embodiment. According to various embodiments, the electronic device (200) may obtain a first audio signal including voice data of a first language spoken by a user while a call connection is established with an external electronic device through a microphone (e.g., the input module (150) of FIG. 1 or the microphone (210) of FIG. 2). The electronic device (200) may obtain a second audio signal obtained by translating the voice data of the first language into a second language, and transmit the obtained second audio signal together with the first audio signal to the external electronic device. The operations of FIG. 10 may be understood as functions performed by at least one processor included in the electronic device (200) (e.g., the processor (120) of FIG. 1 or the processor (240) of FIG. 2).
[0118] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0119] Referring to FIG. 10, in operation 1010, the electronic device (200) may adjust the volume level of the first audio signal. In operation 1010, the electronic device (200) may adjust the volume level of the first audio signal to a specified level (e.g., 9 dB level) or lower, and may adjust the volume level of the second audio signal to be maintained or output at a volume greater than that of the first audio signal. In this case, even if the first audio signal and the second audio signal are output in an overlapping manner, the second audio signal may be recognized as being relatively emphasized.
[0120] According to one embodiment, the electronic device (200) may perform mixing of the adjusted first audio signal and the second audio signal in operation 1020, and transmit the mixed audio signal to the external electronic device in operation 1030. The electronic device (200) may perform mixing of the first audio signal and the second audio signal so that the second audio signal may be output by at least partially overlapping with the first audio signal, and transmit the mixed audio signal to the external electronic device through one channel.
[0121] According to various embodiments, when two channels are used, the electronic device (200) may not perform operation 1020. In this case, the electronic device (200) may transmit the first audio signal and the second audio signal to the external electronic device through two different channels, respectively. For example, in operation 1030, the electronic device (200) may transmit the second audio signal to the external electronic device through a second channel while transmitting the first audio signal to the external electronic device through a first channel.
[0122] FIG. 11 is a diagram illustrating a user interface (1100) provided while using a translation function during a call, according to one embodiment.
[0123] Referring to FIG. 11, the electronic device (200) can display a user interface (1100) including texts corresponding to audio signals transmitted and received between the electronic device (200) and the external electronic device while a call is established with the external electronic device on a display (e.g., the display module (160) of FIG. 1 or the display (230) of FIG. 2).
[0124] In one embodiment, when the electronic device (200) obtains a first audio signal including voice data of a first language spoken by a user through a microphone (e.g., the input module (150) of FIG. 1 or the microphone (210) of FIG. 2), the electronic device (200) may perform voice recognition on the first audio signal to identify a first text (1110) corresponding to the voice data of the first language. The electronic device (200) may display the first text (1110) on the user interface (1100).
[0125] In one embodiment, when the electronic device (200) obtains a second audio signal that translates the voice data of the first language into a second language, the electronic device (200) can confirm (1120) a second text that translates the first text into the second language. The electronic device (200) can display the second text (1120) on the user interface (1100).
[0126] According to various embodiments, the electronic device (200) may additionally display a third text (1130) on the user interface (1100) to verify the translation accuracy (or reliability) of the second audio signal. For example, the electronic device (200) may provide a text that is a second text in a second language translated back into a first language as the third text (1130) through the user interface (1100). As another example, the electronic device (200) may provide a second text that is a third text translated back into a third language as the third text (1130) through the user interface (1100). Here, the third language may be a language (e.g., English) that is recognizable to the user and is different from the first language and the second language.
[0127] FIG. 12 is a flowchart illustrating a method for transmitting a first audio signal and a second audio signal using two channels in an electronic device (200), according to one embodiment. The operations of FIG. 12 may be understood as functions performed by at least one processor included in the electronic device (200) (e.g., the processor (120) of FIG. 1 or the processor (240) of FIG. 2).
[0128] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0129] Referring to FIG. 12, in operation 1210, the electronic device (200) can check channel information supported by the electronic device (200). For example, the electronic device (200) can check whether the audio signal is transmitted using a single channel or two channels while a call connection is established with an external electronic device.
[0130] According to one embodiment, the electronic device (200) may determine whether the audio signal is transmitted to the external electronic device using a single channel as a result of checking the channel information in operation 1220. If it is determined as a result of the determination that the audio signal is transmitted through a single channel (operation 1220-Yes), the electronic device (200) may perform mixing of a first audio signal and a second audio signal in operation 1230. Here, the first audio signal may be an audio signal including voice data spoken by a user in a first language, and the second audio signal may be an audio signal in which voice data in the first language is translated into a second language. In operation 1230, the electronic device (200) may mix the first audio signal and the second audio signal so that the second audio signal may be output by at least partially overlapping with the first audio signal. The electronic device (200) may transmit the mixed signal to the external electronic device using a single channel in operation 1235.
[0131] As a result of the above determination, if it is determined that the audio signal can be transmitted through two channels (operation 1220-No), the electronic device (200) can assign the first audio signal and the second audio signal to two channels, respectively, in operation 1240. The electronic device (200) can transmit the first audio signal and the second audio signal to the external electronic device using the two assigned channels in operation 1245. In operation 1245, the electronic device (200) can transmit the second audio signal through the second channel while transmitting the first audio signal through the first channel.
[0132] FIG. 13 is a flowchart illustrating a method for receiving a first audio signal and a second audio signal transmitted using two channels in an electronic device (200), according to one embodiment. The operations of FIG. 13 may be understood as functions performed by at least one processor included in the electronic device (200) (e.g., the processor (120) of FIG. 1 or the processor (240) of FIG. 2).
[0133] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0134] Referring to FIG. 13, in operation 1310, the electronic device (200) may receive a first audio signal and a second audio signal from an external electronic device with which a call connection is established. Here, the first audio signal may be an audio signal including voice data spoken by a user of the external electronic device in a first language, and the second audio signal may be an audio signal in which voice data in the first language is translated into a second language.
[0135] According to one embodiment, the electronic device (200) may determine whether the first audio signal and the second audio signal received from the external electronic device were transmitted through a single channel in operation 1320. If, as a result of the determination, the first audio signal and the second audio signal were received through a single channel (operation 1320-Yes), the electronic device (200) may output the first audio signal and the second audio signal through the single channel in operation 1340. In this case, the output signal may be an audio signal mixed by the external electronic device.
[0136] As a result of the above determination, if the first audio signal and the second audio signal are received through two channels (operation 1320-No), the electronic device (200) can determine whether stereo output is possible for the electronic device (200) in operation 1330. As a result of the above determination, if stereo output is not possible (operation 1330-No), the electronic device (200) can perform mixing of the first audio signal and the second audio signal in operation 1335, and output the mixed audio signal through a single channel supported by the electronic device (200) in operation 1340.
[0137] As a result of the above determination, if stereo output is possible (operation 1330-Yes), the electronic device (200) may output the first audio signal and the second audio signal using two channels in operation 1350. For example, in operation 1350, the electronic device (200) may output the first audio signal through the left speaker and the second audio signal through the right speaker. In this case, the electronic device (200) may output the second audio signal by at least partially overlapping the first audio signal while outputting the first audio signal. According to various embodiments, the electronic device (200) may output only one audio signal selected by the user among the first audio signal and the second audio signal.
[0138] According to one embodiment, an electronic device (e.g., electronic device (200)) includes one or more microphones (210), a communication circuit (220), at least one processor (240), and a memory (250) storing instructions, wherein the instructions are executed by the at least one processor (240) to cause the electronic device (200) to obtain a first audio signal including voice data of a first language from the one or more microphones (210) while a call is established with an external electronic device, obtain a second audio signal in which the voice data of the first language is translated into a second language specified by a user, identify a voice characteristic of the first audio signal based on a frequency analysis of the voice data, determine setting information related to an output of the first audio signal or the second audio signal based on the voice characteristic, and transmit the first audio signal and the second audio signal to the external electronic device using the communication circuit (220) based on the determined setting information.
[0139] In one embodiment, the instructions may be executed by the at least one processor (240) to cause the electronic device (200) to determine a first voice type corresponding to the first audio signal based on a voice characteristic of the first audio signal, and to set a voice of a different type from the first voice type as the voice of the second audio signal.
[0140] In one embodiment, the instructions may be executed by the at least one processor (240) to cause the electronic device (200) to identify a frequency component of the voice data included in the first audio signal and to set a voice in a frequency band that does not overlap with the identified frequency component as the voice of the second audio signal.
[0141] In one embodiment, the instructions may be executed by the at least one processor (240) to cause the electronic device (200) to identify a first frequency spectrum for the voice data included in the first audio signal, identify a plurality of second frequency spectra corresponding to each of a plurality of voice types configurable for the second audio signal, identify a second voice type corresponding to a frequency spectrum having the lowest similarity with the first frequency spectrum among the plurality of second frequency spectra, and determine the identified second voice type as the voice of the second audio signal.
[0142] In one embodiment, the instructions may be executed by the at least one processor (240) to cause the electronic device (200) to adjust a gain of a specified frequency band for the first audio signal.
[0143] In one embodiment, the instructions may be executed by the at least one processor (240) to cause the electronic device (200) to set the volume level of the first audio signal to a specified level or lower.
[0144] In one embodiment, the instructions may be executed by the at least one processor (240) to cause the electronic device (200) to perform speech recognition on speech data in the first language to identify a first text, obtain a second text in which the first text is translated into the second language, and convert the second text into speech data to obtain the second audio signal.
[0145] In one embodiment, the electronic device further includes a display (230), and the instructions cause the display (230) to display a user interface including the first text, the second text, and a third text, wherein the third text may include at least one of a text translated from the second text into the first language, or a text translated from the first text or the second text into the third language.
[0146] In one embodiment, the instructions may be executed by the at least one processor (240) to cause the electronic device (200) to perform mixing of the first audio signal and the second audio signal and transmit the mixed audio signal to the external electronic device using a first channel.
[0147] In one embodiment, the instructions may be executed by the at least one processor (240) to cause the electronic device (200) to transmit the first audio signal to the external electronic device using a first channel and to transmit the second audio signal to the external electronic device using a second channel.
[0148] According to one embodiment, a method may include: obtaining a first audio signal including voice data of a first language from one or more microphones (210) of an electronic device (200) while a call is established with an external electronic device; obtaining a second audio signal in which the voice data of the first language is translated into a second language specified by a user; identifying a voice characteristic of the first audio signal based on a frequency analysis of the voice data; determining setting information related to output of the first audio signal or the second audio signal based on the voice characteristic; and transmitting the first audio signal and the second audio signal to the external electronic device based on the determined setting information using a communication circuit (220) of the electronic device (300).
[0149] In one embodiment, the operation of determining setting information related to the output of the second audio signal may include an operation of determining a first voice type corresponding to the first audio signal based on a voice characteristic of the first audio signal, and an operation of setting a voice of a type different from the first voice type as the voice of the second audio signal.
[0150] In one embodiment, the operation of determining setting information related to the output of the second audio signal may include the operation of confirming a frequency component of the voice data included in the first audio signal, and the operation of setting a voice in a frequency band that does not overlap with the confirmed frequency component as the voice of the second audio signal.
[0151] In one embodiment, the operation of setting a voice in a frequency band that does not overlap with the confirmed frequency component as the voice of the second audio signal may include the operation of checking a first frequency spectrum for the voice data included in the first audio signal, the operation of checking a plurality of second frequency spectra corresponding to each of a plurality of voice types that can be set for the second audio signal, the operation of identifying a second voice type corresponding to a frequency spectrum having the lowest similarity with the first frequency spectrum among the plurality of second frequency spectra, and the operation of determining the identified second voice type as the voice of the second audio signal.
[0152] In one embodiment, the operation of determining setting information related to the output of the first audio signal may include the operation of adjusting a gain of a specified frequency band for the first audio signal.
[0153] In one embodiment, the operation of determining setting information related to output of the first audio signal may include the operation of setting a volume level for the first audio signal to a specified level or lower.
[0154] In one embodiment, the operation of obtaining the second audio signal may include an operation of performing speech recognition on speech data in the first language to identify a first text, an operation of obtaining a second text in which the first text is translated into the second language, and an operation of converting the second text into speech data to obtain the second audio signal.
[0155] In one embodiment, the method further includes an action of displaying, through a display (230) of the electronic device (200), a user interface including the first text, the second text, and a third text, wherein the third text may include at least one of a text translated from the second text into the first language, or a text translated from the first text or the second text into the third language.
[0156] In one embodiment, the method may further include an operation of mixing the first audio signal and the second audio signal, and an operation of transmitting the mixed audio signal to the external electronic device using a first channel.
[0157] In one embodiment, the method may further include transmitting the first audio signal to the external electronic device using a first channel, and transmitting the second audio signal to the external electronic device using a second channel.
[0158] In one embodiment, a computer-readable recording medium having recorded thereon programs executable on a computer may be caused by the computer to perform the following operations: obtaining a first audio signal including voice data of a first language from one or more microphones (210) of the electronic device (200) while a call is established with an external electronic device; obtaining a second audio signal in which the voice data of the first language is translated into a second language specified by a user; identifying a voice characteristic of the first audio signal based on a frequency analysis of the voice data; determining setting information related to output of the first audio signal or the second audio signal based on the voice characteristic; and transmitting the first audio signal and the second audio signal to the external electronic device based on the determined setting information using a communication circuit (220) of the electronic device (300).
[0159] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains.
[0160] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments disclosed in this document are not limited to the aforementioned devices.
[0161] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments disclosed in this document are not limited to the aforementioned devices.
[0162] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B," "at least one of A and B," "at least one of A or B," "A, B, or C," "at least one of A, B, and C," and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0163] The term "module" as used herein may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0164] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more commands stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one command among the one or more commands stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one command called. The one or more commands may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0165] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smartphones), online (e.g., by download or upload). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0166] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include a single or multiple entities. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device (200), One or more microphones (210); Communication circuit (220); At least one processor (240); and Executed by at least one processor (240), the electronic device (200): While a call is established with an external electronic device, a first audio signal including voice data of a first language is acquired from the one or more microphones (210), Obtaining a second audio signal translated from the speech data of the first language into a second language specified by the user, Confirming the voice characteristics of the first audio signal based on frequency analysis of the above voice data, Determine setting information related to the output of the first audio signal or the second audio signal based on the above voice characteristics, An electronic device comprising a memory (250) storing instructions for transmitting the first audio signal and the second audio signal to the external electronic device using the communication circuit (220) based on the determined setting information.
2. In claim 1, The above instructions are executed by the at least one processor (240), so that the electronic device (200) Determine a first voice type corresponding to the first audio signal based on the voice characteristics of the first audio signal, An electronic device that sets a voice of a different type from the first voice type as the voice of the second audio signal.
3. In claim 1, The above instructions are executed by the at least one processor (240), so that the electronic device (200) Check the frequency components of the voice data included in the first audio signal, An electronic device that sets a voice in a frequency band that does not overlap with the above-mentioned confirmed frequency component as the voice of the second audio signal.
4. In claim 3, The above instructions are executed by the at least one processor (240), so that the electronic device (200) Checking the first frequency spectrum for the voice data included in the first audio signal, Confirming a plurality of second frequency spectra corresponding to each of a plurality of voice types that can be set for the second audio signal, Identifying a second voice type corresponding to a frequency spectrum having the lowest similarity to the first frequency spectrum among the plurality of second frequency spectra, An electronic device that determines the identified second voice type as the voice of the second audio signal.
5. In claim 1, The above instructions are executed by the at least one processor (240), so that the electronic device (200) An electronic device that adjusts the gain of a specified frequency band for the first audio signal.
6. In claim 1, The above instructions are executed by the at least one processor (240), so that the electronic device (200) An electronic device that sets the volume level of the first audio signal to a specified level or lower.
7. In claim 1, The above instructions are executed by the at least one processor (240), so that the electronic device (200) Perform speech recognition on the speech data of the first language to verify the first text, The first text above obtains a second text translated into the second language, An electronic device that converts the second text into voice data to obtain the second audio signal.
8. In claim 7, Further including a display (230), The above instructions are executed by the at least one processor (240), so that the electronic device (200) Through the above display (230), a user interface including the first text, the second text, and the third text is displayed, An electronic device wherein the third text includes at least one of a text translated from the second text into the first language, or a text translated from the first text or the second text into a third language.
9. In claim 1, The above instructions are executed by the at least one processor (240), so that the electronic device (200) Perform mixing of the first audio signal and the second audio signal, An electronic device that transmits the mixed audio signal to the external electronic device using the first channel.
10. In claim 1, The above instructions are executed by the at least one processor (240), so that the electronic device (200) Transmitting the first audio signal to the external electronic device using the first channel, An electronic device that transmits the second audio signal to the external electronic device using a second channel.
11. In the method, An action of obtaining a first audio signal including voice data of a first language from one or more microphones (210) of the electronic device (200) while a call is established with the external electronic device; An operation of obtaining a second audio signal translated from the speech data of the first language into a second language specified by the user; An operation of confirming the voice characteristics of the first audio signal based on frequency analysis of the voice data; An operation of determining setting information related to the output of the first audio signal or the second audio signal based on the voice characteristics; and An operation of transmitting the first audio signal and the second audio signal to the external electronic device based on the determined setting information using the communication circuit (220) of the electronic device (300). A method comprising:
12. In claim 11, The operation of determining the setting information related to the output of the above second audio signal is: An operation of determining a first voice type corresponding to the first audio signal based on the voice characteristics of the first audio signal; and A method comprising an action of setting a voice of a different type from the first voice type as the voice of the second audio signal.
13. In claim 11, The operation of determining the setting information related to the output of the above second audio signal is: An operation of checking the frequency components of the voice data included in the first audio signal; and A method comprising an action of setting a voice in a frequency band that does not overlap with the above-determined frequency component as the voice of the second audio signal.
14. In claim 11, The operation of obtaining the second audio signal comprises: An action of performing speech recognition on speech data of the first language to confirm a first text; An operation of obtaining a second text translated from the first text into the second language; and A method comprising an action of converting the second text into voice data to obtain the second audio signal.
15. In claim 14, Further comprising an action of displaying a user interface including the first text, the second text and the third text through the display (230) of the electronic device (200), A method wherein the third text comprises at least one of a text translated from the second text into the first language, or a text translated from the first text or the second text into a third language.
Citation Information
Patent Citations
Systems, methods, and apparatus for context descriptor transmission
KR1020100113144A
Data augmentation method based on generative adversarial network for improving object detection performance
KR1020240128330A
Method for interpreting
KR102056329B1
Systems and methods for providing real-time automated language translations
US20230096543A1
KR20200027331A