Electronic device, method, and non-transitory computer-readable storage medium for converting voice data related to application

The electronic device addresses language barriers in VoIP calls by converting voice data to text or speech in real-time, ensuring effective communication through integrated voice recognition and translation capabilities.

WO2025183379A1PCT designated stage Publication Date: 2025-09-04SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/001722
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-09
Filing Date
2025-02-05
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing voice over internet protocol (VoIP) systems lack the ability to efficiently translate voice data between different languages in real-time during calls, hindering effective communication across language barriers.

Method used

An electronic device equipped with a processor, memory, and communication circuitry that converts voice data from one language to another using a voice recognizer and translator, displaying the translated text or speech on a user interface during a VoIP call.

Benefits of technology

Enables real-time language translation during VoIP calls, facilitating seamless communication by converting voice data into text or speech in a different language, enhancing user experience and understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025001722_04092025_PF_FP_ABST
    Figure KR2025001722_04092025_PF_FP_ABST
Patent Text Reader

Abstract

According to an embodiment, a method performed by an electronic device comprises an operation of identifying that a call connection with an external electronic device is established through an application related to a voice over Internet protocol (VoIP). The method comprises an operation of, while the call connection is established, obtaining first voice data of a first language type, the first voice data being related to the application. The method comprises an operation of obtaining a second text of a second language type distinguished from the first language type on the basis of a first text of the first language type converted from the first voice data of the first language type. The method comprises an operation of displaying, through a display of the electronic device, a visual object including the first text and the second text, in an overlapping manner on a user interface for the application.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method, and non-transitory computer-readable storage medium for converting voice data for an application

[0001] The following descriptions relate to electronic devices, methods, and non-transitory computer-readable storage media for converting voice data for an application.

[0002] Electronic devices can provide a calling service for transmitting and receiving audio and video signals to external electronic devices located in different locations via applications related to voice over internet protocol (VoIP). Users of the electronic devices can utilize this service to communicate with other users located in different locations.

[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.

[0004] According to one embodiment, an electronic device may include a display, a speaker, a microphone, communication circuitry, at least one processor including a processing circuit, and a memory including one or more storage media storing instructions. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify that a call connection is established with an external electronic device through an application relating to voice over internet protocol (VoIP). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain first voice data of a first language type related to the application while the call connection is established. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain second text of a second language type distinct from the first language type based on first text of the first language type converted from the first voice data of the first language type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, through the display, a visual object comprising the first text and the second text as an overlay on a user interface for the application.

[0005] According to one embodiment, a method performed in an electronic device may include an operation of identifying that a call connection is established with an external electronic device through an application relating to voice over internet protocol (VoIP). The method may include an operation of obtaining, while the call connection is established, first voice data of a first language type related to the application. The method may include an operation of obtaining, based on first text of the first language type, a second text of a second language type that is distinct from the first language type, and which is converted from the first voice data of the first language type. The method may include an operation of displaying, through a display of the electronic device, a visual object including the first text and the second text as an overlay on a user interface for the application.

[0006] According to one embodiment, an electronic device may include at least one processor including a speaker, a microphone, a communication circuit, a processing circuit, and a memory including one or more storage media storing instructions. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain, from an application relating to voice over internet protocol (VoIP), first voice data of a first language type received from an external electronic device, using a voice data manager configured to obtain voice data relating to the application while a call connection with the external electronic device is established through the application. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to provide, using the voice data manager, the first voice data of the first language type to a voice recognizer. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to convert, using the speech recognizer, the first speech data of the first language type into first text of the first language type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to convert, using a translator, the first text of the first language type into second text data of a second language type distinct from the first language type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to convert, using a speech converter, the second text of the second language type into second speech data of the second language type.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to output the second speech data of the second language type through the speaker.

[0007] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.

[0008] FIG. 2 is a simplified block diagram of an electronic device according to one embodiment.

[0009] FIG. 3A illustrates an example of a framework of an electronic device for general telephone services, according to one embodiment.

[0010] Figure 3b illustrates an example of a user interface for a call translation service performed while a normal telephone service is being performed.

[0011] FIG. 4 illustrates an example of a framework of an electronic device for providing call translation services in an application relating to VoIP, according to one embodiment.

[0012] FIG. 5A illustrates a flowchart illustrating the operation of an electronic device according to one embodiment.

[0013] FIG. 5b illustrates a flowchart illustrating the operation of an electronic device according to one embodiment.

[0014] FIG. 6A illustrates an example of operation of an electronic device for providing a currency translation service, according to one embodiment.

[0015] FIG. 6b illustrates an example of operation of an electronic device for providing a currency translation service, according to one embodiment.

[0016] FIG. 7 illustrates an example of operation of an electronic device for initiating a currency translation service, according to one embodiment.

[0017] FIG. 8 illustrates an example of operation of an electronic device for setting a function for a call translation service within a system setting, according to one embodiment.

[0018] FIG. 9 illustrates an example of operation of an electronic device for providing a call translation service while a video call service is provided through an application for VoIP, according to one embodiment.

[0019] FIG. 10 illustrates an example of operation of an electronic device for providing a call translation service while a video call service is provided through an application for VoIP, according to one embodiment.

[0020] FIGS. 11A and 11B illustrate an example of a positional relationship between a first housing and a second housing in an unfolded state and a folded state of an electronic device according to one embodiment.

[0021] FIG. 11c illustrates an example of a screen for a currency translation service displayed based on the positional relationship between the first housing and the second housing, according to one embodiment.

[0022] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.

[0023] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.

[0024] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0025] The processor (120) may control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing, for example, software (e.g., a program (140)), and may perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculation, the processor (120) may store a command or data received from another component (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the command or data stored in the volatile memory (132), and store the resulting data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0026] The auxiliary processor (123) may control at least a part of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0027] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0028] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0029] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0030] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. According to one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0031] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0032] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0033] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0034] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0035] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., the electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0036] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0037] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0038] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0039] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0040] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0041] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0042] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0043] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0044] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0045] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service by itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0046] According to one embodiment, an electronic device (e.g., the electronic device (101) of FIG. 1) may execute an application related to voice over internet protocol (VoIP). The VoIP application may be used to establish a call connection with an external electronic device. While the call connection with the external electronic device is established via the VoIP application, the user of the electronic device may communicate with the user of the external electronic device. The language type used by the user of the electronic device may be different from the language type used by the user of the external electronic device.

[0047] For example, a language type used by a user of an external electronic device may be a first language type, and a language type used by a user of the electronic device may be a second language type. The electronic device may obtain voice data of the first language type from the external electronic device. The electronic device may convert the voice data of the first language type into text of the second language type. The electronic device may provide a service for call translation by displaying the text of the second language type on the display of the electronic device.

[0048] However, for the above-described operation, a component for obtaining voice data from a VoIP application may be required. Therefore, the following specification will describe technical features for providing a call translation service using a component for obtaining voice data from a VoIP application. The same reference numbers used in the following specification may indicate the same description.

[0049] FIG. 2 is a simplified block diagram of an electronic device according to one embodiment.

[0050] Referring to FIG. 2, the electronic device (200) may be used to provide a service for currency translation. For example, the electronic device (200) may include at least some or all of the components of the electronic device (101) of FIG. 1.

[0051] According to one embodiment, the electronic device (200) may include a processor (210), a memory (220), a communication circuit (230), a speaker (240), a microphone (250), a display (260), and / or a camera (270). Depending on the embodiment, the electronic device (200) may include at least one of the processor (210), the memory (220), the communication circuit (230), the speaker (240), the microphone (250), the display (260), and the camera (270). For example, at least some of the processor (210), the memory (220), the communication circuit (230), the speaker (240), the microphone (250), the display (260), and the camera (270) may be omitted depending on the embodiment.

[0052] For example, when a video call service is performed in an application related to VoIP, the electronic device (200) may include a processor (210), a memory (220), a communication circuit (230), a speaker (240), a microphone (250), a display (260), and a camera (270). For example, when a voice call service is performed in an application related to VoIP, the electronic device (200) may not include a camera (270). For example, when the electronic device (200) is connected to an external electronic device for performing the functions of the speaker (240) and / or the microphone (250), the electronic device (200) may not include the speaker (240) and / or the microphone (250).

[0053] For example, the electronic device (200) may include at least some of the components of the electronic device (101) of FIG. 1. Although not illustrated, the electronic device (200) may include various components in addition to the processor (210), memory (220), communication circuit (230), speaker (240), microphone (250), display (260), and camera (270).

[0054] According to one embodiment, the processor (210) may be operatively or operably coupled with or connected with a memory (220), a communication circuit (230), a speaker (240), a microphone (250), a display (260), and a camera (270). The processor (210) being operatively or operably coupled with the memory (220), the communication circuit (230), the speaker (240), the microphone (250), the display (260), and the camera (270) may mean that the processor (210) can control the memory (220), the communication circuit (230), the speaker (240), the microphone (250), the display (260), and the camera (270). For example, the memory (220), communication circuit (230), speaker (240), microphone (250), display (260), and camera (270) may be controlled by the processor (210). For example, the processor (210) may correspond to the processor (120) of FIG. 1.

[0055] Although illustrated based on different blocks, the embodiment is not limited thereto, and some of the hardware of FIG. 2 (e.g., at least a portion of the processor (210), memory (220), communication circuit (230), speaker (240), microphone (250), display (260), and camera (270)) may be included in a single integrated circuit such as a system on a chip (SoC).

[0056] According to one embodiment, the processor (210) may include at least one processor. For example, the processor (210) may include a main processor that performs high-performance processing and a secondary processor that performs low-power processing.

[0057] According to one embodiment, the processor (210) may include a hardware component for processing data based on one or more instructions. The hardware component for processing data may include, for example, an arithmetic and logic unit (ALU), a field programmable gate array (FPGA), and / or a central processing unit (CPU).

[0058] For example, the processor (210) may include an application processor, a supplementary processor (e.g., a sensor hub, a microcontroller unit (MCU)), a central processor unit (CPU), a neural processing unit (NPU), a graphic processing unit (GPU), and / or a processor for IoT (e.g., a processor integrated with a communication module).

[0059] According to one embodiment, the memory (220) of the electronic device (200) may include a circuit and / or a storage medium for storing data and / or instructions input and / or output to the processor (210). The memory (220) may include, for example, volatile memory such as random-access memory (RAM) and / or non-volatile memory such as read-only memory (ROM). The non-volatile memory may be referred to as storage. The volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). The non-volatile memory may include, for example, at least one of programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disc, solid state drive (SSD), and embedded multi media card (eMMC). The processor (210) of the electronic device (200) may execute instructions of the memory (220) within the electronic device (200) to perform functions and / or operations indicated by the instructions. For example, if the electronic device (200) includes at least one processor, the at least one processor may be configured to collectively or individually execute the instructions. For example, the memory (220) may correspond to the memory (130) of FIG. 1.

[0060] According to one embodiment, the communication circuit (230) may correspond to at least a portion of the communication module (190) of FIG. 1. For example, the communication circuit (230) may be used for various radio access technologies (RATs). For example, the communication circuit (230) may be used to perform Bluetooth communication, wireless local area network (WLAN) communication, Zigbee communication, near field communication (NFC), ultra wide band (UWB) communication, UWB communication, or ANT+ communication. For example, the communication circuit (230) may be used to perform cellular communication. According to an embodiment, the communication circuit (230) may be configured to be integrated with the processor (210).

[0061] For example, when a general telephone service is performed, the processor (210) may transmit the user's voice data (or video data) as an analog signal using a communication circuit. The processor (210) (e.g., a CP (communication processor)) may transmit the analog signal to a PSTN (public switched telephone network), and the PSTN may convert the received analog signal into a digital signal and then transmit the digital signal to an external electronic device. The general telephone service may be referred to as a CP (communication processor) call service. A framework for the general telephone service and specific examples of a call translation service performed in the general telephone service will be described later with reference to FIGS. 3A and 3B.

[0062] For example, the processor (210) may establish a call connection with an external electronic device using a voice over internet protocol (VoIP) call path provided by a server using a communication circuit (230). When a call connection with an external electronic device is established using the VoIP call path, the electronic device (200) may transmit voice data (or video data) as a digital signal. A framework for a VoIP call service and specific examples of a call translation service performed in a VoIP call service will be described later with reference to FIGS. 4A and 4B.

[0063] According to one embodiment, a speaker (240) may be used to output (or provide) sound. The processor (210) may receive an audio signal received from an external electronic device and output the received audio signal through the speaker (240). For example, the speaker (240) may correspond to the audio output module (155) of FIG. 1.

[0064] According to one embodiment, the electronic device (200) may include a microphone (250). The microphone (250) may identify an electrical signal corresponding to vibration in the atmosphere. For example, the processor (210) may obtain an audio signal (e.g., the user's voice) from a user using the microphone (250). The processor (210) may transmit the audio signal obtained from the user to an external electronic device.

[0065] According to one embodiment, the display (260) can output visualized information to the user. The display (260) can include a liquid crystal display (LCD), a plasma display panel (PDP), one or more light emitting diodes (LEDs), and / or one or more organic light emitting diodes (OLEDs). According to one embodiment, the display (260) can include a sensor (e.g., a touch sensor panel (TSP)) for detecting an external object (e.g., a user's finger) on the display (260). For example, the display (260) can correspond to the display module (160) of FIG. 1.

[0066] According to one embodiment, the electronic device (200) may include a camera (270). For example, the camera (270) may correspond to at least a portion of the camera module (180) of FIG. 1. For example, the camera (270) may include at least one camera. For example, the camera (270) may include one or more optical sensors (e.g., a charged coupled device (CCD) sensor, a complementary metal oxide semiconductor (CMOS) sensor) that generate electrical signals representing the color and / or brightness of light. The plurality of optical sensors included in the camera (270) may be arranged in the form of a two-dimensional array. The camera (270) may acquire the electrical signals of each of the plurality of optical sensors substantially simultaneously, and generate two-dimensional frame data corresponding to light reaching the optical sensors of the two-dimensional array. For example, photographic data captured using the camera (270) may mean one (a) two-dimensional frame data acquired from the camera (270). For example, video data captured using a camera (270) may mean a sequence of multiple two-dimensional frame data acquired from the camera (270).

[0067] FIG. 3A illustrates an example of a framework of an electronic device for general telephone services, according to one embodiment.

[0068] Referring to FIG. 3a, a CP (communication processor) layer (300), a kernel layer (310), an audio framework layer (320), a multimedia framework layer (330), and an application layer (340) may be configured for a call translation service for a general telephone service (or CP call service).

[0069] According to one embodiment, a general telephone service may be performed in the CP layer (300). For example, when a call connection for a general telephone service is established, the processor (210) may transmit voice data to an external electronic device. Specifically, the processor (210) (e.g., CP) may obtain voice data from a user of the electronic device (200) using a microphone (250). The processor (210) may preprocess the voice data using a preprocessor (301-1). The processor (210) may modulate and / or demodulate the voice data using a vocoder (302-1). The processor (210) may process the voice data using a voice engine (303-1). The processor (210) may transmit the voice data to the external electronic device via a modem (304-1).

[0070] For example, when a call connection for a general telephone service is established, the processor (210) can receive voice data from an external electronic device. Specifically, the processor (210) (e.g., CP) can receive voice data from the external electronic device via a modem (304-2). The processor (210) can process the voice data using a voice engine (303-2). The processor (210) can modulate and / or demodulate the voice data using a vocoder (302-2). The processor (210) can preprocess the voice data using a preprocessor (301-2). The processor (210) can output the voice data using a speaker (240).

[0071] According to one embodiment, while a general telephone service (or CP call service) is performed, the processor (210) may provide a call translation service. In order to provide the call translation service, the processor (210) may obtain second voice data of a second language type based on first voice data of a first language type received from an external electronic device. The processor (210) may output the second voice data of the second language type through a speaker (240). In order to provide the call translation service, the processor (210) may obtain fourth voice data of the first language type based on third voice data of the second language type received through a microphone (250) of the electronic device (200). The processor (210) may transmit the fourth voice data to the external electronic device.

[0072] For example, the processor (210) may transmit the first voice data of the first language type obtained through the vocoder (302-2) to the kernel audio (311). The processor (210) may provide the first voice data to the voice data processor (331-2) through the kernel audio (311), the audio hardware abstraction layer (HAL) (321), and the audio flinger (322). The processor (210) may perform recording on the first voice data using the voice data processor (331-2). The processor (210) may provide the first voice data to the voice recognizer (341-2) through the voice data processor (331-2). The processor (210) may identify (or determine) the language type of the first voice data as the first language type using the voice recognizer (341-2). The processor (210) can obtain a first text of a first language type based on first voice data of a first language type using a voice recognizer (341-2). The processor (210) can convert the first text of the first language type into a second text of a second language type using a translator (342-2). The processor (210) can provide the second text to a voice converter (332-2) via a dialogue processor (343). The processor (210) can obtain second voice data of a second language type based on the second text of the second language type using the voice converter (332-2). The processor (210) can output the second voice data of the second language type through a speaker (240) by providing the second voice data of the second language type to a preprocessor (301-2).

[0073] For example, the processor (210) can transmit third voice data of the second language type acquired through the preprocessor (301-1) to the kernel audio (311). The processor (210) can provide the third voice data to the voice data processor (331-1) through the kernel audio (311), the audio hardware abstraction layer (HAL) (321), and the audio flinger (322). The processor (210) can use the voice data processor (331-1) to perform recording on the third voice data. The processor (210) can provide the third voice data to the voice recognizer (341-1) through the voice data processor (331-1). The processor (210) can use the voice recognizer (341-1) to identify (or determine) the language type of the third voice data as the second language type. The processor (210) can obtain a third text of the second language type based on the third voice data of the second language type using the voice recognizer (341-1). The processor (210) can convert the third text of the second language type into a fourth text of the first language type using the translator (342-1). The processor (210) can provide the fourth text to the voice converter (332-1) via the dialogue processor (343). The processor (210) can obtain fourth voice data of the first language type based on the fourth text of the first language type using the voice converter (332-1). The processor (210) can transmit the fourth voice data of the first language type to an external electronic device by providing the fourth voice data of the first language type to the vocoder (302-1).

[0074] As described above, the processor (210) may convert first voice data of a first language type received from an external electronic device into second voice data of a second language type, and output the second voice data using the speaker (240), in order to perform a call translation service while a general telephone service (or CP call service) is performed. In addition, the processor (210) may convert third voice data of a second language type acquired through a microphone (250) into fourth voice data of the first language type, and transmit the fourth voice data to the external electronic device, in order to perform a call translation service while a general telephone service (or CP call service) is performed. An example of a user interface in which a call translation service is performed while a general telephone service is performed will be described in FIG. 3B.

[0075] Figure 3b illustrates an example of a user interface for a call translation service performed while a normal telephone service is being performed.

[0076] Referring to FIG. 3B, in state (361), the processor (210) of the electronic device (200) may receive a call via the basic call service (or CP call service). In response to receiving a call via the basic call service, the processor (210) may display a screen (370) via the display (260). The screen (370) may be displayed to indicate that a call is received via the basic call service from an external electronic device. The screen (370) may include an element (371) indicating caller information, an object (372) for setting a call assistance function during a call connection, an object (373) for receiving a call, an object (374) for rejecting a call, and / or an object (375) for sending a text message to the external electronic device.

[0077] According to one embodiment, the processor (210) may change the state of the electronic device (200) from state (361) to state (362) based on an input to the object (372).

[0078] In state (362), the processor (210) may display a screen (380) for indicating a call assistance function through the display (260). The screen (380) may include an element (383) indicating caller information. The screen (380) may include an object (381) for performing a service for making a call via text and / or an object (382) for performing a call translation service.

[0079] In one embodiment, the processor (210) may change the state of the electronic device (200) from state (362) to state (363) based on an input to the object (382). For example, the processor (210) may establish a call connection for an incoming call from an external electronic device based on an input to the object (382).

[0080] In state (363), the processor (210) can display a screen (390) for providing a call translation service through the display (260). The screen (390) can include an element (391) indicating caller information, an object (392) for terminating a call connection, an element (393) for indicating the time the call connection was maintained, an object (394) for setting a target language for translation, and / or an area (399) for call translation.

[0081] For example, object (394) may include elements (395) and (396) representing target languages ​​for translation. For example, element (395) may represent English. Element (396) may represent Korean. The target languages ​​for translation specified by elements (395) and (396) may be changed according to user settings.

[0082] For example, the area (399) may be configured to display the speech of a user of the electronic device (200) and the speech of an external user of the external electronic device. Based on the start of the call translation service, the processor (210) may display text (398) indicating that the call translation service is performed in the area (399). The text (398) may be converted into voice data and transmitted to the external electronic device via a call connection.

[0083] The processor (210) can display text (365-1) representing a user's speech of the electronic device (200) in an area (399). If the text (365-1) is in Korean, the processor (210) can translate the Korean text (365-1) into English. The processor (210) can display English text (365-2) corresponding to the Korean text (365-1) in an area (399). The processor (210) can transmit voice data (or the user's speech) corresponding to the text (365-1) and voice data corresponding to the text (365-2) to an external electronic device.

[0084] For example, the processor (210) can receive voice data from an external user. The processor (210) can display text (365-3) for the received voice data. If the text (365-3) is in English, the processor (210) can translate the text (365-3) into Korean. The processor (210) can display Korean text (365-4) corresponding to the English text (365-3) in an area (399). The processor (210) can output voice data (or the external user's speech) corresponding to the text (365-3) and voice data corresponding to the text (365-4) through the speaker (240).

[0085] FIG. 4 illustrates an example of a framework of an electronic device for providing call translation services in an application relating to VoIP, according to one embodiment.

[0086] Referring to FIG. 4, the processor (210) of the electronic device (200) may include a framework for providing a call translation service while a call connection is established in a VoIP application (460) (or an application related to VoIP). For example, a communication processor (CP) layer (400), a kernel layer (410), an audio framework layer (420), a multimedia framework layer (430), and an application layer (440) may be configured to provide a call translation service while a call connection is established in the VoIP application (460) (or an application related to VoIP). For example, the CP layer (400) may include a speaker (240) and / or a microphone (250). The kernel layer (410) may include kernel audio (411). The audio framework layer (420) may include a first audio track (462), a recorder (463), an audio flinger (422), a second audio track (452), and / or a voice data manager (451). The multimedia framework layer (430) may include a voice converter (432-1), a voice converter (432-2), a voice data processor (431-1), and / or a voice data processor (431-2). The application layer (440) may include a VoIP application (460), a translation service manager (461), a translator (442-1), a translator (442-2), a voice recognizer (441-1), a voice recognizer (441-2), and / or a conversation processor (443). Depending on the embodiment, the components (or blocks) included in each layer may be changed. For example, the translation service manager (461) may be included in the multimedia framework layer (430). For example, at least some of the first audio track (462), the second audio track (452), and / or the recorder (463) may be included in the multimedia framework layer (430).

[0087] According to one embodiment, a call connection regarding VoIP may be established in the VoIP application (460). When the call connection regarding VoIP is established, the processor (210) may transmit voice data to an external electronic device through the VoIP application (460). For example, the processor (210) may obtain voice data from a user of the electronic device (200) using a microphone (250). The processor (210) may transmit the voice data to an audio flinger (422) using kernel audio (411). The processor (210) may transmit the voice data obtained from the audio flinger (422) to the VoIP application (460) through a recorder (463). The processor (210) may transmit the voice data to the external electronic device through the call connection regarding VoIP using the VoIP application (460). For example, the kernel audio (411) may be configured to transmit and receive audio data between hardware (e.g., a microphone (250) and / or a speaker (240)) and the audio flinger (422). For example, the audio flinger (422) may be configured to mix the audio data or provide a path for passing the audio data to a call translation service provider (450), the kernel audio (411), and / or a VoIP application (460).

[0088] When a call connection for VoIP is established, the processor (210) can receive voice data from an external electronic device through the VoIP application (460). For example, the processor (210) can transmit the voice data received from the external electronic device to the first audio track (462) using the VoIP application (460). The processor (210) can transmit the voice data to the audio flinger (422) through the first audio track (462). The processor (210) can transmit the voice data to the kernel audio (411) through the audio flinger (422). The processor (210) can transmit the voice data to the speaker (240) through the kernel audio (411). The processor (210) can output the voice data using the speaker (240).

[0089] In general, when only a phone service (or VoIP phone service) is performed through a VoIP application (460), since the VoIP application (460) is a third-party application, a path for transmitting audio data (e.g., incoming voice data and / or outgoing voice data) regarding the VoIP application (460) to the translator (442-1, 442-2) may not be configured. By additionally configuring a voice data manager (451) in the electronic device (200), the processor (210) can obtain voice data regarding the VoIP application (460) and transmit the obtained voice data to the translator (442-1, 442-2). The processor (210) can provide a call translation service by transmitting the voice data to the translator (442-1, 442-2). For example, a path may be provided between a third-party application, such as a VoIP application (460), and a call translation service provider (450) via a voice data manager (451) and / or a second audio track (452).

[0090] Hereinafter, an operation for obtaining voice data for a VoIP application (460) using a voice data manager (451) and / or a second audio track (452) and providing a call translation service based on the obtained voice data will be described.

[0091] According to one embodiment, the processor (210) may obtain second voice data of a second language type based on first voice data of a first language type received from an external electronic device to provide a call translation service while a call connection regarding VoIP is established through the VoIP application (460). The processor (210) may output the second voice data of the second language type through the speaker (240). The processor (210) may obtain fourth voice data of the first language type based on third voice data of the second language type received through the microphone (250) of the electronic device (200) to provide a call translation service. The processor (210) may transmit the fourth voice data to the external electronic device.

[0092] According to one embodiment, when a call connection regarding VoIP is established, the processor (210) may transmit (or provide) information indicating that a call connection regarding VoIP has been established to the call translation service provider (450) through the translation service manager (461). The translation service provider (450) may activate the translation service based on the received information. For example, the translation service manager (461) may transmit information to activate the translation service provider (450) based on an input (or user input) indicating to start the call translation service while the call connection regarding VoIP is established.

[0093] For example, the VoIP application (460) may be one of the applications authorized for call translation services. For example, at least some of the applications providing VoIP call connections may be authorized to provide call translation services. The VoIP application (460) may be authorized to provide call translation services. According to an embodiment, the VoIP application (460) may be one of the applications providing VoIP call connections.

[0094] For example, when the translation service manager (461) transmits information indicating that a call connection regarding VoIP has been established to the call translation service provider (450) upon the start of the interpretation translation service, the processor (210) may identify that the VoIP application (460) is an application authorized to provide call translation services. In some embodiments, the call translation service may not be provided based on identifying that the VoIP application (460) is not an application authorized to provide call translation services.

[0095] In one embodiment, the translation service manager (461) is depicted as being distinct from the VoIP application (460), but is not limited thereto. For example, the translation service manager (461) may be included in the VoIP application (460). In another embodiment, the translation service manager (461) may be included in the translation service provider (450).

[0096] For example, the processor (210) may obtain first voice data from an external electronic device through a VoIP application (460) while a call connection regarding VoIP is established. The processor (210) may transmit (or provide) the first voice data to an audio flinger (422) through a first audio track (462). The processor (210) may transmit (or provide) the first voice data to a voice data manager (451) using the audio flinger (422). The voice data manager (451) may obtain the first voice data from the audio flinger (422). The voice data manager (451) may transmit (or provide) the first voice data to a voice data processor (431-2). The processor (210) may perform recording on the first voice data using the voice data processor (431-2). The processor (210) can transmit (or provide) first voice data to a voice recognizer (441-2) via a voice data processor (431-2). The processor (210) can identify (or determine) the language type of the first voice data as the first language type using the voice recognizer (441-2). The processor (210) can obtain a first text of the first language type based on the first voice data of the first language type using the voice recognizer (441-2). The processor (210) can provide the first text of the first language type to a translator (442-2) via a dialogue processor (443). The processor (210) can convert the first text of the first language type into a second text of the second language type using the translator (442-2). For example, the first language type and the second language type can be changed according to a user's settings. For example, the second language type may be set based on the user's input (or a target language type set by the user). For example, the first language type may be set based on speech data.The processor (210) can determine a first language type based on voice data. For example, the second language type may correspond to a language type set as the default language in the electronic device (200).

[0097] The processor (210) can provide the second text to the voice converter (432-2) through the dialogue processor (443). The processor (210) can obtain second voice data of the second language type based on the second text of the second language type using the voice converter (432-2). The processor (210) can output the second voice data of the second language type through the speaker (240) by providing the second voice data of the second language type to the kernel audio (411) through the second audio track (452) and the audio flinger (422).

[0098] According to one embodiment, the way in which the second voice data of the second language type is output may change according to user settings (e.g., setting information of the translation service manager (461)). For example, the processor (210) may output the second voice data of the second language type through the speaker (240) while outputting the first voice data of the first language type through the speaker (240). For example, the volume at which the first voice data is output may be lower than the volume at which the second voice data is output. For example, the processor (210) may output only the second voice data of the second language type through the speaker (240) without outputting the first voice data of the first language type. For example, the processor (210) may output the second voice data of the second language type through the speaker (240) after the first voice data of the first language type is output.

[0099] According to an embodiment, the processor (210) may transmit second voice data of the second language type to the external electronic device by providing the second voice data of the second language type to the VoIP application (460) via the audio flinger (422). The processor (210) may transmit the second voice data of the second language type to the external electronic device to notify that the first voice data received from the external electronic device is being translated and then provided to the user of the electronic device (200).

[0100] For example, the processor (210) may obtain third voice data from a user of the electronic device (200) through the microphone (250) while a call connection regarding VoIP is established through the VoIP application (460). The processor (210) may transmit (or provide) the third voice data to a voice data manager (451) through the kernel audio (411). The voice data manager (451) may transmit (or provide) the third voice data to a voice data processor (431-1). The processor (210) may perform recording on the third voice data using the voice data processor (431-1). The processor (210) may transmit (or provide) the third voice data to a voice recognizer (441-1) through the voice data processor (431-1). The processor (210) can identify (or determine) the language type of the third voice data as the second language type using the voice recognizer (441-1). The processor (210) can obtain a third text of the second language type based on the third voice data of the second language type using the voice recognizer (441-1). The processor (210) can provide the third text of the second language type to the translator (442-2) through the dialogue processor (443). The processor (210) can convert the third text of the second language type into a fourth text of the first language type using the translator (442-1). The processor (210) can provide the fourth text to the voice converter (432-1) through the dialogue processor (443). The processor (210) can obtain fourth voice data of the first language type based on the fourth text of the first language type using the voice converter (432-1).The processor (210) can transmit the fourth voice data of the first language type to an external electronic device through a call connection regarding VoIP using a VoIP application (460) by providing the fourth voice data of the first language type to the audio flinger (422) through the second audio track (452).

[0101] As described above, the translation service provider (450) may perform a mixing operation so that the translated voice data can be output through the speaker (240) or transmitted to an external electronic device. The dialogue processor (443) may provide the translated text to the voice converter (432-1, 432-2) so that the translated text can be converted into voice data and transmitted to an external electronic device or displayed on the electronic device (200). The voice recognizers (441-1, 441-2) may be configured to convert voice data into text in sentence units. The voice recognizers (441-1, 441-2) may be referred to as automatic speech recognition (ASR). The voice converters (432-1, 432-2) may be configured to convert text into voice data. The voice converters (432-1, 432-2) may be referred to as text to speech (TTS). The voice data processor (431-1, 431-2) may be configured to acquire voice data and provide the acquired voice data to the voice recognizer (441-1, 441-2).

[0102] For example, the audio flinger (422) may perform a mixing operation to output voice data acquired from text through the speaker (240) or transmit it to an external electronic device using the voice converter (432-1, 432-2). The voice data manager (451) may be configured to acquire voice data regarding a VoIP application (460) and transmit the acquired voice data to a call translation service provider (450). The voice data manager (451) may provide a path for transmitting voice data regarding the VoIP application (460) to the translation service provider (450) and transmitting the translated voice data to the VoIP application (460). The processor (210) may use the path provided by the voice data manager (451) to provide a call translation service for a call connection regarding VoIP established through the VoIP application (460), which is a third party application.

[0103] The framework structure of the electronic device (200) illustrated in FIG. 4 is exemplary, and at least some of the components illustrated in FIG. 4 may be modified. For example, the translator (442-1) and the translator (442-2) may be configured as a single translator. For example, the voice recognizer (441-1) and the voice recognizer (441-2) may be configured as a single voice recognizer. For example, the voice converter (432-1) and the voice converter (432-2) may be configured as a single voice converter. For example, the voice data processor (431-1) and the voice data processor (431-2) may be configured as a single voice data processor. For example, the voice recognizer (441-1) and the translator (442-1) may be configured as a single component (or block). For example, the voice recognizer (441-2) and the translator (442-2) may be configured as a single component (or block). For example, a speech recognizer (441-1), a translator (442-1), a recognizer (441-2), and a translator (442-2) may be configured as one component (or block).

[0104] In FIG. 4, the call translation service provider (450) is illustrated as being included in the electronic device (200), but this is merely an example. According to an embodiment, the call translation service provider (450) may be included in an external server (e.g., the server (108) of FIG. 1). For example, the call translation service provider (450) may be stored in the memory of the external server. The electronic device (200) may transmit text extracted from an audio signal to the server via a communication circuit (e.g., the communication module (190) of FIG. 1). Thereafter, the electronic device (200) may receive translated text information from the server via the communication circuit. According to one embodiment, the electronic device (200) may transmit the audio signal together with the text extracted from the audio signal to the server in order to obtain the translated text information.

[0105] According to one embodiment, when the call translation service provider (450) is included in an external server, the processor (210) of the electronic device (200) can transmit first voice data to the external server. The external server can obtain a first text of a first language type based on the first voice data. The external server can convert the first text of the first language type into a second text of a second language type. The external server can obtain second voice data of the second language type based on the second text of the second language type. The external server can transmit the second voice data of the second language type to the electronic device (200). The electronic device (200) can obtain the second voice data of the second language type from the external server.

[0106] Figure 5a illustrates a flowchart illustrating the operation of an electronic device according to one embodiment. In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0107] In operation 501, the processor (210) may identify that a call connection is established with an external electronic device through a VoIP application. For example, the VoIP application may be used to provide at least one of a voice call service and a video call service. The VoIP application may be used to establish a VoIP call connection with an external electronic device. The VoIP application may provide at least one of a voice call service and a video call service through the VoIP call connection. For example, the VoIP application may be a third-party application. For example, the VoIP application may be configured to operate based on IP.

[0108] In operation 502, the processor (210) may obtain first voice data of a first language type related to an application related to VoIP. For example, the processor (210) may obtain first voice data of a first language type (e.g., English) while a call connection is established. The processor (210) may obtain the first voice data from an application related to VoIP. The processor (210) may identify the language type of the first voice data as the first language type.

[0109] According to one embodiment, the processor (210) can identify whether voice data is acquired from an application related to VoIP. Based on whether voice data is acquired from the application related to VoIP, the processor (210) can initialize at least one component for a call translation service (e.g., a translation service provider (450) of FIG. 4). As an example, the processor (210) can initialize a voice recognizer (e.g., voice recognizers (441-1, 441-2) of FIG. 4) to identify a language type of first voice data acquired from the application related to VoIP.

[0110] In operation 503, the processor (210) may obtain a second text of a second language type (e.g., Korean) based on a first text of a first language type converted from first speech data of a first language type (e.g., English). For example, the processor (210) may obtain a second text of a second language type that is distinct from the first language type based on a first text of a first language type converted from first speech data of the first language type.

[0111] According to one embodiment, the processor (210) can obtain a first text of a first language type from first voice data of a first language type. The processor (210) can perform voice recognition on the first voice data. The processor (210) can convert the first voice data into a first text using a voice recognizer (e.g., the voice recognizer (441-1, 441-2) of FIG. 4). The processor (210) can divide the first voice data into sentence units. The processor (210) can convert the first voice data divided into sentence units into a first text. The language type of the first voice data and the language type of the first text can be the same as the first language type.

[0112] According to one embodiment, the processor (210) may obtain a second text of a second language type based on a first text of a first language type. For example, the processor (210) may obtain a second text of a second language type by translating the first text of the first language type using a translator (e.g., the translator (442-1, 442-2) of FIG. 4).

[0113] For example, the second language type may be changed according to the setting information of the electronic device (200). The second language type may be set to a language type set in the system of the electronic device (200). For example, the second language type may be changed according to user input.

[0114] In operation 504, the processor (210) may display a visual object including the first text and the second text. For example, the processor (210) may display the visual object including the first text and the second text as an overlay on a user interface for an application related to VoIP through the display (260).

[0115] For example, the processor (210) may display a visual object for providing a call translation service by overlaying a user interface for the VoIP application using an application that is distinct from an application for VoIP. As an example, since the visual object is displayed as overlay on the user interface for the VoIP application, it may be displayed in a semi-transparent state. As an example, the visual object may be movable while overlayed on the user interface. As an example, the processor (210) may display the visual object on an area that is distinct from a main area for a call (e.g., an area representing a user of an external electronic device and / or a user of the electronic device (200)) among the user interfaces for the VoIP application.

[0116] For example, the processor (210) may display the first text within the visual object in response to outputting the first voice data through the speaker (240). After the first voice data is output through the speaker (240), the processor (210) may output the second voice data through the speaker (240). For example, the size at which the first voice data is output may be smaller than the size at which the second voice data is output. The processor (210) may output the first voice data through the speaker (240) based on a first volume. The processor (210) may output the second voice data through the speaker (240) based on a second volume that is greater than the first volume. However, the present invention is not limited thereto. Depending on the embodiment, the second volume may correspond to the first volume.

[0117] The processor (210) may display second text within a visual object in response to outputting second voice data through the speaker (240). The second text corresponding to the second voice data may be displayed after the first text is displayed within the visual object.

[0118] For example, the display format of the first text and the display format of the second text may be distinguished. The processor (210) may set the display format of the first text and the display format of the second text differently in order to distinguish the first text according to the first voice data acquired from the external electronic device and the second text which is a translation of the first text. The processor (210) may set the display format of the first text and the display format of the second text differently by setting at least one of color, boldness, shading, ambient effect, and display position differently.

[0119] According to one embodiment, a visual object including first text and second text may include an area for providing a handwriting input. The processor (210) may identify a handwriting input from a user of the electronic device (200) within the area. The processor (210) may identify a text according to the handwriting input. The processor (210) may transmit voice data obtained from the text according to the handwriting input to an external electronic device. For example, the language type of the text according to the handwriting input may be a second language type. The processor (210) may obtain text of the first language type by translating the text. The processor (210) may transmit voice data of the first language type obtained based on the text of the first language type to the external electronic device.

[0120] According to an embodiment, the processor (210) may display a virtual keyboard based on an input (e.g., a tap input) to an area for providing handwriting input. The processor (210) may obtain text of a second language type using the virtual keyboard. The processor (210) may obtain text of a first language type by translating the text. The processor (210) may transmit voice data of the first language type obtained based on the text of the first language type to an external electronic device. In this case, a user of the electronic device (200) may communicate with a user of the external electronic device using text input.

[0121] According to one embodiment, a visual object for providing a call translation service may be displayed based on an input for displaying the visual object. For example, the processor (210) may identify an input for displaying the visual object while a call connection is established through an application related to VoIP. In response to the input, the processor (210) may display the visual object for providing a call translation service as an overlay on a user interface. According to an embodiment, the processor (210) may, based on the input for displaying the visual object, display a user interface for an application related to VoIP and the visual object in parallel within a display area of ​​the display (260). For example, the processor (210) may display a user interface for an application related to VoIP within a first display area of ​​the display area of ​​the display (260), and display a visual object for providing a call translation service within a second display area of ​​the display area of ​​the display (260). For example, the processor (210) may reduce the size of a user interface for an application related to VoIP based on the input and display a visual object for providing a call translation service together with the user interface.

[0122] According to one embodiment, the processor (210) can obtain third voice data of a second language type through the microphone (250). The processor (210) can obtain third text of a second language type from the third voice data of the second language type. Based on the third text of the second language type, the processor (210) can obtain fourth text of a first language type using a translator. Based on the fourth text of the first language type, the processor (210) can obtain fourth voice data of the first language type. The processor (210) can transmit the fourth voice data to an external electronic device through a call connection related to VoIP.

[0123] Figure 5b illustrates a flowchart illustrating the operation of an electronic device according to one embodiment. In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0124] In operation 511, the processor (210) may obtain first voice data of a first language type received from an external electronic device using a voice data manager (e.g., the voice data manager (451) of FIG. 4) while a call connection regarding VoIP is established with an external electronic device through an application regarding VoIP. For example, the processor (210) may obtain first voice data of a first language type received from an external electronic device through an application regarding VoIP using a voice data manager configured to obtain voice data regarding an application regarding VoIP while a call connection regarding VoIP is established with the external electronic device through an application regarding VoIP.

[0125] According to one embodiment, the processor (210) may identify that a call connection regarding VoIP has been established with an external electronic device through a VoIP application. Using the VoIP application, the processor (210) may transmit voice data acquired through a microphone (250) to the external electronic device and output voice data acquired from the external electronic device through a speaker (240).

[0126] The processor (210) may obtain voice data related to a VoIP application to provide a call translation service. For example, the processor (210) may identify whether voice data is exchanged with an external electronic device through the VoIP application. The processor (210) may obtain the voice data based on the identification that voice data is exchanged with the external electronic device through the VoIP application. According to an embodiment, if only image data that does not include voice data is exchanged with the external electronic device through the VoIP application, the processor (210) may not obtain the voice data. Even if sound data is exchanged with the external electronic device through the VoIP application, the processor (210) may not provide the call translation service if the sound data includes information distinguishable from a human voice. The processor (210) may identify (or detect) voice data related to the VoIP application and provide the call translation service based on the identification (or detection) of the voice data. According to one embodiment, the processor (210) may initialize a component (or module) for providing a call translation service based on the identification (or detection) of voice data. For example, the component for providing a call translation service may include a translation service provider (e.g., the translation service provider (450) of FIG. 4). As an example, the processor (210) may initialize at least one of a text converter, a translator, and / or a voice converter included in the translation service provider based on the identification (or detection) of voice data.

[0127] For example, voice data for an application related to VoIP may include first voice data received from an external electronic device and third voice data acquired through a microphone (250).

[0128] For example, the processor (210) can obtain first voice data and third voice data using a voice data manager (e.g., the voice data manager (451) of FIG. 4). The processor (210) can transmit (or provide) the first voice data and third voice data to a translation service provider (e.g., the translation service provider (450) of FIG. 4) through the voice data manager.

[0129] According to one embodiment, the processor (210) may transmit (or provide) the first voice data and the third voice data to the translation service provider in a separated state without performing mixing of the first voice data and the third voice data. According to another embodiment, when an additional operation, such as call recording, is performed, the processor (210) may transmit (or provide) the first voice data and the third voice data to the translation service provider in a state in which mixing of the first voice data and the third voice data is performed.

[0130] According to one embodiment, the processor (210) can identify the language type of the first voice data. The processor (210) can identify the language type of the first voice data as a first language type (e.g., English). Based on identifying the language type of the first voice data as the first language type, the processor (210) can determine the target language type for the call translation service. Based on identifying the language type according to the system settings as a second language type, the processor (210) can determine the target language type for the call translation service.

[0131] According to an embodiment, the processor (210) may determine a target language type for the call translation service based on identifying the language type of the third voice data as the second language type. According to an embodiment, the target language type for the call translation service may be set based on at least one of acquired voice data, user settings, specified input, and / or system setting information.

[0132] According to one embodiment, the processor (210) may determine the target language types for the call translation service as a first language type and a second language type. For example, the processor (210) may provide a call translation service that provides voice data of a second language type based on voice data of the first language type, and provides voice data of the first language type based on voice data of the second language type.

[0133] In operation 512, the processor (210) may provide first voice data of a first language type to a voice recognizer (e.g., voice recognizers (441-1, 441-2) of FIG. 4) using a voice data manager. The processor (210) may provide first voice data obtained from an application related to VoIP (or third voice data obtained through a microphone (250)) to the voice recognizer using the voice data manager.

[0134] In operation 513, the processor (210) may convert first voice data of a first language type into first text of the first language type using a voice recognizer. The processor (210) may segment the first voice data of the first language type into sentence units using the voice recognizer. The processor (210) may obtain the first text based on the first voice data segmented into sentence units. For example, the processor (210) may convert the first voice data into the first text based on the first language type. For example, if the first voice data is composed of English, the processor (210) may convert the first voice data into the first text based on English.

[0135] The processor (210) can convert third voice data of a second language type acquired through the microphone (250) into third text of a second language type. For example, if the third voice data is in Korean, the processor (210) can convert the third voice data into third text based on Korean.

[0136] In operation 514, the processor (210) may convert a first text of a first language type into a second text of a second language type using a translator (e.g., translators (442-1, 442-2) of FIG. 4). The processor (210) may translate the first text sentence by sentence. Based on the translation of the first text, the processor (210) may obtain a second text.

[0137] According to one embodiment, the processor (210) may use a translator to convert a third text in a second language type into a fourth text in a first language type. The processor (210) may translate the third text sentence by sentence. Based on the translation of the third text, the processor (210) may obtain the fourth text.

[0138] In operation 515, the processor (210) may convert a second text of a second language type into second voice data of a second language type using a voice converter (e.g., voice converters (432-1, 432-2) of FIG. 4). For example, the voice converter may be referred to as TTS.

[0139] According to one embodiment, the processor (210) may use a translator to convert a fourth text of a first language type into fourth speech data of the first language type.

[0140] In operation 516, the processor (210) may output second voice data of a second language type through the speaker (240). In response to outputting the first voice data, the processor (210) may output the second voice data. Depending on the embodiment, the processor (210) may not output the first voice data, but may output only the second voice data.

[0141] According to one embodiment, the processor (210) may display second text of a second language type while outputting second voice data. For example, the processor (210) may display first text corresponding to the first voice data using the display (260) while outputting first voice data using the speaker (240). The processor (210) may display second text corresponding to the second voice data using the display (260) while outputting second voice data using the speaker (240).

[0142] According to an embodiment, the processor (210) may not output the second voice data after the first voice data is output through the speaker (240). Based on the output of the first voice data, the processor (210) may display a first text corresponding to the first voice data and a second text corresponding to the second voice data. Based on identifying an input for the second text, the processor (210) may output the second voice data using the speaker (240).

[0143] According to an embodiment, the processor (210) may transmit second voice data to an external electronic device. The processor (210) may transmit second voice data to the external electronic device to indicate that the first voice data received from the external electronic device has been translated and provided by the electronic device (200).

[0144] According to one embodiment, the processor (210) may display a fourth text of a first language type while outputting fourth voice data. For example, the processor (210) may display a third text corresponding to the third voice data using the display (260) while transmitting the third voice data to an external electronic device. The processor (210) may display a fourth text corresponding to the fourth voice data using the display (260) while outputting the fourth voice data using the speaker (240).

[0145] According to the above-described embodiment, the voice data manager can distinguish between the first voice data received from the external electronic device and the third voice data acquired through the microphone (250). The processor (210) can use the voice data manager to transmit the second voice data acquired based on the first voice data to the kernel audio so that the second voice data can be output through the speaker (240). The processor (210) can use the voice data to transmit the fourth voice data acquired based on the third voice data to the application related to VoIP so that the fourth voice data can be transmitted to the external electronic device.

[0146] FIG. 6A illustrates an example of operation of an electronic device for providing a currency translation service, according to one embodiment.

[0147] Referring to FIG. 6A, the processor (210) of the electronic device (200) can establish a call connection regarding VoIP with an external electronic device (600) through an application regarding VoIP. The processor (210) can display a user interface (601) for the application regarding VoIP through the display (260). For example, the user interface (601) can include an element (610) indicating caller information, an element (611) indicating the duration of the call connection, and / or elements (609) for settings regarding the call connection regarding VoIP (e.g., terminating the call connection, setting the speakerphone, muting). The external electronic device (600) can also display a user interface (650) corresponding to the user interface (601) while the call connection is established through the application regarding VoIP.

[0148] According to one embodiment, the processor (210) may perform a call translation service while a call connection regarding VoIP is established. The processor (210) may display a visual object (602) for the call translation service as an overlay on the user interface (601). For example, the transparency of the visual object (602) may be changed. The processor (210) may display an element (e.g., a slider) for changing the transparency of the visual object (602) together with the visual object (602). The processor (210) may change the transparency of the visual object (602) based on a user input for the element. FIG. 6A illustrates an example in which the visual object (602) is displayed as an overlay on the user interface (601), but is not limited thereto. According to an embodiment, the visual object (602) may be displayed based on a split screen. For example, a screen corresponding to a visual object (602) may be displayed in a first area of ​​the display (260), and a screen corresponding to a user interface (601) may be displayed in a second area of ​​the display (260).

[0149] For example, the visual object (602) may include an object (603) for setting a target language for translation. For example, the object (603) may represent target language types for translation. The object (603) may include an object (604) representing a first language type and an object (605) representing a second language type. The processor (210) may identify a language type of the first voice data received from the external electronic device (600) as the first language type. One of the target language types for translation may be set as the first language type. The processor (210) may identify a language type set in the system of the electronic device (200) as the second language type. The processor (210) may set the remaining one of the target language types for translation as the second language type. Depending on the embodiment, the target language type for translation may be changed based on an input to the object (604) and / or the object (605).

[0150] According to one embodiment, the processor (210) may display text (606) indicating that a call translation service is performed on a visual object (602). The processor (210) may transmit voice data corresponding to the text (606) to an external electronic device (600). The processor (210) may transmit voice data corresponding to the text (606) to the external electronic device (600) to notify that a call translation service is performed on the electronic device (200).

[0151] According to one embodiment, the processor (210) may receive first voice data of a first language type from an external electronic device (600). The processor (210) may convert the first voice data of the first language type into first text (607-1) of the first language type. Based on outputting the first voice data through the speaker (240), the processor (210) may display the first text (607-1) on a visual object (602). The processor (210) may use a translator to convert the first text (607-1) of the first language type into second text (607-2) of the second language type. The processor (210) may display the second text (607-2) on the visual object (602). The processor (210) may output second voice data of a second language type corresponding to the second text (607-2) of the second language type through the speaker (240) based on displaying the second text (607-2). Although not illustrated, the processor (210) may also transmit the second voice data of the second language type to an external electronic device (600).

[0152] The processor (210) can obtain third voice data of a second language type using a microphone (250). The processor (210) can convert the third voice data of the second language type into third text (608-1) of the second language type. Based on obtaining the third voice data, the processor (210) can display the third text (608-1) on a visual object (602). Based on obtaining the third voice data, the processor (210) can transmit the third voice data to an external electronic device (600).

[0153] The processor (210) can use a translator to convert a third text (608-1) of a second language type into a fourth text (608-2) of a first language type. The processor (210) can display the fourth text (608-2) on a visual object (602). The processor (210) can obtain fourth voice data of the first language type corresponding to the fourth text (608-2) of the first language type. The processor (210) can transmit the fourth voice data to an external electronic device (600).

[0154] Although not shown, the visual object (602) may further include an area for providing a handwriting input. The processor (210) may identify a handwriting input within the input. The processor (210) may identify a text of a second language type based on the handwriting input. The processor (210) may display the text of the second language type and the text of the first language type translated from the text on the visual object (602). The processor (210) may obtain voice data corresponding to the text of the first language type. The processor (210) may transmit the obtained voice data to an external electronic device (600).

[0155] According to the above-described embodiment, the processor (210) can translate the first voice data and provide the second text and / or the second voice data to the user of the electronic device (200), regardless of whether the external electronic device (600) provides a call translation function. The processor (210) can provide the third voice data and the fourth voice data translated from the third voice data to the external electronic device (600), regardless of whether the external electronic device (600) provides a call translation function.

[0156] According to an embodiment, both the electronic device (200) and the external electronic device (600) may provide a call translation service. The processor (210) may identify that the call translation service is provided by the external electronic device (600). For example, the processor (210) may identify that the call translation service is provided by the external electronic device (600) based on receiving voice data corresponding to text (606) from the external electronic device (600).

[0157] The processor (210) may not transmit the fourth voice data based on identifying that a call translation service is provided by the external electronic device (600). The electronic device (200) and the external electronic device (600) may not provide translated voice data to each other. For example, the processor (210) of the electronic device (200) may receive only the first voice data from the external electronic device (600). The processor (210) may provide the user of the electronic device (200) with the first text and / or the second voice data obtained based on the first voice data. In addition, the processor (210) may not transmit the fourth voice data translated from the third voice data to the external electronic device (600). Although not illustrated, when a call translation service is provided by the external electronic device (600), the external electronic device (600) may display a visual object for the call translation service as an overlay on the user interface (650).

[0158] FIG. 6b illustrates an example of operation of an electronic device for providing a currency translation service, according to one embodiment.

[0159] Referring to FIG. 6b, the processor (210) of the electronic device (200) can establish a call connection regarding VoIP with an external electronic device (600) through an application regarding VoIP. The call connection regarding VoIP can be established for a video call.

[0160] The processor (210) may display a user interface (621) for an application related to VoIP through the display (260). For example, the user interface (621) may include an area (618) for displaying image data acquired through a camera (270) of the electronic device (200). The user interface (621) may include an object (619) for overlapping the area (618) and representing image data received from an external electronic device (600). According to an embodiment, the processor (210) may display image data received from the external electronic device (600) in the area (618) and display image data acquired through the camera (270) through the object (619).

[0161] The user interface (621) may include elements (620) for setting up a call connection related to VoIP (e.g., changing video effects, terminating a video call connection, muting). An external electronic device (600) may also display a user interface (660) corresponding to the user interface (621).

[0162] According to one embodiment, the processor (210) may perform a call translation service while a call connection regarding VoIP is established. The processor (210) may display a visual object (612) for the call translation service as an overlay on the user interface (621). For example, the visual object (612) may be configured to be movable. The display position of the visual object (612) may be changed based on a user input. For example, the processor (210) may change the display position of the visual object (612) based on a drag (or drag and drop) input to the visual object (612). The visual object (612) may correspond to the visual object (602) of FIG. 6A. According to one embodiment, the visual object (612) may be displayed based on a split screen. For example, a screen corresponding to a visual object (612) may be displayed in a first area of ​​the display (260), and a screen corresponding to a user interface (621) may be displayed in a second area of ​​the display (260). Depending on the embodiment, the visual object (612) may be displayed based on one of a split screen and / or a pop-up window. If the visual object (612) is displayed based on a pop-up window, the visual object (612) may be displayed as an overlay on the user interface (621). If the visual object (612) is displayed based on a split screen, a screen corresponding to the visual object (612) may be displayed in a first area of ​​the display (260), and a screen corresponding to the user interface (621) may be displayed in a second area of ​​the display (260). For example, the manner in which the visual object (612) is displayed (e.g., a split screen or a pop-up window) may change depending on a user input.

[0163] For example, the visual object (612) may include an object (613) for setting a target language for translation. For example, the object (613) may represent target language types for translation. The object (613) may include an object (614) representing a first language type and an object (615) representing a second language type. Depending on the embodiment, the target language type for translation may be changed based on input to the object (614) and / or the object (615).

[0164] According to one embodiment, the processor (210) may display text (616) indicating that a call translation service is performed on a visual object (612). The processor (210) may transmit voice data corresponding to the text (616) to an external electronic device (600) to notify that a call translation service is performed on the electronic device (200).

[0165] According to one embodiment, the processor (210) may receive first voice data of a first language type from an external electronic device (600). For example, the processor (210) may receive the first voice data together with image data.

[0166] The processor (210) can convert first voice data of a first language type into first text (617-1) of a first language type. Based on outputting the first voice data through the speaker (240), the processor (210) can display the first text (617-1) on a visual object (612). The processor (210) can convert the first text (617-1) of the first language type into second text (617-2) of a second language type using a translator. The processor (210) can display the second text (617-2) on the visual object (612). Based on displaying the second text (617-2), the processor (210) can output second voice data of the second language type corresponding to the second text (617-2) of the second language type through the speaker (240). Although not shown, the processor (210) may also transmit second voice data of a second language type to an external electronic device (600).

[0167] The processor (210) can obtain third voice data of a second language type using a microphone (250). The processor (210) can convert the third voice data of the second language type into third text (618-1) of the second language type. Based on obtaining the third voice data, the processor (210) can display the third text (618-1) on a visual object (612). Based on obtaining the third voice data, the processor (210) can transmit the third voice data to an external electronic device (600).

[0168] The processor (210) can use a translator to convert a third text (618-1) of a second language type into a fourth text (618-2) of a first language type. The processor (210) can display the fourth text (618-2) on a visual object (612). The processor (210) can obtain fourth voice data of the first language type corresponding to the fourth text (618-2) of the first language type. The processor (210) can transmit the fourth voice data to an external electronic device (600).

[0169] Although not shown, the visual object (612) may further include an area for providing a handwriting input. The processor (210) may identify a handwriting input within the input. The processor (210) may identify a text of a second language type based on the handwriting input. The processor (210) may display the text of the second language type and the text of the first language type translated from the text on the visual object (612). The processor (210) may obtain voice data corresponding to the text of the first language type. The processor (210) may transmit the obtained voice data to an external electronic device (600).

[0170] FIG. 7 illustrates an example of operation of an electronic device for initiating a currency translation service, according to one embodiment.

[0171] Referring to FIG. 7, in state (710), the processor (210) can establish a call connection with an external electronic device through an application related to VoIP. The processor (210) can display a user interface (711) for the application related to VoIP through the display (260). For example, the user interface (711) can include an element (712) indicating caller information (e.g., name), an element (713) indicating the duration of the call connection, and / or elements (714) for setting a call connection related to VoIP (e.g., terminating a call connection, setting a speakerphone, muting).

[0172] According to one embodiment, while the user interface (711) is displayed, a user input (715) may be identified. For example, the user input (715) may include a swipe input from top to bottom within a display area of ​​the display (260) of the electronic device (200), but is not limited thereto. The user input (715) may include at least one of an input to an object displayed on the display (260), an input to a physical button of the electronic device (200), a drag input, an input in a specified direction, and / or a swipe input from a specified first location to a specified second location. The processor (210) may change the state of the electronic device (200) from state (710) to state (720) in response to the user input (715).

[0173] In state (720), the processor (210) can display an object (725) as an overlay on the user interface (711). The object (725) can represent setting information of the electronic device (200). For example, the object (725) can include an object (721) for activating or deactivating a call translation function, an object (722) for activating or deactivating a microphone mode, an object (723) for activating or deactivating a wireless LAN function, an object (724) for activating or deactivating a Bluetooth function, an area (727) for objects for activating or deactivating specified functions, and / or an area (728) for configuring the display (260).

[0174] According to one embodiment, the processor (210) may identify an input to an object (721) for activating or deactivating a call translation function. Based on the input to the object (721), the processor (210) may activate or deactivate the call translation function. According to an embodiment, the processor (210) may display the object (721) only when an application related to VoIP is running.

[0175] For example, when the currency translation function is activated, the processor (210) may display a visual effect (e.g., displaying a specified color) for the object (721). When the currency translation function is deactivated, the processor (210) may remove the visual effect for the object (721). For example, the processor (210) may control activation or deactivation of the currency translation function based on an input (e.g., a tap input or a swipe input) for the object (721).

[0176] According to one embodiment, the processor (210) may activate a currency translation function based on an input to the object (721). In response to the input to the object (721), the processor (210) may change the state of the electronic device (200) from state (720) to state (730).

[0177] In state (730), the processor (210) may display a visual object (731) as an overlay on a user interface (711) for an application related to VoIP. The visual object (731) may be displayed to provide a call translation service. The visual object (731) may be displayed through an application that is distinct from the application related to VoIP. According to one embodiment, the visual object (731) may correspond to the visual object (602) of FIG. 6A or the visual object (612) of FIG. 6B.

[0178] According to one embodiment, the processor (210) may provide a call translation service from the time an input for the object (721) is identified. According to another embodiment, the processor (210) may also provide a call translation service based on voice data exchanged with an external electronic device from the time a call connection is established, based on the input for the object (721).

[0179] In one embodiment, the processor (210) may disable the currency translation function based on other inputs to the object (721). The processor (210) may remove a visual object (731) displayed as an overlay on the user interface (711) based on other inputs to the object (721).

[0180] FIG. 8 illustrates an example of operation of an electronic device for setting a function for a call translation service within a system setting, according to one embodiment.

[0181] Referring to FIG. 8, the processor (210) may display a screen (810) for system settings of the electronic device (200). Although omitted, the screen (810) for system settings may include a plurality of menus. The screen (810) may include a menu (811) for settings related to a call translation service. In response to identifying an input for the menu (811), the processor (210) may change the screen (810) to a screen (820). The screen (810) of FIG. 8 is merely an example for convenience of explanation, and embodiments of the present disclosure are not limited thereto.

[0182] The processor (210) may, in response to identifying an input to the menu (811), display a screen (820) for specific settings regarding the call translation service.

[0183] For example, the screen (820) may include an object (821) for activating or deactivating a call translation service. The object (821) may include a toggle. The processor (210) may change the activation or deactivation of the call translation service based on an input to the toggle. For example, the screen (820) may include an object (822) for indicating how a visual object for the call translation service is displayed.

[0184] For example, the screen (820) may include a menu (824) for selecting a language type (or target language type) for voice data received from the electronic device (200) and / or a voice for outputting the voice data. The screen (820) may include a menu (825) for selecting a language type (or target language type) for voice data acquired from an external electronic device and / or a voice for outputting the voice data.

[0185] For example, the screen (820) may include a menu (826) that displays a list of applications related to VoIP for providing call translation services. The applications displayed in the menu (826) may be changed. For example, based on the installation and / or deletion of applications related to VoIP, the applications may be added to the menu (826). Based on the addition of the applications, the applications displayed in the menu (826) may be updated. In some embodiments, at least some or all of the applications displayed in the menu (826) may be changed based on user input. The processor (210) may add applications to the applications displayed in the menu (826) or remove at least some or all of the applications displayed in the menu (826) based on the user input.

[0186] FIG. 9 illustrates an example of operation of an electronic device for providing a call translation service while a video call service is provided through an application for VoIP, according to one embodiment.

[0187] Referring to FIG. 9, the processor (210) of the electronic device (200) can establish a VoIP call connection with a plurality of external electronic devices through a VoIP application. The VoIP call connection can be established for a video call.

[0188] The processor (210) can display a user interface (970) for an application related to VoIP through the display (260). For example, the user interface (970) can display image data acquired through the camera (270) of the electronic device (200) in an area (971). The processor (210) can display image data received from a first external electronic device among a plurality of external electronic devices in an area (972). The processor (210) can display image data received from a second external electronic device among a plurality of external electronic devices in an area (973). The processor (210) can display image data received from a third external electronic device among a plurality of external electronic devices in an area (974).

[0189] The user interface (970) may include elements (975) for setting up a call connection regarding VoIP (e.g., changing video effects, terminating a video call connection, muting).

[0190] According to one embodiment, the processor (210) may obtain voice data of a second language type using a microphone (250). The processor (210) may obtain voice data of a second language type from a first external electronic device and a second external electronic device. On the other hand, the processor (210) may obtain voice data of a first language type from a third external electronic device.

[0191] The processor (210) may provide a call translation service for voice data received from a third external electronic device based on obtaining voice data of a first language type from the third external electronic device. The processor (210) may display a visual object (976) for the call translation service in an overlapping manner on an area (974) for displaying image data received from the third external electronic device. Depending on the embodiment, the location where the visual object (976) is displayed may be changed. The processor (210) may not provide a translation service for voice data obtained from the first external electronic device and the second external electronic device. The processor (210) may only provide a translation service for voice data obtained from the third external electronic device.

[0192] For example, the processor (210) may obtain first voice data of a first language type from a third external electronic device. The processor (210) may obtain text of a first language type corresponding to the first voice data of the first language type. Based on the text of the first language type, the processor (210) may obtain text of a second language type using a translator. The processor (210) may display the text of the first language type and the text of the second language type on a visual object (976).

[0193] According to an embodiment, the processor (210) may not transmit voice data corresponding to text of the second language type to a plurality of external electronic devices. The processor (210) may not transmit voice data corresponding to text of the second language type to a plurality of external electronic devices based on the establishment of a call connection with the plurality of external electronic devices.

[0194] Although FIG. 9 illustrates an example of providing a translation service for voice data acquired from a third external electronic device, the present invention is not limited thereto. According to an embodiment, the electronic device (200) may provide translation services for three or more language types.

[0195] For example, voice data of a second language type (e.g., Spanish) may be acquired from a first external electronic device, voice data of a third language type (e.g., Japanese) may be acquired from a second external electronic device, and voice data of a fourth language type (e.g., French) may be acquired from a third external electronic device. The processor (210) may acquire and output voice data of a first language type (e.g., English) based on the voice data of the second language type acquired from the first external electronic device. The processor (210) may acquire and output voice data of a first language type (e.g., English) based on the voice data of the third language type acquired from the third external electronic device.

[0196] FIG. 10 illustrates an example of operation of an electronic device for providing a call translation service while a video call service is provided through an application for VoIP, according to one embodiment.

[0197] Referring to FIG. 10, in state (1010), the electronic device (200) can operate in portrait mode. In state (1020), the electronic device (200) can operate in landscape mode. The processor (210) can use a gyro sensor included in the electronic device (200) to identify the display mode of the electronic device (200) as one of portrait mode and landscape mode.

[0198] In states (1010) and (1020), the processor (210) can establish a call connection with an external electronic device through an application related to VoIP. The processor (210) can display a user interface (1050, 1060) for the application related to VoIP through the display (260).

[0199] In state (1010), the processor (210) may display a user interface (1050) for an application related to VoIP. The user interface (1050) may include an area (1052) for displaying image data acquired through a camera (270) of the electronic device (200). The user interface (1050) may be displayed as an overlay on the area (1052) and may include an object (1051) for representing image data received from an external electronic device. According to an embodiment, the processor (210) may display image data received from the external electronic device in the area (1052) and display image data acquired through the camera (270) through the object (1051). For example, the user interface (1050) may correspond to the user interface (601) of FIG. 6B .

[0200] According to one embodiment, the processor (210) can identify that the display mode of the electronic device (200) has changed from a portrait mode to a landscape mode. Based on identifying that the display mode of the electronic device (200) has changed from a portrait mode to a landscape mode, the processor (210) can change the state of the electronic device (200) from a state (1010) to a state (1020).

[0201] In state (1020), the processor (210) can display a user interface (1060) for an application related to VoIP in a first display area (1021) among the display areas of the display (260). The processor (210) can display a visual object (1063) for a call translation service in a second display area (1022) among the display areas of the display (260). The processor (210) can display the user interface (1060) and the visual object (1063) separately while the display mode is a landscape mode.

[0202] For example, the user interface (1060) may include an area (1062) for displaying image data acquired through the camera (270) of the electronic device (200). The user interface (1060) may be displayed as an overlay on the area (1062) and may include an object (1061) for representing image data received from an external electronic device. According to an embodiment, the processor (210) may display image data received from the external electronic device in the area (1062) and display image data acquired through the camera (270) through the object (1061). For example, the user interface (1060) may correspond to the user interface (601) of FIG. 6B.

[0203] For example, a visual object (1063) may be configured to provide a currency translation service. The visual object (1063) may include information for providing a currency translation service.

[0204] For example, the visual object (1063) may include an area (1064) for providing a handwriting input. The processor (210) may identify a handwriting input through an external electronic device (e.g., a stylus or a digital pen) or a part of the user's body (e.g., a finger) in the area (1064). The processor (210) may identify text according to the handwriting input and transmit voice data obtained from the text according to the handwriting input to the external electronic device. The visual object (1053) displayed in portrait mode may not include an area for providing a handwriting input. The visual object (1063) displayed in landscape mode may include an area (1064) for providing a handwriting input. Depending on the embodiment, the visual object (1053) may also include an area for providing a handwriting input.

[0205] FIGS. 11A and 11B illustrate an example of a positional relationship between a first housing and a second housing in an unfolded state and a folded state of an electronic device according to one embodiment.

[0206] FIG. 11c illustrates an example of a screen for a currency translation service displayed based on the positional relationship between the first housing and the second housing, according to one embodiment.

[0207] Referring to FIGS. 11A and 11B, the electronic device (200) may be an example of the electronic device (101) of FIG. 1. For example, the display (260) of the electronic device (200) may include a first display (260-1) and a second display (260-2).

[0208] A first housing (1110), a second housing (1120), and a folding housing (1165) may be included in an electronic device (200). The electronic device (200) may include a first housing (1110) including a first side (1111) and a second side (1112) opposite the first side. The electronic device (200) may include a second housing (1120) including a third side (1121) and a fourth side (not shown) opposite the third side (1121). The electronic device (200) may include a folding housing (1165) that pivotally connects the first housing (1110) and the second housing (1120) about a folding axis (1137). At least a portion of the first display (260-1) may be disposed on one side of the first housing (1110) (e.g., the first side (1111)) and one side of the second housing (1120) (e.g., the third side (1121)). For example, at least a portion of the first display (260-1) may be disposed on the first side (1111) and the third side (1121) across the folding housing (1165). The first display area (1131), the second display area (1132), and the third display area (1133) may be included in the first display (260-1). The folding housing (1165) may include a hinge structure. The second display (260-2) may be disposed on the second side (1112). The electronic device (200) may include a camera oriented in the direction in which the second side (1112) faces. The above camera may be placed within a portion area (1150) of the second surface (1112).

[0209] The above-described housings (e.g., the first housing (1110), the second housing (1120), and the folding housing (1165)) may be referred to as housing parts. For example, the first housing (1110) may be referred to as a first housing part. For example, the second housing (1120) may be referred to as a second housing part. For example, the folding housing (1165) may be referred to as a folding housing part.

[0210] According to one embodiment, the electronic device (200) can provide an unfolding state in which the first housing (1110) and the second housing (1120) are fully folded out by the folding housing (1165). For example, referring to FIG. 11A, the electronic device (200) can be in a state (1100) which is the unfolding state. For example, the state (1100) can mean a state in which a first direction (1191) toward which the first surface (1111) faces corresponds to a second direction (1192) toward which the third surface (1121) faces. For example, within the state (1100), the first direction (1191) can be substantially parallel to the second direction (1192). For example, within state (1100), the first direction (1191) may be substantially identical to the second direction (1192).

[0211] According to one embodiment, within the state (1100), the first surface (1111) may form substantially one plane with the third surface (1121). For example, the angle (1105-1) between the first surface (1111) and the third surface (1121) within the state (1100) may be approximately 180 degrees. For example, the state (1100) may refer to a state in which the entire display area of ​​the first display (260-1) may be provided substantially on one plane. For example, the state (1100) may refer to a state in which the first display area (1131), the second display area (1132), and the third display area (1133) may all be provided on one plane. For example, within the state (1100), the third display area (1133) may not include a curved surface. For example, the unfolded state may be referred to as an outspread state or outspreading state. Different states of the electronic device (200) based on angles (1105-2, 1105-3, 1105-4) are described below.

[0212] Referring to FIG. 11B, an electronic device (200) according to one embodiment can provide a folding state in which a first housing (1110) and a second housing (1120) are folded in by a folding housing (1165). For example, the electronic device (200) can be in the folding state including a state (1101), a state (1102), and a state (1103). For example, the folding state including a state (1101), a state (1102), and a state (1103) can mean a state in which a first direction (1191) toward which a first surface (1111) faces is distinct from a second direction (1192) toward which a third surface (1121) faces. For example, in state (1101), the angle between the first direction (1191) and the second direction (1192) is approximately 45 degrees, and the first direction (1191) and the second direction (1192) can be distinguished from each other. For example, in state (1102), the angle between the first direction (1191) and the second direction (1192) is approximately 150 degrees, and the first direction (1191) and the second direction (1192) can be distinguished from each other. For example, in state (1103), the angle between the first direction (1191) and the second direction (1192) is substantially 180 degrees, and the first direction (1191) and the second direction (1192) can be distinguished from each other.

[0213] In one embodiment, the angle between the first side (1111) and the third side (1121) within the folded state may be greater than or equal to about 0 degrees and less than 180 degrees. For example, in state (1101), the angle (1105-2) between the first side (1111) and the third side (1121) may be about 135 degrees. In state (1102), the angle (1105-3) between the first side (1111) and the third side (1121) may be about 30 degrees. In state (1103), the angle (1105-4) between the first side (1111) and the third side (1121) may be substantially 0 degrees. For example, the folded state may be referred to as a folded state.

[0214] In one embodiment, the folded state may include a plurality of sub-folding states, unlike the unfolded state. For example, referring to FIG. 11B, the folded state may include a plurality of sub-folding states, including a state (1103) that is a fully folded state in which the first face (1111) substantially overlaps the third face (1121) by rotation provided through the folding housing (1165), and a state (1101) and a state (1102) that are intermediate folding states between the state (1103) and the unfolded state (e.g., the state (1100) of FIG. 11A). For example, the electronic device (200) can provide a state (1103) in which the entire area of ​​the first display area (1131) is substantially completely overlapped with the entire area of ​​the second display area (1132) by causing the first side (1111) and the third side (1121) to face each other by the folding housing (1165). For example, the electronic device (200) can provide a state (1103) in which the first direction (1191) is substantially opposite to the second direction (1192). For example, the state (1103) may mean a state in which the first display (260-1) is covered within the field of view of a user looking at the electronic device (200). However, the present invention is not limited thereto.

[0215] According to one embodiment, the first display (260-1) can be bent by the rotation provided through the folding housing (1165). For example, in the first display (260-1), unlike the first display area (1131) and the second display area (1132), the third display area (1133) can be bent according to the folding operation. For example, the third display area (1133) can be in a bent state to prevent damage to the first display (260-1) within the fully folded state. In the fully folded state, unlike the third display area (1133) being bent, the entire first display area (1131) can be completely overlapped over the entire second display area (1132).

[0216] Referring to FIGS. 11A and 11B , an example is illustrated in which the first display (260-1) of the electronic device (200) includes one folding display area (e.g., the third display area (1133)) or the electronic device (200) includes one folding housing (e.g., the folding housing (1165)), but this is for convenience of explanation. According to embodiments, the first display (260-1) of the electronic device (200) may include a plurality of folding display areas. For example, the first display (260-1) of the electronic device (200) may include two or more folding display areas, and the electronic device (200) may include two or more folding housings for providing the two or more folding areas, respectively.

[0217] Referring to FIG. 11c, a currency translation service may be provided in states (1160), (1170), and (1180). For the description of FIG. 11c, reference numerals of FIGS. 11a and 11b may be used.

[0218] For example, state (1160) may correspond to state (1100) of FIGS. 11A and 11B. For example, in state (1160), a first direction (1191) toward which a first face (1111) faces may correspond to a second direction (1192) toward which a third face (1121) faces.

[0219] For example, state (1170) may correspond to state (1101) of FIGS. 11A and 11B. For example, in state (1170), the angle between the first direction (1191) toward which the first face (1111) faces and the second direction (1192) toward which the third face (1121) faces may be within a specified angle.

[0220] For example, state (1180) may correspond to state (1103) of FIGS. 11A and 11B. For example, in state (1180), the first direction (1191) toward which the first face (1111) faces and the second direction (1192) toward which the third face (1121) faces may be opposite to each other.

[0221] In state (1160), the processor (210) may display a user interface (1161) for an application related to VoIP using the first display (260-1). The processor (210) may display a visual object (1162) for providing a call translation service as an overlay on the user interface (1161). The screen displayed in state (1160) may correspond to the screen displayed on the electronic device (200) of FIG. 6B.

[0222] In state (1170), the processor (210) can display a user interface (1171) for an application related to VoIP on a first display area (1131) of the first display (260-1) (or at least a portion of the first display area (1131) and the third display area (1133). The processor (210) can display a visual object (1172) for providing a call translation service on a second display area (1132) of the first display (260-1) (or at least a portion of the second display area (1132) and the third display area (1133). For example, the visual object (1172) can further include an area (1173) for providing a handwriting input.

[0223] In state (1180), the processor (210) may display a user interface (1181) for an application related to VoIP using the second display (260-2). The processor (210) may display a visual object (1182) for providing a call translation service by overlaying the user interface (1181). The screen displayed in state (1180) may be a screen that is transformed from the screen displayed in state (1160) according to the size of the second display (260-2).

[0224] According to one embodiment, an electronic device (e.g., electronic device (200)) may include a display (e.g., display (260)), a speaker (e.g., speaker (240)), a microphone (e.g., microphone (250)), communication circuitry (e.g., communication circuitry (230)), at least one processor (e.g., processor (210)) including processing circuitry, and a memory (e.g., memory (220)) including one or more storage media storing instructions. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify that a call connection is established with an external electronic device via an application relating to voice over internet protocol (VoIP). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain first voice data of a first language type related to the application while the call connection is established. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a second text of a second language type, distinct from the first language type, based on the first text of the first language type converted from the first speech data of the first language type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, through the display, a visual object including the first text and the second text as an overlay on a user interface for the application.

[0225] According to one embodiment, the first voice data of the first language type can be received from the external electronic device.

[0226] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain second speech data of the second language type based on the second text of the second language type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to output the second speech data through the speaker after the first speech data has been output.

[0227] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain third voice data of the second language type through the microphone. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain fourth text of the first language type based on third text of the second language type converted from the third voice data of the second language type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit fourth voice data of the first language type, obtained based on the fourth text of the first language type, to the external electronic device through the call connection.

[0228] According to one embodiment, the third text and the fourth text may be included in the visual object.

[0229] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the first voice data from the application through a voice data manager configured to obtain voice data relating to the application while the call connection is established.

[0230] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to provide the first speech data to a translator for changing a language type, the translator being connected to the speech data manager. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the second text of the second language type, based on the first speech data of the first language type, using the translator.

[0231] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify an input for displaying the visual object while the user interface is displayed. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the visual object as an overlay on the user interface based on the input for displaying the visual object.

[0232] According to one embodiment, the electronic device may include a first housing, a second housing, and a hinge structure rotatably connecting the first housing to the second housing with respect to a folding axis. The display may include a first display area corresponding to one side of the first housing and a second display area corresponding to one side of the second housing, the first display area being divided with respect to the folding axis. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the visual object as an overlay on the user interface displayed in the first display area and the second display area, in a first state in which a first direction toward which the first display area faces corresponds to a second direction toward which the second display area faces.

[0233] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the user interface in the first display area and to display the visual object in the second display area of ​​the display based on a state of the electronic device that has changed from the first state to a second state in which an angle between the first direction and the second direction is within a specified angle.

[0234] In one embodiment, the electronic device may include another display disposed on the other side of the first housing. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the visual object as an overlay on the user interface through the other display based on a state of the electronic device changed from the first state to a third state in which the first direction and the second direction are opposite to each other.

[0235] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a language type of the first speech data as a first language type based on the first speech data.

[0236] According to one embodiment, the visual object may include an area for providing handwriting input.

[0237] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify text according to the handwriting input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit voice data obtained from the text according to the handwriting input to the external electronic device via the call connection.

[0238] According to one embodiment, the call connection may be established for at least one of a video call and a voice call.

[0239] According to one embodiment, a method performed in an electronic device may include an operation of identifying that a call connection is established with an external electronic device through an application relating to voice over internet protocol (VoIP). The method may include an operation of obtaining, while the call connection is established, first voice data of a first language type related to the application. The method may include an operation of obtaining, based on first text of the first language type, a second text of a second language type that is distinct from the first language type, and which is converted from the first voice data of the first language type. The method may include an operation of displaying, through a display of the electronic device, a visual object including the first text and the second text as an overlay on a user interface for the application.

[0240] According to one embodiment, the first voice data of the first language type can be received from the external electronic device.

[0241] According to one embodiment, the method may include an operation of obtaining third voice data of the second language type through a microphone of the electronic device. The method may include an operation of obtaining a fourth text of the first language type based on a third text of the second language type converted from the third voice data of the second language type. The method may include an operation of transmitting the fourth voice data of the first language type obtained based on the fourth text of the first language type to the external electronic device through the call connection.

[0242] According to one embodiment, the third text and the fourth text may be included in the visual object.

[0243] According to one embodiment, a non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by a processor of an electronic device having a display, a speaker, a microphone, and a communication circuit, cause the electronic device to identify that a call connection is established with an external electronic device through an application relating to voice over internet protocol (VoIP). The one or more programs may include instructions that, when executed by the processor, cause the electronic device to obtain first voice data of a first language type related to the application while the call connection is established. The one or more programs may include instructions that, when executed by the processor, cause the electronic device to obtain second text of a second language type distinct from the first language type based on first text of the first language type converted from the first voice data of the first language type. The one or more programs may include instructions that, when executed by the processor, cause the electronic device to display, through the display, a visual object including the first text and the second text as an overlay on a user interface for the application.

[0244] According to one embodiment, an electronic device may include at least one processor including a speaker, a microphone, a communication circuit, a processing circuit, and a memory including one or more storage media storing instructions. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain, from an application relating to voice over internet protocol (VoIP), first voice data of a first language type received from an external electronic device, using a voice data manager configured to obtain voice data relating to the application while a call connection with the external electronic device is established through the application. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to provide, using the voice data manager, the first voice data of the first language type to a voice recognizer. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to convert, using the speech recognizer, the first speech data of the first language type into first text of the first language type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to convert, using a translator, the first text of the first language type into second text data of a second language type distinct from the first language type. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to convert, using a speech converter, the second text of the second language type into second speech data of the second language type.The above instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to output the second speech data of the second language type through the speaker.

[0245] Electronic devices according to embodiments disclosed herein may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to embodiments disclosed herein are not limited to the aforementioned devices.

[0246] The embodiments of this document and the terminology used herein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among the phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0247] In one embodiment of this document, the term "module" used may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0248] One embodiment of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0249] According to one embodiment, the method according to one embodiment disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0250] According to one embodiment, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to one embodiment, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to one embodiment, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In electronic devices, display; speaker; microphone; communication circuit; At least one processor comprising a processing circuit; and A memory comprising one or more storage media for storing instructions, The above instructions, when individually or collectively executed by the at least one processor, Identifies that a call connection is established with an external electronic device through an application for voice over internet protocol (VoIP), While the above call connection is established, first voice data of a first language type related to the above application is acquired, Obtaining a second text of a second language type that is distinct from the first language type based on the first text of the first language type converted from the first speech data of the first language type, Causing the electronic device to display, through the display, a visual object including the first text and the second text as an overlay on a user interface for the application; Electronic devices.

2. In the first paragraph, the first voice data of the first language type is, Received from the above external electronic device, Electronic devices.

3. In the second paragraph, when the instructions are individually or collectively executed by the at least one processor, Based on the second text of the second language type, second speech data of the second language type is obtained, Further causing the electronic device to output the second voice data after the first voice data is output through the speaker. Electronic devices.

4. In the second paragraph, when the instructions are individually or collectively executed by the at least one processor, Through the above microphone, third voice data of the second language type is acquired, Obtaining a fourth text of the first language type based on the third text of the second language type converted from the third speech data of the second language type, Further causing the electronic device to transmit, to the external electronic device, fourth voice data of the first language type obtained based on the fourth text of the first language type through the call connection. Electronic devices.

5. In the fourth paragraph, the third text and the fourth text are, Included in the above visual object, Electronic devices.

6. In the first paragraph, when the instructions are individually or collectively executed by the at least one processor, Further causing the electronic device to obtain the first voice data from the application, through a voice data manager configured to obtain voice data regarding the application while the call connection is established. Electronic devices.

7. In the sixth paragraph, when the instructions are individually or collectively executed by the at least one processor, Connected to the above voice data manager, and providing the first voice data to a translator for changing the language type, Further causing the electronic device to obtain the second text of the second language type using the translator based on the first speech data of the first language type. Electronic devices.

8. In the first paragraph, when the instructions are individually or collectively executed by the at least one processor, While the above user interface is displayed, identify an input for displaying the above visual object, Further causing the electronic device to display the visual object as an overlay on the user interface based on the input for displaying the visual object. Electronic devices.

9. In the first paragraph, the electronic device, 1st housing; Second housing; and Further comprising a hinge structure that rotatably connects the first housing to the second housing based on the folding axis, The above display is, A first display area corresponding to one side of the first housing and a second display area corresponding to one side of the second housing are included based on the folding axis, The above instructions, when individually or collectively executed by the at least one processor, Further causing the electronic device to display the visual object in an overlapping manner on the user interface displayed in the first display area and the second display area, within a first state in which a first direction toward which the first display area faces corresponds to a second direction toward which the second display area faces. Electronic devices.

10. In the 9th paragraph, when the instructions are individually or collectively executed by the at least one processor, Further causing the electronic device to display the user interface in the first display area and display the visual object in the second display area of ​​the display based on a state of the electronic device changed from the first state to a second state in which the angle between the first direction and the second direction is within a specified angle. Electronic devices.

11. In the 10th paragraph, the electronic device, Further comprising another display arranged on the other side of the first housing, The above instructions, when individually or collectively executed by the at least one processor, Further causing the electronic device to display the visual object as an overlay on the user interface through the other display based on the state of the electronic device changed from the first state to a third state in which the first direction and the second direction are opposite to each other. Electronic devices.

12. In the first paragraph, when the instructions are individually or collectively executed by the at least one processor, Further causing the electronic device to identify the language type of the first voice data as the first language type based on the first voice data. Electronic devices.

13. In the first paragraph, the visual object is, Including an area for providing handwriting input, Electronic devices.

14. In a method performed in an electronic device, An action that identifies that a call connection has been established with an external electronic device via an application relating to voice over internet protocol (VoIP); An operation of obtaining first voice data of a first language type related to the application while the call connection is established; An operation of obtaining a second text of a second language type that is distinct from the first language type, based on the first text of the first language type converted from the first speech data of the first language type; and An action of displaying a visual object including the first text and the second text as an overlay on a user interface for the application through a display of the electronic device, method.

15. In a non-transitory computer-readable storage medium storing one or more programs, the one or more programs, when executed by a processor of an electronic device having a display, a speaker, a microphone, and a communication circuit, Identifies that a call connection is established with an external electronic device through an application for voice over internet protocol (VoIP), While the above call connection is established, first voice data of a first language type related to the above application is acquired, Obtaining a second text of a second language type that is distinct from the first language type based on the first text of the first language type converted from the first speech data of the first language type, Instructions for causing the electronic device to display, through the display, a visual object including the first text and the second text as an overlay on a user interface for the application; Non-transitory computer-readable storage medium.

Citation Information

Patent Citations

  • Communication system and communication method thereof

    JP2020120356A

  • Two-way speech translation system, two-way speech translation method and program

    JP2023022150A

  • Saturated air pressure bucket

    KR1020210142297A

  • Manufacturing method of light emitting element

    KR102039090B1

  • Handle device for folding door

    KR102683191B1