Electronic device and method for providing translation function by using same
The electronic device addresses the challenge of real-time voice translation by learning speaker recognition models, facilitating immediate communication through real-time translation.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-10-28
- Publication Date
- 2026-06-04
AI Technical Summary
Existing voice translation technologies fail to provide real-time translation data, hindering immediate communication between multiple speakers.
An electronic device learns a speaker recognition model for voice signals, enabling real-time translation by recognizing speakers and providing translated data when user interaction is detected.
Enables immediate communication between multiple speakers by providing real-time translation using speaker recognition and translation models.
Smart Images

Figure KR2025017334_04062026_PF_FP_ABST
Abstract
Description
Electronic device and method of providing translation function using the same
[0001] The embodiments of the present disclosure relate to an electronic device and a method for providing a translation function using the same.
[0002] Recently, voice translation functions through translation applications are being widely used. For example, users of electronic devices can use voice translation functions when conversing with speakers of different languages. For instance, an electronic device can collect the entire voice signal received through a microphone via the voice translation function, perform speaker separation, and provide translation data (e.g., text or sound) corresponding to each speaker's voice signal.
[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.
[0004] However, since translation data is provided after the entire voice signal received through the microphone is collected, translation data may not be available in real time. Consequently, immediate communication between multiple speakers may not be possible.
[0005] An electronic device according to an embodiment of the present disclosure can learn a speaker recognition model for the characteristics of a speaker and language corresponding to a voice signal received through a microphone when user interaction is detected. When the learning of the speaker recognition model of the speaker is completed, the electronic device can automatically recognize the speaker corresponding to the voice signal received through the microphone using the learned speaker recognition model of the speaker and provide translated data corresponding to the voice signal to the user in real time.
[0006] According to one embodiment of the present disclosure, an electronic device may include a microphone, a display, at least one processor including processing circuitry, and a memory for storing instructions. According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may display a user interface on the display that provides translation functions for a plurality of speakers. According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may learn a speaker recognition model for the characteristics and language of the speaker corresponding to the first voice signal based on the first voice signal when a first voice signal is received through the microphone and a first user interaction is detected. According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may check whether a second voice signal is received through the microphone and a second user interaction is detected. According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may check whether there exists a trained speaker recognition model corresponding to the second voice signal if the second user interaction is not detected. According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may, if there exists a trained speaker recognition model corresponding to the second voice signal, translate the second voice signal into a target language based on the trained speaker recognition model and display it on the display.
[0007] According to one embodiment of the present disclosure, a method for providing a translation function may include an operation of displaying a user interface that provides translation functions for a plurality of speakers on a display. According to one embodiment, a method for providing a translation function may include an operation of receiving a first voice signal through a microphone and, when a first user interaction is detected, learning a speaker recognition model for the characteristics of a speaker and language corresponding to the first voice signal based on the first voice signal. According to one embodiment, a method for providing a translation function may include an operation of receiving a second voice signal through the microphone and checking whether a second user interaction is detected. According to one embodiment, if the second user interaction is not detected, a method for providing a translation function may include an operation of checking whether a speaker recognition model that has completed learning corresponding to the second voice signal exists. According to one embodiment, if a speaker recognition model that has completed learning corresponding to the second voice signal exists, a method for providing a translation function may include an operation of translating the second voice signal into a target language based on the speaker recognition model that has completed learning and displaying it on the display.
[0008] According to one embodiment of the present disclosure, a non-transient computer-readable storage medium (or computer program product) storing one or more programs may be described. One or more programs according to one embodiment may include instructions for displaying a user interface on a display that provides translation functions for a plurality of speakers when executed by at least one processor of an electronic device. One or more programs according to one embodiment may include instructions for learning a speaker recognition model for the characteristics of a speaker and language corresponding to the first voice signal based on the first voice signal when a first voice signal is received through a microphone and a first user interaction is detected when executed by at least one processor of an electronic device. One or more programs according to one embodiment may include instructions for checking whether a second voice signal is received through the microphone and a second user interaction is detected when executed by at least one processor of an electronic device. One or more programs according to one embodiment may include a command to check whether there exists a speaker recognition model that has completed training corresponding to the second voice signal when executed by at least one processor of an electronic device and the second user interaction is not detected. One or more programs according to one embodiment may include a command to translate the second voice signal into a target language and display it on the display based on the speaker recognition model that has completed training if there exists a speaker recognition model that has completed training corresponding to the second voice signal when executed by at least one processor of an electronic device.
[0009] An electronic device according to an embodiment of the present disclosure can, when user interaction is detected, learn a speaker recognition model corresponding to a voice signal received through a microphone, and use the learned speaker recognition model to provide speaker recognition and translated data corresponding to the voice signal in real time. Accordingly, it can support immediate communication between multiple speakers.
[0010] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment of the present disclosure.
[0011] FIG. 2 is a block diagram illustrating an electronic device according to one embodiment of the present disclosure.
[0012] FIG. 3 is a flowchart illustrating a method for providing a translation function according to one embodiment of the present disclosure.
[0013] FIGS. 4a and FIGS. 4b are flowcharts illustrating a method for providing a translation function according to one embodiment of the present disclosure.
[0014] FIG. 5 is a drawing for illustrating a user interface that provides a translation function for a plurality of speakers according to one embodiment of the present disclosure.
[0015] FIG. 6 is a drawing for explaining a method of providing a translation function according to one embodiment of the present disclosure.
[0016] FIG. 7 is a drawing for explaining a method of providing a translation function according to one embodiment of the present disclosure.
[0017] FIGS. 8 to 12 are drawings for explaining a method of providing a translation function according to one embodiment of the present disclosure.
[0018] FIG. 13 is a flowchart illustrating a method for generating a summary according to one embodiment of the present disclosure.
[0019] FIG. 14 is a flowchart illustrating a method for providing a translation function according to one embodiment of the present disclosure.
[0020] FIG. 15 is a drawing for explaining a method of providing a translation function according to one embodiment of the present disclosure.
[0021] FIG. 16 is a flowchart illustrating a method for providing a translation function according to one embodiment of the present disclosure.
[0022] FIG. 17 is a drawing for explaining a method of providing a translation function according to one embodiment of the present disclosure.
[0023] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.
[0024] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to one embodiment of the present disclosure.
[0025] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0026] The processor (120) can control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., a program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0027] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0028] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0029] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0030] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0031] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0032] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0033] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0034] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0035] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0036] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0037] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0038] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0039] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0040] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0041] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0042] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) may support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.
[0043] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a printed circuit board (PCB)). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0044] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0045] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0046] According to one embodiment, commands or data may be transmitted or received between an electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0047] FIG. 2 is a block diagram illustrating an electronic device (101) according to one embodiment of the present disclosure.
[0048] Referring to FIG. 2, an electronic device (e.g., electronic device (101) of FIG. 1) may include a communication circuit (210) (e.g., communication module (190) of FIG. 1), a memory (220) (e.g., memory (130) of FIG. 1), a microphone (230) (e.g., input module (150) of FIG. 1), a display (240) (e.g., display module (160) of FIG. 1), and / or a processor (250) (e.g., processor (120) of FIG. 1).
[0049] According to one embodiment of the present disclosure, a communication circuit (210) (e.g., communication module (190) of FIG. 1) can control a communication connection between an electronic device (101) and at least one external electronic device (e.g., electronic device (102), electronic device (104) of FIG. 1) (and / or a server (e.g., server (108) of FIG. 1)) under the control of a processor (250).
[0050] According to one embodiment of the present disclosure, a memory (220) (e.g., memory (130) of FIG. 1) performs the function of storing a program (e.g., program (140) of FIG. 1) for processing and controlling the processor (250) of the electronic device (101), an operating system (OS) (e.g., operating system (142) of FIG. 1), various applications, and / or input / output data, and can store a program that controls the overall operation of the electronic device (101). The memory (220) can store various configuration information required for processing functions related to various embodiments of the present disclosure in the electronic device (101). The memory (220) can store executable instructions. For example, the memory (220) can store instructions that cause the electronic device (101) to perform operations when executed by the processor (250). For example, instructions may be stored on a computer-readable recording medium. The recording medium may be tangible and non-transitory. The memory (220) and / or the recording medium may store one or more programs containing instructions.
[0051] In one embodiment, the memory (220) can store instructions for learning a speaker recognition model of the speaker corresponding to the voice signal when a voice signal is received through the microphone (230) and user interaction is detected. The memory (220) can store conditions under which the learning of the speaker recognition model of the speaker is completed.
[0052] In one embodiment, the memory (220) may store instructions for checking whether there exists a trained speaker recognition model corresponding to the voice signal when a voice signal is received through the microphone (230) and no user interaction is detected. If there exists a trained speaker recognition model corresponding to the voice signal, the memory (220) may store instructions for translating the voice signal into a target language based on the trained speaker recognition model corresponding to the voice signal.
[0053] In one embodiment, the memory (220) may store a mapping of the translated text and speaker information (e.g., speaker's name information and / or language information used by the speaker). Based on the translated text and speaker information, the memory (220) may store instructions for generating a summary related to a conversation (e.g., meeting content) of multiple speakers.
[0054] According to one embodiment of the present disclosure, a microphone (230) (e.g., input module (150) of FIG. 1) can receive a voice signal spoken by at least one of a plurality of speakers. In one embodiment, the microphone (230) may include a plurality of microphones.
[0055] According to one embodiment of the present disclosure, the display (240) displays an image under the control of a processor (250) and may be implemented as any one of a liquid crystal display (LCD), a light-emitting diode (LED) display, a μLED (micro LED) display, an organic light-emitting diode (OLED) display, an active matrix organic light-emitting diode (AMOLED) display, a micro electro mechanical systems (MEMS) display, an electronic paper display, a flexible display, a foldable display, or a rollable display. However, it is not limited thereto.
[0056] In one embodiment, the display (240) may display a user interface that provides translation functions for multiple speakers under the control of the processor (250). The user interface may include multiple objects representing multiple speakers. Under the control of the processor (250), the display (240) may display text in which the speaker's information and voice signal are translated into a target language.
[0057] In one embodiment, the display (240) can distinguishly display, under the control of the processor (250), at least one object representing at least one speaker whose speaker recognition model has been trained among a plurality of objects included in the user interface, and at least one other object representing at least one other speaker whose speaker recognition model has not been trained among a plurality of speakers.
[0058] According to one embodiment of the present disclosure, the processor (250) may include, for example, a microcontroller unit (MCU) and may control a plurality of hardware components connected to the processor (250) by running an operating system (OS) or an embedded software program. The processor (250) may control a plurality of hardware components according to, for example, instructions stored in memory (220) (e.g., program (140) of FIG. 1).
[0059] In one embodiment, the processor (250) may include at least one component (or module) for operation according to an embodiment of the present disclosure. For example, the processor (250) may include at least one functional part such as a speaker recognition model learning module (251), a speaker recognition module (252), a language recognition module (253), a speech recognition module (254), a language translation module (255), and / or a summary generation module (256). According to one embodiment, at least some of the functional parts may be included in the processor (250) as a hardware module (e.g., circuitry) and / or implemented as software including one or more instructions that can be executed by the processor (250). For example, operations performed by the processor (250) may be stored in memory (220) and, at execution, executed by instructions that cause the processor (250) to operate.
[0060] In one embodiment, a speaker recognition model learning module (251) of a processor (250) can learn the speaker's voice to identify the speaker. The speaker recognition model learning module (251) can map information of the speaker (e.g., labeling information, name information, and / or language information) in which user interaction (e.g., touch input or gesture input) is detected along with the speaker's voice to extract characteristics such as the frequency and / or amplitude of the speaker (e.g., the speaker's voice). The speaker recognition model learning module (251) can learn a speaker recognition model for the characteristics of the speaker based on the extracted characteristics such as the frequency and / or amplitude of the speaker. The extracted characteristics such as the frequency and / or amplitude of the speaker can be used to distinguish the speaker. For example, if the speaker's name is "Hong Gil-dong," labeling information (or name information) such as "Hong Gil-dong" can be mapped to the training data. The speaker recognition model learning module (251) can learn a speaker recognition model for the characteristics of a speaker in a way that enables real-time learning through the feature vector of a speech signal. The learned speaker recognition model for the characteristics of a speaker can be used to identify the speaker when a new speech signal is received. The speaker recognition model learning module (251) can learn a speaker recognition model for the characteristics of a speaker based on the characteristics of a speaker extracted through a real-time learning method such as continual learning. However, it is not limited to this.
[0061] In one embodiment, the speaker recognition module (252) of the processor (250) can recognize (e.g., identify, distinguish, differentiate) the speaker of a voice signal received through the microphone (230). For example, recognizing the speaker may be a technique for recognizing (e.g., identify, distinguish, differentiate) the voice signal of a specific speaker among voice signals (e.g., voices) corresponding to multiple speakers. The speaker recognition module (252) can recognize (e.g., identify, distinguish, differentiate) the speaker of a voice signal whenever a new voice signal is received through the microphone (230) by using a model that has already been learned (e.g., a speaker recognition model). Even when multiple voice signals corresponding to multiple speakers are received through the microphone (230), the speaker recognition module (252) can recognize (e.g., identify, distinguish, differentiate) the voice signal corresponding to each speaker.
[0062] In one embodiment, the speaker recognition module (252) can analyze the speaker's voice signal to extract characteristics such as frequency and / or amplitude. For example, the speaker recognition module (252) can receive the voice signal and emphasize high-frequency components in the pre-emphasis stage to remove noise and prevent distortion of the voice signal. The speaker recognition module (252) can reduce spectrum leakage by applying a window function after dividing the voice signal. The speaker recognition module (252) can obtain spectrum information through frequency domain transformation using the Discrete Fourier Transform (DFT), separate sub-bands using a filter bank, and then perform filtering. Finally, the speaker recognition module (252) can perform model training or classification tasks by generating feature vectors through energy calculation of each sub-band. In the feature vector extraction step, algorithms such as MFCC (mel frequency cepstral coefficients) and LPC (linear predictive coding) may be used. However, it is not limited to this.
[0063] In one embodiment, the speaker recognition module (252) can classify a speaker using a machine learning algorithm based on characteristics such as extracted frequency and / or amplitude. For example, the speaker recognition module (252) can learn by extracting vectors from a received speech signal through a modeling process such as a Gaussian mixture model (GMM), a support vector machine (SVM), or a deep neural network (DNN). The speaker recognition model (252) can separate features between vectors to predict an input signal through the learned model. However, it is not limited thereto.
[0064] In one embodiment, the language recognition module (253) of the processor (250) can determine which country's language a speaker's voice signal received through the microphone (230) is in by analyzing the speaker's voice signal. The language recognition module (253) can determine which language the received voice signal is written in by using a naive Bayes classifier or a KNN (K-nearest neighbors) algorithm. The language recognition module (253) can determine that the voice signal has the language corresponding to the pattern by analyzing the pattern of a specific language and, if the pattern is detected, determining that the voice signal has the language corresponding to the pattern. For example, the language recognition module (253) can be trained to identify differences between different languages, such as English, Korean, Chinese, and Japanese. In one embodiment, the language recognition module (253) can transmit the identified country information to the voice recognition module (254).
[0065] In one embodiment, the speech recognition module (254) of the processor (250) may select (e.g., determine) a model for recognizing speech signals based on country information received from the language recognition module (253). The speech recognition module (254) may perform speech recognition based on the selected model and transmit the result to a language translation model (255). For example, the speech recognition module (254) may preprocess the speech signal and convert it into text. The converted text may be used for translation, but is not limited thereto.
[0066] In one embodiment, the speech recognition module (254) can digitize a speaker's voice signal (e.g., sound) and convert it into text. For example, the speech recognition module (254) can perform a preprocessing operation to remove noise from the voice signal and extract only the necessary voice signal. The speech recognition module (254) can perform the preprocessing operation using techniques such as high-frequency filtering, low-frequency filtering, and / or spectral flattening. After performing the preprocessing operation, the speech recognition module (254) can perform a feature extraction operation to extract a feature vector representing the characteristics of the voice signal (e.g., sound). For example, the speech recognition module (254) can perform a feature extraction operation to extract the characteristics of the voice signal using MFCC (mel-frequency cepstral coefficients) or LPC (linear predictive coefficients) algorithms. However, it is not limited thereto. Based on the extracted feature vector, the speech recognition module (254) can perform a classification operation to recognize the voice signal. For example, the speech recognition module (254) can perform classification operations using HMM (hidden Markov models) or DNN (deep neural networks) algorithms. However, it is not limited to this.
[0067] In one embodiment, a language translation module (255) of a processor (250) can translate text data (or voice signal) received from a speech recognition module (254) into another language (e.g., a target language). For example, the language translation module (255) can translate using a machine translation method or a human translation method. The translated result may be provided to a user or used in other processes. In one embodiment, the language translation module (255) can convert to the target language while maintaining the meaning, context, and / or style of the original text as much as possible. To this end, the language translation module (255) can translate the text data (or voice signal) into the target language using statistical machine translation (SMT) and / or neural machine translation (NMT) algorithms. However, it is not limited thereto.
[0068] In one embodiment, the processor (250) can translate multiple voice signals of multiple speakers in real time and store them in memory (220). For example, the processor (250) can map information of at least one speaker and translated text associated with at least one speaker and store them in memory (220). The information of at least one speaker and the translated text associated with at least one speaker stored in memory (220) can be used to summarize the content of a conversation (e.g., meeting content) between multiple speakers.
[0069] In one embodiment, a summary generation module (256) of a processor (250) can generate a summary (e.g., a summary of conversation content between multiple speakers (e.g., meeting content)) based on information of at least one speaker among multiple speakers stored in memory (220) and translated text related to at least one speaker. For example, the summary generation module (256) can derive main keywords and topics using natural language processing technology after performing preprocessing operations such as removing unnecessary parts or highlighting necessary parts from the translated text. For example, the summary generation module (256) can generate a summary using algorithms such as latent semantic analysis (LSA) which generates a summary considering topic distribution, textrank which utilizes the connection relationships of words within the text, BERT (bidirectional encoder representations from transformers), or deep learning using GPT (generative pre-trained transformer). The summary generation module (256) can generate a summary of conversation content between multiple speakers (e.g., meeting content) based on derived key keywords and topics. The summary generation module (256) can store the generated summary of conversation content between multiple speakers (e.g., meeting content) in memory (220). The summary of conversation content between multiple speakers (e.g., meeting content) can be shared with at least one of the multiple speakers. If the summary of conversation content between multiple speakers is related to meeting content, the multiple speakers participating in the meeting can clearly understand the matters decided at the meeting through the summary.
[0070] In one embodiment, the processor (250) may display a user interface on the display (240) that provides translation functions for a plurality of speakers. The plurality of speakers may include speakers who use different languages. When the processor (250) (e.g., speaker recognition model learning module (251)) receives a first voice signal through the microphone (230) and detects a first user interaction, it may learn a speaker recognition model for the characteristics of the speaker and the language corresponding to the first voice signal based on the first voice signal. For example, the characteristics of the speaker may include a feature vector including the frequency and / or amplitude of the first voice signal. In one embodiment, the first user interaction may be a trigger input for generating a speaker recognition model for the speaker corresponding to the first voice signal or for learning a previously generated speaker recognition model (e.g., a speaker recognition model that has not yet been trained).
[0071] In one embodiment, when a first voice signal is received and a first user interaction is detected, the processor (250) (e.g., speaker recognition model learning module (251)) maps information of the speaker corresponding to the first voice signal and the selected object, and learns a speaker recognition model for the characteristics and language of the speaker corresponding to the first voice signal. In one embodiment, the processor (250) (e.g., language recognition module (253)) extracts a language feature vector (e.g., language pattern) from the first voice signal and recognizes the language of the first voice signal based on the extracted language feature vector. The processor (250) (e.g., speech recognition module (254)) can convert the first voice signal into text through automatic speech recognition (ASR). The processor (250) (e.g., language translation module (255)) can translate the converted text into a target language using a language translation model corresponding to the first voice signal. The processor (250) can map the translated text (e.g., translated text corresponding to the first voice signal) and the speaker's information (e.g., speaker's information corresponding to the first voice signal) and store them in memory (220).
[0072] In one embodiment, the processor (250) can receive a second voice signal through the microphone (230) and check whether a second user interaction is detected. The second user interaction may be a trigger input for creating a speaker recognition model for the speaker corresponding to the second voice signal or for learning a pre-created speaker recognition model (e.g., a speaker recognition model that has not yet been trained). In one embodiment, the object in which the aforementioned first user interaction is detected and the object in which the second user interaction is detected may be the same or different. In other words, the speaker corresponding to the first voice signal received through the microphone (230) and the speaker corresponding to the second voice signal may be the same or different.
[0073] In one embodiment, when a second voice signal is received and a second user interaction is detected, the processor (250) determines to learn a speaker recognition model of the speaker corresponding to the second voice signal, and can learn a speaker recognition model for the characteristics and language of the speaker corresponding to the second voice signal based on the second voice signal. When the second voice signal is received and a second user interaction is not detected, the processor (250) can check whether there is a speaker recognition model that has completed learning corresponding to the second voice signal. If there is a speaker recognition model that has completed learning corresponding to the second voice signal, the processor (250) can translate the second voice signal into a target language based on the speaker recognition model that has completed learning corresponding to the second voice signal and display it on the display (240). The processor (250) can map the translated text (e.g., translated text corresponding to the second voice signal) and the speaker's information (e.g., speaker's information corresponding to the second voice signal) and store them in memory (220).
[0074] In one embodiment, if there is no speaker recognition model corresponding to the second voice signal, the processor (250) (e.g., language recognition module (253)) can identify a language recognition model (ASR) and a language translation model (e.g., translator) corresponding to the second voice signal. The processor (250) (e.g., language translation module (255)) can translate the second voice signal into a target language and display it on the display (240) based on the language translation model corresponding to the language of the recognized second voice signal.
[0075] An electronic device (101) according to one embodiment of the present disclosure may include a microphone (230), a display (240), at least one processor (250) including processing circuitry, and a memory (220) for storing instructions. In one embodiment, instructions may cause the electronic device (101) to display a user interface on the display (240) that provides translation functions for a plurality of speakers when executed individually or collectively by at least one processor (250). In one embodiment, instructions may cause the electronic device (101) to learn a speaker recognition model for the characteristics and language of the speaker corresponding to the first voice signal based on the first voice signal when a first voice signal is received through the microphone (230) and a first user interaction is detected. Instructions according to one embodiment, when executed individually or collectively by at least one processor (250), may cause the electronic device (101) to check whether a second voice signal is received through the microphone (230) and whether a second user interaction is detected. Instructions according to one embodiment, when executed individually or collectively by at least one processor (250), may cause the electronic device (101) to check whether a trained speaker recognition model corresponding to the second voice signal exists if the second user interaction is not detected. Instructions according to one embodiment, when executed individually or collectively by at least one processor (250), may cause the electronic device (101) to translate the second voice signal into a target language based on the trained speaker recognition model and display it on the display (240) if a trained speaker recognition model corresponding to the second voice signal exists.
[0076] Instructions according to one embodiment, when executed individually or collectively by at least one processor (250), may cause an electronic device (101) to map a first voice signal and information of a speaker corresponding to the first voice signal based on receiving a first voice signal and detecting a first user interaction. Instructions according to one embodiment, when executed individually or collectively by at least one processor (250), may cause an electronic device (101) to store the mapped first voice signal and information of a speaker corresponding to the first voice signal in a memory (220).
[0077] Instructions according to one embodiment, when executed individually or collectively by at least one processor (250), may cause an electronic device (101) to learn a speaker recognition model for the characteristics and language of a speaker corresponding to the second voice signal based on the second voice signal when a second voice signal is received through a microphone (230) and a second user interaction is detected. According to one embodiment, the speaker corresponding to the first voice signal and the speaker corresponding to the second voice signal may be the same or different.
[0078] Instructions according to one embodiment, when executed individually or collectively by at least one processor (250), may cause the electronic device (101) to identify the language of the second voice signal if there is no speaker recognition model corresponding to the second voice signal. Instructions according to one embodiment, when executed individually or collectively by at least one processor (250), may cause the electronic device (101) to translate the second voice signal into a target language and display it on the display (240) based on a language translation model corresponding to the language of the second voice signal.
[0079] When instructions according to one embodiment are executed individually or collectively by at least one processor (250), the electronic device (101) may display a user interface on a display (240) that includes a plurality of objects representing a plurality of speakers. A first user interaction or a second user interaction according to one embodiment may include an input for selecting one of the plurality of objects. A speaker corresponding to a first voice signal according to one embodiment may be a speaker corresponding to the selected object.
[0080] Instructions according to one embodiment, when executed individually or collectively by at least one processor (250), may cause the electronic device (101) to learn a speaker recognition model for the speaker's characteristics and language, and then translate the first voice signal into a target language using a language translation model corresponding to the language of the first voice signal. Instructions according to one embodiment, when executed individually or collectively by at least one processor (250), may cause the electronic device (101) to display the speaker's information corresponding to the first voice signal and the text translated into the target language on a display (240).
[0081] When instructions according to one embodiment are executed individually or collectively by at least one processor (250), the electronic device (101) may map speaker information corresponding to a first voice signal and text corresponding to the first voice signal translated into a target language and store them in memory (220). When instructions according to one embodiment are executed individually or collectively by at least one processor (250), the electronic device (101) may map speaker information corresponding to a second voice signal and text corresponding to the second voice signal translated into a target language and store them in memory (220).
[0082] Instructions according to one embodiment, when executed individually or collectively by at least one processor (250), may cause the electronic device (101) to generate a summary of a conversation involving a plurality of speakers based on speaker information corresponding to a first voice signal stored in memory (220), text corresponding to the first voice signal translated into a target language, speaker information corresponding to a second voice signal, and text corresponding to the second voice signal translated into a target language.
[0083] Instructions according to one embodiment, when executed individually or collectively by at least one processor (250), may enable the electronic device (101) to identify the characteristics of a speaker, including the frequency and amplitude of the first voice signal, based on the first voice signal. Instructions according to one embodiment, when executed individually or collectively by at least one processor (250), may enable the electronic device (101) to identify the language based on the pattern of the first voice signal. Instructions according to one embodiment, when executed individually or collectively by at least one processor (250), may enable the electronic device (101) to learn a speaker recognition model of the speaker corresponding to the first voice signal based on the frequency, amplitude, and language of the first voice signal.
[0084] Instructions according to one embodiment, when executed individually or collectively by at least one processor (250), may cause the electronic device (101) to check whether the learning of the speaker recognition model of the speaker corresponding to the first voice signal has been completed. Instructions according to one embodiment, when executed individually or collectively by at least one processor (250), may cause the electronic device (101) to distinguishly display an object representing the speaker and at least one object representing at least one speaker among a plurality of speakers whose learning of the speaker recognition model has not been completed, when the learning of the speaker recognition model of the speaker corresponding to the first voice signal has been completed.
[0085] FIG. 3 is a flowchart illustrating a method for providing a translation function according to one embodiment of the present disclosure.
[0086] In the following embodiments, each operation of FIG. 3 may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation of FIG. 3 may be changed, and at least two operations may be performed in parallel.
[0087] According to one embodiment, operations 305 through 325 of FIG. 3 can be understood as being performed in a processor (e.g., processor (250) of FIG. 2) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0088] Referring to FIG. 3, the processor (250) can display a user interface that provides translation functions for multiple speakers on a display (e.g., the display (240) of FIG. 2) in operation 305.
[0089] In one embodiment, the plurality of speakers may include speakers who use different languages.
[0090] In one embodiment, the user interface may include a plurality of objects representing a plurality of speakers. Not limited thereto, the user interface may further include an object for setting the language of each of the plurality of speakers and / or an object for setting a target language. For example, the target language may be a language that the user of the electronic device (101) wishes to receive through translation.
[0091] In one embodiment, the processor (250) may, in operation 310, receive a first voice signal through a microphone (e.g., the microphone (230) of FIG. 2) and, when a first user interaction is detected, learn a speaker recognition model for the characteristics and language of the speaker corresponding to the first voice signal based on the first voice signal. For example, the processor (250) may extract a feature vector of the speaker from the first voice signal and learn a speaker recognition model for the speaker based on the extracted feature vector. For example, the feature vector may include the frequency and / or amplitude of the first voice signal, but is not limited thereto.
[0092] In one embodiment, the first user interaction may be a trigger input for generating a speaker recognition model for a speaker corresponding to the first voice signal or for training a previously generated speaker recognition model (e.g., a speaker recognition model that has not yet been trained). For example, the first user interaction may include a touch input or a gesture input for selecting one of a plurality of objects representing a plurality of speakers included within a user interface. However, it is not limited thereto. For example, the selected object may include an object representing a speaker corresponding to the first voice signal.
[0093] In one embodiment, when a first voice signal is received and a first user interaction is detected, the processor (250) can map the first voice signal to the speaker information corresponding to the selected object. The speaker information according to one embodiment may include the speaker's name and / or language information (e.g., language information used by the speaker). However, it is not limited thereto. For example, when a user of the electronic device (101) receives the first voice signal through the microphone (230), the user may select an object representing the speaker corresponding to the first voice signal among a plurality of speakers. When a first user interaction is detected in which the object representing the speaker corresponding to the first voice signal is selected, the processor (250) can map the first voice signal to the speaker information corresponding to the selected object (e.g., the speaker's name and / or language information used by the speaker) and learn a speaker recognition model for the characteristics and language of the speaker corresponding to the first voice signal. For example, if the processor (250) does not have a speaker recognition model for a speaker corresponding to the first voice signal, it can create a speaker recognition model for a speaker corresponding to the first voice signal and then learn the speaker recognition model. The learned speaker recognition model can be stored in memory (e.g., memory (220) of FIG. 2). As another example, if the processor (250) has a speaker recognition model for a speaker corresponding to the first voice signal, it can perform learning using the speaker recognition model for a speaker corresponding to the first voice signal stored in memory (220).
[0094] The aforementioned 310 operation can be performed by the speaker recognition model learning module (251) of FIG. 2.
[0095] In one embodiment, although not illustrated, a processor (250) (e.g., a language recognition module (253)) may extract a language feature vector from a first voice signal and recognize the language of the first voice signal based on the extracted language feature vector. For example, the language feature vector may include a language pattern. The processor (250) (e.g., a speech recognition module (254)) may convert the first voice signal into text through automatic speech recognition. For example, the processor (250) (e.g., a speech recognition module (254)) may convert the first voice signal into text based on the first voice signal and information of the speaker corresponding to the selected object, e.g., language information used by the speaker. Not limited thereto, the processor (250) (e.g., a speech recognition module (254)) may identify language information based on a language recognition model and convert the first voice signal into text using a speech recognition model associated with the identified language information. For example, if the first voice signal is recognized as Korean, the processor (250) (e.g., voice recognition module (254)) can convert the first voice signal into text using a Korean voice recognition model (e.g., Korean Automatic Speech Recognition (ASR)). As another example, if the first voice signal is recognized as English, the first voice signal can be converted into text using an English voice recognition model (e.g., English Automatic Speech Recognition (ASR)).
[0096] In one embodiment, a processor (250) (e.g., a language translation module (255)) can translate converted text into a target language using a language translation model corresponding to a first voice signal. The target language according to one embodiment may be a state pre-set by the user of the electronic device (101) in the user interface displayed in the operation 305. The processor (250) can display information of a speaker (e.g., a speaker corresponding to an object selected by the first user interaction) along with text corresponding to the first voice signal translated into the target language on the display (240). The processor (250) can map the translated text and the speaker's information and store them in memory (220).
[0097] In one embodiment, the processor (250) can check whether a second voice signal is received through the microphone (230) and whether a second user interaction is detected in operation 315.
[0098] In one embodiment, the second user interaction may be a trigger input for generating a speaker recognition model for a speaker corresponding to the second voice signal or for training a previously generated speaker recognition model (e.g., a speaker recognition model that has not yet been trained). In one embodiment, the second user interaction may include a touch input or a gesture input for selecting one of a plurality of objects representing a plurality of speakers included in the user interface (e.g., an object representing a speaker corresponding to the second voice signal). However, it is not limited thereto.
[0099] In one embodiment, the object in which the aforementioned first user interaction is detected and the object in which the second user interaction is detected may be the same or different. In other words, the speaker corresponding to the first voice signal received through the microphone (230) and the speaker corresponding to the second voice signal may be the same or different.
[0100] In one embodiment, if a second voice signal is received and a second user interaction is not detected, the processor (250) determines not to learn the speaker recognition model of the speaker corresponding to the second voice signal and can perform the 320 operation described below.
[0101] In one embodiment, when a second voice signal is received and a second user interaction is detected, the processor (250) determines to learn a speaker recognition model of the speaker corresponding to the second voice signal, and can learn a speaker recognition model for the characteristics and language of the speaker corresponding to the second voice signal based on the second voice signal.
[0102] In one embodiment, in operation 320, the processor (250) can check whether there is a trained speaker recognition model corresponding to the second voice signal if the second voice signal is received and the second user interaction is not detected. In operation 325, if there is a trained speaker recognition model corresponding to the second voice signal, the processor (250) can translate the second voice signal into a target language based on the trained speaker recognition model corresponding to the second voice signal and display it on the display (240).
[0103] In one embodiment, a processor (250) (e.g., speaker recognition module (252)) can recognize a speaker by extracting characteristics of a second voice signal (e.g., frequency and / or amplitude of the second voice signal). A processor (250) (e.g., language recognition module (253)) can extract a feature vector of the language of the second voice signal (e.g., pattern of language) and recognize the language of the second voice signal based on the extracted feature vector of the language of the second voice signal (e.g., pattern of language). Based on the characteristics and language of the second voice signal, the processor (250) can determine whether there is a speaker recognition model corresponding to the second voice signal among the speaker recognition models that have completed training. If a speaker recognition model corresponding to the second voice signal exists, the processor (250) (e.g., language translation module (255)) can translate the second voice signal into a target language using the speaker recognition model corresponding to the second voice signal. The processor (250) can display the speaker's information along with the translated text on the display (240). The processor (250) can map the translated text and the speaker's information and store them in memory (220).
[0104] In one embodiment, although not illustrated, if there is no speaker recognition model corresponding to the second voice signal, the processor (250) (e.g., language recognition module (253)) can identify a language recognition model (ASR) and a language translation model (e.g., translator) corresponding to the second voice signal. For example, the processor (250) (e.g., language recognition module (253)) can extract a feature vector (e.g., language pattern) of the language of the second voice signal and recognize the language of the second voice signal based on the extracted feature vector (e.g., language pattern) of the language of the second voice signal. The processor (250) (e.g., voice recognition module (254)) can convert the second voice signal into text based on the voice recognition model corresponding to the recognized language of the second voice signal. The processor (250) (e.g., language translation module (255)) can translate the second voice signal into a target language based on the language translation model corresponding to the recognized language of the second voice signal and display it on the display (240). For example, the processor (250) can display text corresponding to a second voice signal translated into a target language on the display (240).
[0105] In one embodiment, although not illustrated, when a plurality of voice signals are received through the microphone (230) at the time when a user interaction (e.g., a first user interaction or a second user interaction) is detected, the processor (250) may not perform the operation of learning the speaker recognition model of the corresponding speaker. However, it is not limited thereto.
[0106] In FIG. 2 according to one embodiment, it is described that text translated into a target language is displayed on a display (240), but it is not limited thereto. For example, if there is a wearable electronic device (e.g., an ear-type wireless audio device) that is connected to the electronic device (101) in communication, the translated text may be provided as sound through the speaker of the wearable electronic device.
[0107] FIGS. 4a and FIGS. 4b are flowcharts illustrating a method for providing a translation function according to one embodiment of the present disclosure.
[0108] In the following embodiments, each operation of FIGS. 4a and FIGS. 4b may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation of FIGS. 4a and FIGS. 4b may be changed, and at least two operations may be performed in parallel.
[0109] According to one embodiment, operations 405 through 495 of FIGS. 4a and 4b can be understood as being performed in a processor (e.g., processor (250) of FIG. 2) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0110] FIGS. 4a and FIG. 4b according to various embodiments may be drawings embodying the operations of FIG. 3 described above.
[0111] Referring to FIGS. 4a and 4b, the processor (250) may display a user interface on a display (e.g., the display (240) of FIG. 2) in operation 405, the user interface may further include, in addition to the aforementioned multiple objects, an object for setting the language of each of the multiple speakers and / or an object for setting a target language (e.g., a language that the user of the electronic device (101) wishes to receive as a translation).
[0112] In one embodiment, the plurality of speakers may include speakers who use different languages.
[0113] In one embodiment, the processor (250) can check whether a first voice signal is received through a microphone (e.g., the microphone (230) of FIG. 2) in operation 410. If the first voice signal is received through the microphone (230) (e.g., YES in operation 410), the processor (250) can check whether a first user interaction with the first object among a plurality of objects is detected in operation 415. For example, the first user interaction may be a trigger input for creating a first speaker recognition model for the first speaker corresponding to the first voice signal or for learning a first speaker recognition model that has already been created (e.g., a first speaker recognition model that has not yet been trained). For example, the first user interaction may include a touch input or a gesture input for selecting a first object representing the first speaker corresponding to the first voice signal among a plurality of objects included in the user interface. However, it is not limited thereto.
[0114] In one embodiment, when a first user interaction with a first object among a plurality of objects is detected (e.g., YES of operation 415), the processor (250) (e.g., speaker recognition model learning module (251) of FIG. 2) can learn a first speaker recognition model for a first characteristic of a first speaker and a first language corresponding to a first voice signal based on a first voice signal in operation 420. For example, the processor (250) (e.g., speaker recognition model learning module (251)) can extract the characteristics of the first speaker (e.g., frequency and / or amplitude of the first voice signal) and the language of the first speaker from the first voice signal, and learn a first speaker recognition model for the first speaker based on the extracted characteristics and language. For example, when a first voice signal is received and a first user interaction is detected, a processor (250) (e.g., speaker recognition model learning module (251)) can map information of a first speaker corresponding to the first voice signal and a selected first object (e.g., name information of the first speaker and / or language information of the first speaker (e.g., first language)) and learn a first speaker recognition model for the first characteristics of the first speaker and the first language corresponding to the first voice signal.
[0115] In one embodiment, the processor (250) can, in operation 425, translate the first voice signal into a target language and display it on the display (240) based on a first language translation model corresponding to the first voice signal (e.g., a first language translation model corresponding to the language of the first voice signal). For example, the processor (250) (e.g., the voice recognition module (254) of FIG. 2) can convert the first voice signal into text by performing automatic voice recognition on the first language of the first speaker. The processor (250) (e.g., the language translation module (255) of FIG. 2) can translate the converted text into a target language using a first language translation model corresponding to the first voice signal (e.g., a first language translation model corresponding to the language of the first voice signal). The processor (250) can display information (e.g., the name of the first speaker) of the first speaker corresponding to the first voice signal (e.g., the first speaker corresponding to the first object selected by the first user interaction) along with text translated into the target language on the display (240).
[0116] In one embodiment, although not illustrated, the processor (250) may map the text translated into the target language and the information of the first speaker corresponding to the first voice signal (e.g., name information of the first speaker and / or language information of the first speaker (e.g., first language)) and store it in memory (220).
[0117] In one embodiment, the processor (250) can check whether the training of the first speaker recognition model is complete in operation 430. If the training of the first speaker recognition model is complete (e.g., YES in operation 430), the processor (240) can terminate the training of the first speaker recognition model.
[0118] In one embodiment, although not illustrated, the processor (250) may apply a visual effect to a first object representing the first speaker among a plurality of objects when the training of the first speaker recognition model is completed, so as to display it differently from at least one object for which the training of the speaker recognition model is not completed. Accordingly, the user of the electronic device (101) can intuitively identify the first speaker for which the training of the first speaker recognition model is completed, and may not perform a first user interaction (e.g., a first user interaction for the first object) to perform the training of the first speaker recognition model of the first speaker.
[0119] In one embodiment, if the learning of the first speaker recognition model is not completed (e.g., NO of operation 430), the processor (240) may repeat operations 410 through 425. For example, when the first voice signal of the first speaker and the first user interaction are detected, the processor (250) may repeat the operation of learning the first speaker recognition model for the first characteristic of the first speaker and the first language to complete the learning of the first speaker recognition model.
[0120] In one embodiment, if a first user interaction with a first object among a plurality of objects is not detected (e.g., NO of operation 415), the processor (250) can check in operation 435 whether there is a trained speaker recognition model corresponding to the first voice signal. If there is a trained speaker recognition model corresponding to the first voice signal (e.g., YES of operation 435), the processor (250) can translate the first voice signal into a target language and display it on the display (240) based on the trained speaker recognition model in operation 440.
[0121] In one embodiment, if there is no speaker recognition model that has completed learning corresponding to the first voice signal (e.g., NO of operation 435), the processor (250) may, in operation 445, translate the first voice signal into a target language and display it on the display (240) based on the first language recognition model and the first language translation model corresponding to the first voice signal (e.g., the first language translation model corresponding to the language of the first voice signal). For example, if the first user interaction is not detected, the processor (250) may decide not to learn the first speaker recognition model for the first speaker corresponding to the first voice signal. In other words, the processor (250) may not perform the operation of learning the first speaker recognition model for the first speaker corresponding to the first voice signal, but may only perform the operation of translating the first voice signal into a target language using the first language recognition model and the first language translation model, and displaying the translated target language on the display (240).
[0122] In one embodiment, if the first voice signal is not received through the microphone (230) (e.g., NO of operation 410), the processor (250) can check whether the second voice signal is received through the microphone (230) in operation 460.
[0123] In FIG. 4a and FIG. 4b according to one embodiment, the speaker corresponding to the second voice signal and the speaker corresponding to the first voice signal are assumed to be different for the explanation.
[0124] In one embodiment, if a second voice signal is not received through the microphone (230) (e.g., NO of operation 460), the processor (250) may branch to operation 405 and maintain an operation that displays a user interface including multiple objects representing multiple speakers.
[0125] In one embodiment, when a second voice signal is received through the microphone (230) (e.g., YES of operation 460), the processor (250) can determine whether a second user interaction with the second object among a plurality of objects is detected in operation 465. For example, the second user interaction may be a trigger input for creating a second speaker recognition model for the second speaker corresponding to the second voice signal, or for learning a previously created second speaker recognition model (e.g., a second speaker recognition model that has not yet been trained). For example, the second user interaction may include a touch input or a gesture input for selecting a second object representing the second speaker corresponding to the second voice signal among a plurality of objects included in the user interface. However, it is not limited thereto.
[0126] In one embodiment, when a second user interaction with a second object among a plurality of objects is detected (e.g., YES of operation 465), the processor (250) (e.g., speaker recognition model learning module (251)) can learn a second speaker recognition model for a second characteristic and a second language of a second speaker corresponding to a second voice signal based on a second voice signal in operation 470. For example, the processor (250) (e.g., speaker recognition model learning module (251)) can extract the characteristics of the second speaker (e.g., frequency and / or amplitude of the second voice signal) and the language of the second speaker from the second voice signal, and learn a second speaker recognition model for the second speaker based on the extracted characteristics and language of the second speaker. For example, when a second voice signal is received and a second user interaction is detected, the processor (250) (e.g., speaker recognition model learning module (251)) maps information of the second speaker corresponding to the second voice signal and the selected second object (e.g., name information of the second speaker and / or language information of the second speaker (e.g., second language)) and learns a second speaker recognition model for the second characteristics of the second speaker and the second language corresponding to the second voice signal.
[0127] In one embodiment, the processor (250) can, in operation 475, translate the second voice signal into a target language and display it on the display (240) based on a second language translation model corresponding to the second voice signal (e.g., a second language translation model corresponding to the language of the second voice signal). For example, the processor (250) (e.g., a voice recognition module (254)) can perform automatic voice recognition for the second language of the second speaker and convert the second voice signal into text. The processor (250) (e.g., a language translation module (255)) can translate the converted text into a target language using a second language translation model corresponding to the second voice signal (e.g., a second language translation model corresponding to the language of the second voice signal). The processor (250) can display information (e.g., the name of the second speaker) of the second speaker corresponding to the second voice signal (e.g., the second speaker corresponding to the second object selected by the second user interaction) along with the text translated into the target language on the display (240).
[0128] In one embodiment, the processor (250) can check whether the training of the second speaker recognition model is completed in operation 480. If the training of the second speaker recognition model is completed (e.g., YES in operation 480), the processor (240) can terminate the operation of training the second speaker recognition model.
[0129] Although not shown, the processor (250) may apply a visual effect to a second object representing the second speaker among multiple objects when the training of the second speaker recognition model is completed, so as to display it differently from at least one object for which the training of the speaker recognition model is not completed. Accordingly, the user of the electronic device (101) can intuitively identify the second speaker for which the training of the second speaker recognition model is completed, and may not perform a second user interaction (e.g., a second user interaction for the second object) to perform the training of the second speaker recognition model of the second speaker.
[0130] In one embodiment, if the learning of the second speaker recognition model is not completed (e.g., NO of operation 480), the processor (240) may repeat operations 460 through 475. For example, when the second voice signal of the second speaker and the second user interaction are detected, the processor (250) may repeat the operation of learning the second speaker recognition model for the second characteristics and the second language of the second speaker to complete the learning of the second speaker recognition model for the second speaker.
[0131] In one embodiment, if a second user interaction with a second object among a plurality of objects is not detected (e.g., NO in operation 465), the processor (250) can check in operation 485 whether there is a trained speaker recognition model corresponding to the second voice signal. If there is a trained speaker recognition model corresponding to the second voice signal (e.g., YES in operation 485), the processor (250) can translate the second voice signal into a target language and display it on the display (240) based on the trained speaker recognition model in operation 490.
[0132] In one embodiment, if there is no speaker recognition model that has completed learning corresponding to the second voice signal (e.g., NO of operation 485), the processor (250) may, in operation 495, translate the second voice signal into a target language and display it on the display (240) based on a second language recognition model corresponding to the second voice signal and a second language translation model (e.g., a second language translation model corresponding to the language of the second voice signal). For example, if the second user interaction is not detected, the processor (250) may decide not to learn a speaker recognition model for the second speaker corresponding to the second voice signal. In other words, the processor (250) may not perform the operation of learning a second speaker recognition model for the second speaker corresponding to the second voice signal, but may only perform the operation of translating the second voice signal into a target language using the second language recognition model and the second language translation model and displaying the translated target language on the display (240).
[0133] In FIG. 4a and FIG. 4b according to various embodiments, it is described that a speaker recognition model of one speaker or a speaker recognition model of each of two speakers is learned based on a first voice signal and a second voice signal, but is not limited thereto. When multiple speakers exceeding two participate in a conversation, the processor (250) can learn a speaker recognition model of each of the corresponding speakers exceeding two based on the voice signal received through the microphone (230) and the detection of user interaction.
[0134] FIG. 5 is a drawing for illustrating a user interface that provides a translation function for a plurality of speakers according to one embodiment of the present disclosure.
[0135] Referring to FIG. 5, the processor (e.g., processor (250) of FIG. 2) of an electronic device (e.g., electronic device (101) of FIG. 1) <510> As illustrated in the figure, a first user interface including a first area (515) for setting the number of people participating in a meeting (or conversation) and a second area (520) for setting the target language can be displayed on a display (e.g., the display (240) of FIG. 2). In each of the first area (515) and the second area (520), information on the number of people participating in the meeting (e.g., 6 people) and information on the target language (e.g., Korean) can be displayed.
[0136] In one embodiment, when the number of participants in the meeting and the target language are set, the processor (250) <550> As illustrated in the figure, a second user interface including multiple objects representing multiple speakers participating in a meeting can be displayed on the display (240). Each of the multiple objects may include information about the speaker (e.g., Speaker A, Speaker B (not shown), Speaker C (not shown), Speaker D, Speaker E (not shown), Speaker F (not shown)). This is not limited thereto, and each of the multiple objects may further include language information for each speaker (e.g., Korean, English, Japanese, etc.).
[0137] In one embodiment, the processor (250) may detect an input for setting language information for each speaker. For example, the processor (240) may set language information for the speaker corresponding to the object in which the user interaction is detected, based on the detection of a user interaction selecting an object representing each speaker. For example, when the processor (250) detects an input (560) selecting a first object (555) (e.g., a first object representing speaker A), the processor (250) may display a list of configurable language information (565) in close proximity to the first object (555). When a specific language is selected from the list of configurable language information (565), the processor (250) may map the information of speaker A (e.g., identification information (e.g., name information)) corresponding to the first object (555) with the selected specific language information and display the selected specific language information together with the information of speaker A. In one embodiment, for Speaker D, it can be confirmed that Speaker D's language information is set to English, and accordingly, the processor (250) can display Speaker D's language information (575) (e.g., English) together with Speaker D's information.
[0138] In FIG. 5 according to one embodiment, it is described that language information of each speaker is set through a language information list (565), but this is not limited thereto. For example, the processor (250) may set language information of a speaker corresponding to a voice signal by identifying the language based on the pattern of the voice signal received through a microphone (e.g., the microphone (230) of FIG. 2).
[0139] FIG. 6 is a drawing for explaining a method of providing a translation function according to one embodiment of the present disclosure.
[0140] Referring to FIG. 6, a processor (e.g., processor (250) of FIG. 2) of an electronic device (e.g., electronic device (101) of FIG. 1) may display a user interface (620) comprising a plurality of objects (621) representing a plurality of speakers on a display (e.g., display (240) of FIG. 2). The plurality of speakers may include Speaker A (601), Speaker B (603), Speaker C (605), Speaker D (607), Speaker E (609), and Speaker F (611). The plurality of speakers may include speakers who use different languages.
[0141] In one embodiment, the user interface (620) may include the aforementioned plurality of objects (621), for example, a first object (601a) representing speaker A (601), a second object (603a) representing speaker B (603), a third object (605a) representing speaker C (605), a fourth object (607a) representing speaker D (607), a fifth object (609a) representing speaker E (609), and a sixth object (611a) representing speaker F (611).
[0142] In one embodiment, the user interface (620) may further display target language information (615) (e.g., Korean).
[0143] In one embodiment, the processor (250) can determine whether a voice signal (625) is received through a microphone (e.g., the microphone (230) of FIG. 2). For example, the voice signal (625) may include “Hello, I work for AI team.” When the voice signal (625) (e.g., “Hello, I work for AI team”) is received through the microphone (230), the processor (250) can determine whether a user interaction (610) is detected for one of the objects (621). For example, the user interaction (610) may include a touch input or gesture input for selecting a first object (601a) representing Speaker A (601) corresponding to the voice signal (625) among the objects (621) included in the user interface (620). For example, a touch input or gesture input for selecting a first object (601a) representing speaker A (601) corresponding to a voice signal (625) may be a trigger input for creating a speaker recognition model for speaker A (601) corresponding to the voice signal (625) or for learning a speaker recognition model that has already been created (e.g., a speaker recognition model that has not yet been trained).
[0144] In one embodiment, when user interaction (610) is detected for a first object (601a) representing speaker A (601) corresponding to a voice signal (625) among a plurality of objects (621) included in a user interface (620), the processor (250) (e.g., speaker recognition model learning module (251) of FIG. 2) can learn a speaker recognition model for speaker A (601) based on the characteristics of speaker A (601) (e.g., frequency and / or amplitude of the voice signal (625)) and the language of speaker A (601) (e.g., English).
[0145] In one embodiment, a processor (250) (e.g., the speech recognition module (254) of FIG. 2) can perform automatic speech recognition for the language (e.g., English) of speaker A (601) to convert the speech signal (625) into text. The processor (250) (e.g., the language translation module (255) of FIG. 2) can translate the converted text into a target language (e.g., Korean) using a language translation model (e.g., an English translation model) corresponding to the speech signal (625). The processor (250) can display information (635) (e.g., A) of speaker A (601) corresponding to the speech signal (625) on a display (240) along with the text (630) (e.g., Hello. I am working at the AI team.) translated into the target language (e.g., Korean).
[0146] In one embodiment, the processor (250) can map information (635) (e.g., name information, language information) of speaker A (601) corresponding to the translated text (630) and voice signal (625) and store it in memory (e.g., memory (220) of FIG. 2).
[0147] FIG. 7 is a drawing for explaining a method of providing a translation function according to one embodiment of the present disclosure.
[0148] Referring to FIG. 7, a processor (e.g., processor (250) of FIG. 2) of an electronic device (e.g., electronic device (101) of FIG. 1) may display a user interface (620) comprising a plurality of objects (621) representing a plurality of speakers on a display (e.g., display (240) of FIG. 2). The plurality of speakers may include Speaker A (601), Speaker B (603), Speaker C (605), Speaker D (607), Speaker E (609), and Speaker F (611). The plurality of speakers may include speakers who use different languages.
[0149] In one embodiment, the user interface (620) may include the aforementioned plurality of objects (621), for example, a first object (601a) representing speaker A (601), a second object (603a) representing speaker B (603), a third object (605a) representing speaker C (605), a fourth object (607a) representing speaker D (607), a fifth object (609a) representing speaker E (609), and a sixth object (611a) representing speaker F (611).
[0150] In one embodiment, the user interface (620) may further display target language information (615) (e.g., Korean).
[0151] In one embodiment, when a voice signal of Speaker A (601) of FIG. 6 described above and a user interaction (e.g., a user interaction with the first object (601a)) are detected, the processor (250) may repeat the operation of learning a speaker recognition model for the characteristics and language of Speaker A (601) until the learning of the speaker recognition model for Speaker A (601) is completed. Although not shown, the learning of a speaker recognition model for Speaker B (603), Speaker C (605), Speaker D (607), Speaker E (609), and / or Speaker F (611) may also be performed in the same manner as the learning of the speaker recognition model for Speaker A (601).
[0152] In one embodiment, the processor (250) may display at least one object corresponding to at least one speaker for whom the training of the speaker recognition model is completed among a plurality of objects differently from at least one other object corresponding to at least one other speaker for whom the training is not completed. For example, the processor (250) may apply a visual effect to at least one object corresponding to at least one speaker for whom the training of the speaker recognition model is completed among a plurality of objects so as to distinguish it from at least one other object corresponding to at least one other speaker for whom the training is not completed.
[0153] In FIG. 7 according to various embodiments, the speakers who have completed training are assumed to be Speaker B (603), Speaker D (607), and Speaker F (611). In this case, the processor (250) may apply visual effects (e.g., making the outline thickness of each object (603a, 607a, and 611a) thicker) to the second object (603a) representing Speaker B (603) who has completed training in the speaker recognition model, the fourth object (607a) representing Speaker D (607), and the sixth object (611a) representing Speaker F (611) to display them. The processor (250) may not apply visual effects to the first object (601a) representing Speaker A (601) for whom the speaker recognition model training is not complete, the third object (605a) representing Speaker C (605), and the fifth object (609a) representing Speaker E (609). Accordingly, the user of the electronic device (101) can intuitively identify the speaker for whom the speaker recognition model training is complete, and may not perform user interactions (e.g., user interactions with objects) to perform training on the speaker recognition model for which training is complete.
[0154] In one embodiment, for a speaker whose speaker recognition model has been trained, when a voice signal is received through the microphone (230), the processor (250) can identify the speaker and language corresponding to the voice signal using the trained speaker recognition model without separate input, and automatically display the speaker's information and text translated into the target language on the display (240). For example, when speaker D (607) speaks, the processor (250) receives the voice signal (710) spoken by speaker D (607) through the microphone (230) (e.g., today The second proposal of the discussion can be received. When speaker D (607) speaks, the processor (250) can display a second visual effect applied to a fourth object (607a) representing speaker D (607) so that the user can recognize the speaker speaking (e.g., applying an effect that fills the inside of the fourth object (607a)).
[0155] In one embodiment, when a voice signal (710) is received and no user interaction is detected for a specific object among a plurality of objects, the processor (250) can check whether there is a trained speaker recognition model corresponding to the voice signal (710), for example, a speaker recognition model for speaker D (607). If there is a trained speaker recognition model for speaker D (607), the processor (250) can translate the voice signal (710) into a target language (e.g., Korean) based on the trained speaker recognition model for speaker D (607) and display it on the display (240). For example, based on the trained speaker recognition model for speaker D (607), it can be confirmed that the language used by speaker D (607) is “Japanese,” and the voice signal (710) can be translated into a target language (e.g., Korean) using a language translation model corresponding to “Japanese” (e.g., a Japanese translation model). The processor (250) can display information (730) (e.g., D) of speaker D (607) corresponding to a voice signal (710) on the display (240) along with text (720) translated into a target language (e.g., Korean) (e.g., The second agenda item of today's meeting is as follows).
[0156] In one embodiment, the processor (250) can map information (730) of speaker D (607) corresponding to the translated text (720) and voice signal (710) and store it in memory (e.g., memory (220) of FIG. 2).
[0157] FIGS. 8 to 12 are drawings for explaining a method of providing a translation function according to one embodiment of the present disclosure.
[0158] FIGS. 8 to 12 according to various embodiments may be additional operations of FIG. 7 described above.
[0159] Referring to FIG. 8, when a processor (e.g., processor (250) of FIG. 2) of an electronic device (e.g., electronic device (101) of FIG. 1) receives a voice signal (810) (e.g., “When shall we schedule the next meeting”) through a microphone (e.g., microphone (230) of FIG. 2), the processor (250) [receives] a specific object among a plurality of objects (621) (e.g., a first object (601a) representing speaker A (601), a second object (603a) representing speaker B (603), a third object (605a) representing speaker C (605), a fourth object (607a) representing speaker D (607), a fifth object (609a) representing speaker E (609), and a sixth object (611a) representing speaker F (611)) (e.g., a first object representing speaker A (601) corresponding to the voice signal (810). It can be checked whether user interaction is detected for an object (601a)). If user interaction is not detected for a specific object among the multiple objects (621) (e.g., a first object (601a) representing Speaker A (601) corresponding to the voice signal (810), the processor (250) can check whether there is a trained speaker recognition model corresponding to the voice signal (810). As previously discussed, the speaker recognition model of Speaker A (601) corresponding to the voice signal (810) may not be in a state where training is complete. If there is no trained speaker recognition model corresponding to the voice signal (810), the processor (250) can translate the voice signal (810) into a target language (e.g., Korean) based on a language translation model (e.g., an English translation model) corresponding to the voice signal (810). The processor (250) can then translate the text (820) (e.g., next meeting) into the target language (e.g., Korean). Only "When should we set the schedule?" can be displayed on the display (240).For example, as the processor (250) does not perform the operation of learning a speaker recognition model for speaker A (601) corresponding to the voice signal (810), the processor (250) may not display information of speaker A (601) on the display (240). In this case, the processor (250) may store only text (820) translated into a target language (e.g., Korean) (e.g., When should we schedule the next meeting?) in memory (e.g., memory (220) of FIG. 2) (e.g., information of speaker A (601) may not be stored in memory (220)).
[0160] Referring to FIG. 9, the processor (250) can receive an audio signal (910) (e.g., "私は火曜日が好きです。") through the microphone (230). The processor (250) can check whether a user interaction with a specific object (e.g., the fourth object (607a) representing the speaker D (607) corresponding to the audio signal (910)) among a plurality of objects (621) (e.g., the first object (601a) representing the speaker A (601), the second object (603a) representing the speaker B (603), the third object (605a) representing the speaker C (605), the fourth object (607a) representing the speaker D (607), the fifth object (609a) representing the speaker E (609), and the sixth object (611a) representing the speaker F (611)) is detected. If a user interaction with a specific object (e.g., the fourth object (607a) representing the speaker D (607) corresponding to the audio signal (910)) among the plurality of objects (621) is not detected, the processor (250) can check whether there is a speaker recognition model for which learning corresponding to the audio signal (910) is completed. As described above, the speaker recognition model of the speaker D (607) corresponding to the audio signal (910) may be in a state where learning is completed. If there is a speaker recognition model for which learning corresponding to the audio signal (910) is completed, the processor (250) can translate the audio signal (910) into a target language (e.g., Korean) based on the speaker recognition model of the speaker D (607) for which learning is completed. For example, the processor (250) can display the information (915) (e.g., D) of the speaker D (607) and the text (920) (e.g., "저는 화요일이 좋습니다.") translated into the target language (e.g., Korean) on the display (240). In this case, the processor (250) can map the information (915) (e.g., D) of the speaker D (607) and the text (920) (e.g., "저는 화요일이 좋습니다.") translated into the target language (e.g., Korean) and store them in the memory (220).
[0161] Referring to FIG. 10, the processor (250) can receive a voice signal (1010) (e.g., "I'm fine on Tuesday too") through a microphone (230). The processor (250) can determine whether user interaction is detected for a specific object (e.g., the third object (605a) representing speaker C (605) corresponding to the voice signal (1010)) among a plurality of objects (621) (e.g., a first object (601a) representing speaker A (601), a second object (603a) representing speaker B (603), a third object (605a) representing speaker C (605), a fourth object (607a) representing speaker D (607), a fifth object (609a) representing speaker E (609), and a sixth object (611a) representing speaker F (611)). If user interaction is not detected for a specific object among the plurality of objects (621) (e.g., a third object (605a) representing speaker C (605) corresponding to the voice signal (1010)), the processor (250) can check whether there is a trained speaker recognition model corresponding to the voice signal (1010). As previously discussed, the speaker recognition model of speaker C (605) corresponding to the voice signal (1010) may not be in a state where training is complete. If there is no trained speaker recognition model corresponding to the voice signal (1010), the processor (250) can translate the voice signal (1010) into a target language (e.g., Korean) based on a language translation model corresponding to the voice signal (1010). The processor (250) can confirm that the language of the voice signal (1010) is Korean, and accordingly, may not perform the operation of translating the voice signal (1010). In this case, the processor (250) may only perform the operation of converting the voice signal (1010) into text. The converted text may or may not be displayed on the display (240).
[0162] Referring to FIG. 11, the processor 250 can receive an audio signal 1110 (e.g., "I have other schedules on Tuesday.") through the microphone 230. The processor 250 can check whether a user interaction with a specific object (e.g., the fourth object 607a representing the speaker D 607 corresponding to the audio signal 910) among a plurality of objects 621 (e.g., the first object 601a representing the speaker A 601, the second object 603a representing the speaker B 603, the third object 605a representing the speaker C 605, the fourth object 607a representing the speaker D 607, the fifth object 609a representing the speaker E 609, and the sixth object 611a representing the speaker F 611) is detected. If a user interaction with a specific object (e.g., the sixth object 611a representing the speaker F 611 corresponding to the audio signal 1110) among the plurality of objects 621 is not detected, the processor 250 can check whether there is a speaker recognition model for which learning corresponding to the audio signal 1110 is completed. As described above, the speaker recognition model for the speaker F 611 corresponding to the audio signal 1110 may be in a state where learning is completed. If there is a speaker recognition model for which learning corresponding to the audio signal 1110 is completed, the processor 250 can translate the audio signal 1110 into a target language (e.g., Korean) based on the speaker recognition model of the speaker F 611 for which learning is completed. For example, the processor 250 can display the information 1115 (e.g., F) of the speaker F 611 and the text 1120 (e.g., "I have other schedules on Tuesday.") translated into the target language (e.g., Korean) on the display 240. In this case, the processor 250 can map and store the information 1115 (e.g., F) of the speaker F 611 and the text 1120 (e.g., "I have other schedules on Tuesday.") translated into the target language (e.g., Korean) in the memory 220.
[0163] Referring to FIG. 12, the processor (250) can receive a voice signal (1220) (e.g., how about 3 pm on Wednesday?) through the microphone (230) and determine whether user interaction (1210) is detected for one of a plurality of objects (621). For example, the user interaction (1210) may include a touch input or gesture input for selecting a first object (601a) representing Speaker A (601) corresponding to the voice signal (1220) among a plurality of objects (601a, 603a, 605a, 607a, 609a, and 611a) included in the user interface (621).
[0164] In one embodiment, when a user interaction (1210) is detected for a first object (601a) representing speaker A (601) corresponding to a voice signal (1220) among a plurality of objects (601a, 603a, 605a, 607a, 609a, and 611a) included in a user interface (621), the processor (250) (e.g., speaker recognition model learning module (251)) can learn a speaker recognition model for speaker A (601) based on the characteristics of speaker A (601) (e.g., frequency and / or amplitude of the voice signal (625)) and the language of speaker A (601) (e.g., English).
[0165] In one embodiment, a processor (250) (e.g., a speech recognition module (254)) can perform automatic speech recognition for the language (e.g., English) of speaker A (601) to convert the speech signal (1220) into text. A processor (250) (e.g., a language translation module (255)) can translate the converted text into a target language (e.g., Korean) using a language translation model (e.g., an English translation model) corresponding to the speech signal (1220). The processor (250) can display information (1240) (e.g., A) of speaker A (601) corresponding to the speech signal (625) on a display (240) along with the text (1230) (e.g., How about Wednesday at 3 PM?) translated into the target language (e.g., Korean). In this case, the processor (250) can store in memory (220) the information (1240) (e.g., A) of speaker A (601) translated into a target language (e.g., Korean) text (1230) (e.g., How about Wednesday at 3 PM?).
[0166] As described in FIGS. 3 to 12 according to various embodiments, the electronic device (101) can learn a speaker recognition model corresponding to a voice signal received through a microphone (230) when user interaction is detected. When the learning of the speaker recognition model is completed, the electronic device (101) can analyze the voice signal received through the microphone (230) using the learned speaker recognition model, even if no additional user interaction is detected, to recognize the speaker and provide translated data corresponding to the voice signal in real time. Accordingly, the electronic device (101) can support immediate communication between multiple speakers.
[0167] FIG. 13 is a flowchart illustrating a method for generating a summary according to one embodiment of the present disclosure.
[0168] In the following embodiments, each operation of FIG. 13 may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation of FIG. 13 may be changed, and at least two operations may be performed in parallel.
[0169] According to one embodiment, the 1305 operation and 1310 operation of FIG. 13 can be understood as being performed in a processor (e.g., processor (250) of FIG. 2) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0170] Referring to FIG. 13, the processor (250) can, in operation 1305, map information of at least one speaker among multiple speakers and translated text associated with at least one speaker and store them in memory (e.g., memory (220) of FIG. 2).
[0171] In one embodiment, a processor (250) (e.g., summary generation module (256) of FIG. 2) may generate a summary in operation 1310 based on information of at least one speaker stored in memory (220) and translated text related to at least one speaker. For example, the processor (250) (e.g., summary generation module (256) of FIG. 2) may perform a preprocessing operation (e.g., removing unnecessary parts or highlighting necessary parts) based on the translated text and then extract key keywords and topics. Based on the extracted key keywords and topics, the processor (250) (e.g., summary generation module (256) of FIG. 2) may generate a summary of conversation content between multiple speakers (e.g., meeting content).
[0172] In one embodiment, although not shown, the processor (250) may generate a summary that includes only the translated text, excluding the speaker's information, in the case of text among the translated texts that is not mapped to the speaker's information.
[0173] In one embodiment, although not shown, the processor (250) can store the generated summary in memory (220).
[0174] As seen in FIG. 13 according to various embodiments, the processor (250) can generate a summary of conversation content (e.g., meeting content) between multiple speakers based on information of at least one speaker stored in memory (220) and translated text related to at least one speaker. The generated summary can be shared with at least one speaker among the multiple speakers. The at least one speaker who receives the summary can intuitively and clearly recognize the conversation content through the summary of conversation content (e.g., meeting content) between the multiple speakers.
[0175] FIG. 14 is a flowchart illustrating a method for providing a translation function according to one embodiment of the present disclosure.
[0176] In the following embodiments, each operation of FIG. 14 may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation of FIG. 14 may be changed, and at least two operations may be performed in parallel.
[0177] According to one embodiment, operations 1405 through 1415 of FIG. 14 may be understood to be performed in a processor (e.g., processor (250) of FIG. 2) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0178] Referring to FIG. 14, the processor (250) can set at least one language to be translated and a target language in operation 1405. For example, the at least one language to be translated may be the language to be translated, and the target language may be the language output as a translation result.
[0179] In one embodiment, a processor (250) (e.g., the language recognition module (253) of FIG. 2) can determine, in operation 1410, whether the language of the received voice signal corresponds to at least one language to be translated when a voice signal is received through a microphone (e.g., the microphone (230) of FIG. 2). For example, the processor (250) (e.g., the language recognition module (253) of FIG. 2) can determine the language of the voice signal by analyzing the pattern of the voice signal received through the microphone (230). For example, the processor (250) (e.g., the language recognition module (253) of FIG. 2) can analyze the pattern of a specific language (e.g., Japanese, Chinese, or English), and if the pattern of the specific language is detected in the voice signal received through the microphone (230), it can determine a language translation model corresponding to the specific language.
[0180] In one embodiment, the processor (250) can, in operation 1415, translate the voice signal into a target language using a language translation model corresponding to the voice signal and display it on a display (e.g., the display (240) of FIG. 2) if the language of the received voice signal corresponds to at least one language to be translated.
[0181] FIG. 15 is a drawing for explaining a method of providing a translation function according to one embodiment of the present disclosure.
[0182] Referring to FIG. 15, a processor (e.g., processor (250) of FIG. 2) of an electronic device (e.g., electronic device (101) of FIG. 1) can detect at least one language information to be translated and an input that sets a target language. For example, the processor (250) <1510> As illustrated in FIG. 15, a user interface (1515) for setting at least one language information to be translated and a target language can be displayed on a display (e.g., the display (240) of FIG. 2). A processor (250) can detect an input setting at least one language information (1521) to be translated, e.g., Japanese, Chinese, and English, in a first area (1520) of the user interface (1515). When an input selecting to add language information to be translated, e.g., an object “Add” (1523), is detected, the processor (250) can add language information to be translated other than Japanese, Chinese, and English. A processor (250) can detect an input setting a target language in a second area (1530) of the user interface (1515). In FIG. 15 according to various embodiments, the target language is assumed to be “Korean” for the description.
[0183] In one embodiment, after setting at least one language information to be translated and a target language, the processor (250) can check whether the received voice signal corresponds to the set at least one language information to be translated when a voice signal is received through a microphone (e.g., the microphone (230) of FIG. 2). For example, the processor (250) can check whether the received voice signal corresponds to the set at least one language information to be translated by checking whether a pattern of a specific language (e.g., Japanese, Chinese, or English) is detected from the voice signal.
[0184] In one embodiment, when the received voice signal does not match at least one language information set for translation, the processor (250) may not perform the operation of translating the language information into the target language. When the received voice signal matches at least one language information set for translation, the processor (250) may translate the voice signal into the target language using the language translation model corresponding to the identified language information and display it on the display (240).
[0185] For example, as shown in <1550>, the processor (250) may receive a first voice signal (e.g., Hello. I will start the meeting.). The language information corresponding to the first voice signal is Korean, and the processor (250) may confirm that the language information (e.g., Korean) corresponding to the first voice signal does not match at least one language information (e.g., Japanese, Chinese, and English) set for translation. The processor (250) may not perform the translation of the first voice signal, convert the first voice signal into text, and display the converted text (1560) on the display (240).
[0186] For example, the processor (250) may receive a second voice signal (e.g., 本日 The second item on the agenda is as follows.). The language information corresponding to the second voice signal is Chinese, and the processor (250) may confirm that the language information (e.g., Chinese) corresponding to the second voice signal matches at least one language information (e.g., Japanese, Chinese, and English) set for translation. The processor (250) may translate the second voice signal into the target language and display the text translated into the target language on the display (240). For example, the text corresponding to the second voice signal (e.g., 本日 The second item on the agenda is as follows. The translated text (1565) in the target language (e.g., The second item on today's agenda is as follows.) can be displayed on the display (240) together with [[ID=]]. This is not restrictive, and the processor (250) may display only the translated text in the target language (e.g., The second item on today's agenda is as follows.) on the display (240).
[0187] For example, the processor (250) can receive a third voice signal (e.g., Please open the material I sent you yesterday.). The language information corresponding to the third voice signal is English, and the processor (250) can confirm that the language information (e.g., English) corresponding to the third voice signal conforms to at least one language information (e.g., Japanese, Chinese, and English) for which translation is to be performed and which is set. The processor (250) can translate the third voice signal into the target language and display the translated text in the target language on the display (240). For example, the translated text (1570) in the target language (e.g., Please open the document I sent you yesterday.) can be displayed on the display (240) together with the text corresponding to the third voice signal (e.g., Please open the material I sent you yesterday.). This is not restrictive, and the processor (250) may display only the translated text in the target language (e.g., Please open the document I sent you yesterday.) on the display (240).
[0188] As seen in FIGS. 14 and 15 according to various embodiments, the electronic device (101) can translate and provide only voice signals corresponding to the language that the user of the electronic device (101) wants to translate into the target language, and this translation and provision function can be usefully used in the listening mode.
[0189] FIG. 16 is a flowchart illustrating a method for providing a translation function according to one embodiment of the present disclosure.
[0190] In the following embodiments, each operation of FIG. 16 may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation of FIG. 16 may be changed, and at least two operations may be performed in parallel.
[0191] According to one embodiment, operations 1605 through 1620 of FIG. 16 can be understood as being performed in a processor (e.g., processor (250) of FIG. 2) of an electronic device (e.g., electronic device (101) of FIG. 1).
[0192] Referring to FIG. 16, the processor (250) can display the screen of the translation application on a display (e.g., the display (240) of FIG. 2) in operation 1605. The processor (250) can set a target language to be translated in operation 1610. According to one embodiment, the target language may be a language that is output as a translation result.
[0193] In one embodiment, when user interaction is detected in operation 1615, the processor (250) can identify a language translation model corresponding to a voice signal among a plurality of language translation modules based on a voice signal received through a microphone (e.g., the microphone (230) of FIG. 2). The user interaction according to one embodiment may be a trigger input for translating the voice signal received through the microphone (230) into a target language.
[0194] In one embodiment, the processor (250) can translate a voice signal into a target language using an identified language translation model in operation 1620 and display it on the display (240). For example, the processor (250) can convert the voice signal into text and translate the converted text into a target language using a language translation model corresponding to the voice signal. The processor (250) can display the text translated into the target language on the display (240).
[0195] In one embodiment, the processor (250) can store text translated into a target language in memory (e.g., memory (220) of FIG. 2).
[0196] FIG. 17 is a drawing for explaining a method of providing a translation function according to one embodiment of the present disclosure.
[0197] Referring to FIG. 17, the processor (e.g., processor (250) of FIG. 2) of an electronic device (e.g., electronic device (101) of FIG. 1) is, <1710> As illustrated in the figure, a screen (1715) of a translation application can be displayed on a display (e.g., the display (240) of FIG. 2). The screen (1715) of the translation application may include a first area (1720) where user interaction is detected to translate a voice signal received through a microphone (e.g., the microphone (230) of FIG. 2) and a second area (1730) for setting a target language to be translated.
[0198] In one embodiment, as shown in <1750>, the processor (250) can receive a first voice signal (e.g., Hello. I will start the meeting.). When the first voice signal is received, no user interaction may be detected in the first area (1720). In this case, the processor (250) may not perform translation of the first voice signal, convert the first voice signal into text, and display the converted text (1760) (e.g., Hello. I will start the meeting.) on the display (240).
[0199] In one embodiment, a user interaction (1755) can be detected in the first area (1720). In response to detecting the user interaction (1755), the processor (250) can identify a language translation model corresponding to the voice signal based on the voice signal (e.g., 本日 The second item on the agenda is as follows.) received through the microphone (230) among a plurality of language translation modules. The processor (250) can use the identified language translation model (e.g., Japanese translation model) to translate the voice signal (e.g., 本日 The second item on the agenda is as follows.) into a target language (e.g., Korean) and display it on the display (240). For example, the processor (250) can convert the voice signal into text, and translate the converted text (e.g., 本日 The second item on the agenda is as follows.) into the target language (e.g., Korean) using the language translation model corresponding to the voice signal. The processor (250) can display the converted text (e.g., 本日 The second item on the agenda is as follows.) together with the text (1765) translated into the target language (e.g., Korean) (e.g., The second item on today's meeting agenda is as follows.) on the display (240).
[0200] As seen in FIGS. 16 and 17 according to various embodiments, when user interaction is detected, the electronic device (101) can translate the voice signal into a target language using a language translation model corresponding to the voice signal received through the microphone (230) and provide it to the user. Accordingly, the user of the electronic device (101) can simply receive only the desired language as a translation result in a conversation between multiple speakers, and the load on the electronic device (101) can also be reduced.
[0201] A method for providing a translation function according to one embodiment of the present disclosure may include an operation of displaying a user interface that provides translation functions for a plurality of speakers on a display (240). A method for providing a translation function according to one embodiment may include an operation of receiving a first voice signal through a microphone (230) and, when a first user interaction is detected, learning a speaker recognition model for the characteristics of a speaker and language corresponding to the first voice signal based on the first voice signal. A method for providing a translation function according to one embodiment may include an operation of receiving a second voice signal through a microphone (230) and checking whether a second user interaction is detected. A method for providing a translation function according to one embodiment may include an operation of checking whether a speaker recognition model that has completed learning corresponding to the second voice signal exists if the second user interaction is not detected. A method for providing a translation function according to one embodiment may include an operation of translating the second voice signal into a target language based on the speaker recognition model that has completed learning corresponding to the second voice signal and displaying it on the display (240) if a speaker recognition model that has completed learning corresponding to the second voice signal exists.
[0202] A method for providing a translation function according to one embodiment may include an operation of mapping a first voice signal and information of a speaker corresponding to the first voice signal based on receiving a first voice signal and detecting a first user interaction. A method for providing a translation function according to one embodiment may include an operation of storing the mapped first voice signal and information of a speaker corresponding to the first voice signal in a memory (220).
[0203] A method for providing a translation function according to one embodiment may include an operation of learning a speaker recognition model for the characteristics of a speaker and language corresponding to the second voice signal based on the second voice signal when a second voice signal is received through a microphone (230) and a second user interaction is detected. The speaker corresponding to the first voice signal and the speaker corresponding to the second voice signal according to one embodiment may be the same or different.
[0204] A method for providing a translation function according to one embodiment may include an operation to check the language of the second voice signal when there is no speaker recognition model corresponding to the second voice signal. A method for providing a translation function according to one embodiment may include an operation to translate the second voice signal into a target language and display it on a display (240) based on a language translation model corresponding to the language of the second voice signal.
[0205] A method for providing a translation function according to one embodiment may include an operation of displaying a user interface on a display (240) that includes a plurality of objects representing a plurality of speakers. A first user interaction or a second user interaction according to one embodiment may include an input for selecting one of the plurality of objects. A speaker corresponding to a first voice signal according to one embodiment may be a speaker corresponding to the selected object.
[0206] A method for providing a translation function according to one embodiment may include an operation of learning a speaker recognition model for the characteristics of a speaker and language, and then translating a first voice signal into a target language using a language translation model corresponding to the language of the first voice signal. A method for providing a translation function according to one embodiment may include an operation of displaying information of a speaker corresponding to the first voice signal and text translated into the target language on a display (240).
[0207] A method for providing a translation function according to one embodiment may include an operation of mapping speaker information corresponding to a first voice signal and text corresponding to the first voice signal translated into a target language and storing them in a memory (220). A method for providing a translation function according to one embodiment may include an operation of mapping speaker information corresponding to a second voice signal and text corresponding to the second voice signal translated into a target language and storing them in a memory (220).
[0208] A method for providing a translation function according to one embodiment may include generating a summary of a conversation related to a plurality of speakers based on information of a speaker corresponding to a first voice signal stored in memory (220), text corresponding to the first voice signal translated into a target language, information of a speaker corresponding to a second voice signal, and text corresponding to the second voice signal translated into a target language.
[0209] A method for providing a translation function according to one embodiment may include an operation to check whether the learning of a speaker recognition model corresponding to a first voice signal has been completed. A method for providing a translation function according to one embodiment may include, when the learning of a speaker recognition model corresponding to a first voice signal has been completed, an operation to distinguishly display an object representing the speaker and at least one object representing at least one speaker among a plurality of speakers for which the learning of the speaker recognition model has not been completed.
[0210] A non-transitory computer-readable medium storing instructions that cause at least one processor (250) of an electronic device (101) according to one embodiment of the present disclosure to perform operations when executed by at least one processor (250) may perform an operation of displaying a user interface that provides translation functions for a plurality of speakers on a display (240). A non-transitory computer-readable medium storing instructions that cause at least one processor (250) to perform operations when executed by at least one processor (250) of an electronic device (101) according to one embodiment may perform an operation of learning a speaker recognition model for the characteristics of a speaker and language corresponding to the first voice signal based on the first voice signal when a first voice signal is received through a microphone (230) and a first user interaction is detected. A non-transient computer-readable recording medium storing instructions that cause at least one processor (250) to perform operations when executed by at least one processor (250) of an electronic device (101) according to one embodiment may perform an operation to check whether a second voice signal is received through a microphone (230) and whether a second user interaction is detected. A non-transient computer-readable recording medium storing instructions that cause at least one processor (250) to perform operations when executed by at least one processor (250) of an electronic device (101) according to one embodiment may perform an operation to check whether a learned speaker recognition model corresponding to the second voice signal exists when the second user interaction is not detected.A non-transient computer-readable recording medium storing instructions that cause at least one processor (250) of an electronic device (101) according to one embodiment to perform operations when executed by at least one processor (250) may enable the second voice signal to be translated into a target language and displayed on a display (240) based on a speaker recognition model that has completed learning corresponding to the second voice signal, if such model exists.
[0211] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.
[0212] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as “A or B,” “at least one of A and B,” “at least one of A or B,” “A, B, or C,” “at least one of A, B, and C,” and “at least one of A, B, or C” each may include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as “first,” “second,” or “first” or “second” may be used simply to distinguish a component from another corresponding component and do not limit the components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0213] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. According to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0214] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.
[0215] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0216] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
In the electronic device (101), Microphone (230); Display (240); At least one processor (250) including processing circuitry; and It includes a memory (220) that stores instructions, When the above instructions are executed individually or collectively by the at least one processor (250), the electronic device (101), A user interface that provides translation functions for multiple speakers is displayed on the display (240), and When a first voice signal is received through the microphone (230) and a first user interaction is detected, a speaker recognition model for the characteristics of the speaker and language corresponding to the first voice signal is learned based on the first voice signal, and Check whether a second voice signal is received through the microphone (230) and whether a second user interaction is detected, and If the above second user interaction is not detected, check whether there exists a speaker recognition model that has completed training corresponding to the above second voice signal, and An electronic device that, if there is a speaker recognition model that has completed learning corresponding to the second voice signal, translates the second voice signal into a target language based on the speaker recognition model that has completed learning and displays it on the display (240). In Article 1, When the above instructions are executed individually or collectively by the at least one processor (250), the electronic device (101), Based on the reception of the first voice signal and the detection of the first user interaction, the first voice signal and information of the speaker corresponding to the first voice signal are mapped, and An electronic device that stores the mapped first voice signal and information of the speaker corresponding to the first voice signal in the memory (220). In Article 1, When the above instructions are executed individually or collectively by the at least one processor (250), the electronic device (101), When a second voice signal is received through the microphone (230) and the second user interaction is detected, a speaker recognition model for the characteristics of the speaker and language corresponding to the second voice signal is learned based on the second voice signal. The speaker corresponding to the first voice signal and the speaker corresponding to the second voice signal are the same or different electronic devices. In Article 1, When the above instructions are executed individually or collectively by the at least one processor (250), the electronic device (101), If a speaker recognition model corresponding to the second voice signal does not exist, the language of the second voice signal is verified, and An electronic device that translates the second voice signal into the target language and displays it on the display (240) based on a language translation model corresponding to the language of the second voice signal. In Article 1, When the above instructions are executed individually or collectively by the at least one processor (250), the electronic device (101), A user interface including a plurality of objects representing the plurality of speakers is displayed on the display (240), and The first user interaction or the second user interaction includes an input for selecting one of the plurality of objects, and The speaker corresponding to the first voice signal is an electronic device that is the speaker corresponding to the selected object. In Article 1, When the above instructions are executed individually or collectively by the at least one processor (250), the electronic device (101), After training a speaker recognition model for the characteristics and language of the speaker, the first voice signal is translated into the target language using a language translation model corresponding to the language of the first voice signal, and Information of the speaker corresponding to the first voice signal and text translated into the target language are displayed on the display (240), and Information of the speaker corresponding to the first voice signal and text corresponding to the first voice signal translated into the target language are mapped and stored in the memory (220), and Information of the speaker corresponding to the second voice signal and text corresponding to the second voice signal translated into the target language are mapped and stored in the memory (220), and An electronic device that generates a summary of a conversation related to a plurality of speakers based on information of a speaker corresponding to a first voice signal stored in the memory (220), text corresponding to the first voice signal translated into the target language, information of a speaker corresponding to the second voice signal, and text corresponding to the second voice signal translated into the target language. In Article 1, When the above instructions are executed individually or collectively by the at least one processor (250), the electronic device (101), Based on the first voice signal, the characteristics of the speaker, including the frequency and amplitude of the first voice signal, are identified, and Based on the pattern of the first voice signal, the language is identified, and An electronic device for learning a speaker recognition model of a speaker corresponding to the first voice signal based on the frequency, amplitude, and language of the first voice signal. In Article 1, When the above instructions are executed individually or collectively by the at least one processor (250), the electronic device (101), Check whether the training of the speaker recognition model of the speaker corresponding to the first voice signal is completed, and An electronic device that, when the learning of a speaker recognition model corresponding to the first voice signal is completed, distinguishably displays an object representing the speaker and at least one object representing at least one speaker among the plurality of speakers for which the learning of the speaker recognition model is not completed. Regarding the method of providing translation functions, The action of displaying a user interface that provides translation functions for multiple speakers on a display (240); When a first voice signal is received through a microphone (230) and a first user interaction is detected, the operation of learning a speaker recognition model for the characteristics of the speaker and language corresponding to the first voice signal based on the first voice signal; An operation to check whether a second voice signal is received through the microphone (230) and whether a second user interaction is detected; If the above second user interaction is not detected, an operation to check whether there exists a speaker recognition model that has completed learning corresponding to the above second voice signal; and A method comprising, if there exists a speaker recognition model that has completed learning corresponding to the second voice signal, translating the second voice signal into a target language based on the speaker recognition model that has completed learning and displaying it on the display (240). In Article 9, An operation of mapping the first voice signal and the speaker information corresponding to the first voice signal based on the reception of the first voice signal and the detection of the first user interaction; The operation of storing the mapped first voice signal and the speaker information corresponding to the first voice signal in memory (220); and When a second voice signal is received through the microphone (230) and the second user interaction is detected, the method further includes an operation to learn a speaker recognition model for the characteristics of the speaker and language corresponding to the second voice signal based on the second voice signal. The speaker corresponding to the first voice signal and the speaker corresponding to the second voice signal are the same or different. In Article 9, If a speaker recognition model corresponding to the second voice signal does not exist, an operation to verify the language of the second voice signal; and A method further comprising the operation of translating the second voice signal into the target language and displaying it on the display (240) based on a language translation model corresponding to the language of the second voice signal. In Article 9, The operation further includes displaying a user interface comprising a plurality of objects representing the plurality of speakers on the display (240). The first user interaction or the second user interaction includes an input for selecting one of the plurality of objects, and A method in which the speaker corresponding to the first voice signal is the speaker corresponding to the selected object. In Article 9, After learning a speaker recognition model regarding the characteristics and language of the speaker, the operation of translating the first voice signal into the target language using a language translation model corresponding to the language of the first voice signal; The operation of displaying the speaker's information corresponding to the first voice signal and the text translated into the target language on the display (240); The operation of mapping speaker information corresponding to the first voice signal and text corresponding to the first voice signal translated into the target language and storing them in memory (220); and The operation of mapping speaker information corresponding to the second voice signal and text corresponding to the second voice signal translated into the target language and storing them in the memory (220); and A method further comprising the operation of generating a summary of a conversation related to the plurality of speakers based on information of a speaker corresponding to a first voice signal stored in the memory (220), text corresponding to the first voice signal translated into the target language, information of a speaker corresponding to the second voice signal, and text corresponding to the second voice signal translated into the target language. In Article 9, An operation to check whether the training of the speaker recognition model of the speaker corresponding to the first voice signal is completed; and A method further comprising, when the learning of a speaker recognition model corresponding to the first voice signal is completed, a distinguishing display of an object representing the speaker and at least one object representing at least one speaker among the plurality of speakers for which the learning of the speaker recognition model is not completed. In a non-transitory computer-readable medium storing instructions that cause the at least one processor (250) to perform operations when executed by at least one processor (250) of an electronic device (101), The action of displaying a user interface that provides translation functions for multiple speakers on a display (240); When a first voice signal is received through a microphone (230) and a first user interaction is detected, the operation of learning a speaker recognition model for the characteristics of the speaker and language corresponding to the first voice signal based on the first voice signal; An operation to check whether a second voice signal is received through the microphone (230) and whether a second user interaction is detected; If the above second user interaction is not detected, an operation to check whether there exists a speaker recognition model that has completed learning corresponding to the above second voice signal; and A computer-readable recording medium that, if there exists a speaker recognition model that has completed learning corresponding to the second voice signal, performs the operation of translating the second voice signal into a target language based on the speaker recognition model that has completed learning and displaying it on the display (240).