Electronic device, method, and non-transitory computer-readable storage medium for generating summary
The electronic device addresses the challenge of real-time speech translation by generating and audibly outputting summaries, enhancing user experience through simultaneous visual and auditory feedback.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2026-04-02
AI Technical Summary
Existing speech translation technologies face challenges in providing real-time and simultaneous visual and auditory outputs, leading to prolonged waiting times for users.
An electronic device generates a summary of translated text and provides it audibly while displaying the full translation, using a processor to convert voice signals into text, translate the text, and output a summary through a speaker.
Improves the real-time and simultaneous delivery of translation services by allowing users to visually check detailed translations and audibly receive key content summaries, reducing waiting times.
Smart Images

Figure KR2025013605_02042026_PF_FP_ABST
Abstract
Description
Electronic device, method, and non-transient computer-readable storage medium for generating a summary
[0001] Embodiments of the present disclosure relate to an electronic device, a method, and a non-transient computer-readable storage medium for generating a summary.
[0002] Interpretation or translation may be performed for a conversation between a first user (e.g., first language) and a second user (e.g., second language) using different languages. Here, interpretation may be the conversion of a voice signal formed in the first language into a voice signal formed in the second language, which is 'speech,' and translation may be the conversion of a voice signal formed in the first language into 'text' formed in the second language. Hereinafter, recognizing a voice signal to interpret or translate may all be referred to as 'speech translation.'
[0003] Recently, with the advancement of automatic speech recognition and machine translation technologies, electronic devices can perform speech translation technology (e.g., speech translation programs) that recognizes speech signals and automatically translates and outputs them.
[0004] The electronic device can perform 'voice translation' through a translation application and provide text and voice signals based on the 'voice translation' to the user. The electronic device can perform a language translation function based on a language detection model configured in correspondence with the translation application.
[0005] The information described above may be provided as background art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art in relation to the present disclosure.
[0006] According to one embodiment, the electronic device aims to generate a summary based on a translated sentence and provide the generated summary to a user in order to reduce the time that occurs from the time it receives a voice signal to be translated until it outputs a translated voice signal according to 'voice translation'.
[0007] In the case of a text translation based on 'voice translation', the electronic device can provide the user with visualized information through a display. While providing the text translation, the electronic device can generate a summary of the text translation. In the case of the generated summary, the electronic device can provide the user with auditory information through a speaker. According to one embodiment, the user can visually check the text translation while audibly checking the summary generated based on the text translation. The electronic device can provide an effective translation service to the user visually or audibly.
[0008] The technical tasks intended to be accomplished in this document are not limited to those mentioned above, and other technical tasks not mentioned will be clearly understood by those skilled in the art to which this document belongs from the description below.
[0009] According to one embodiment, the electronic device may include a processor comprising a display, a speaker, a microphone, and a processing circuit, and a memory for storing instructions. When the instructions are executed individually or collectively by the processor, the electronic device may receive a first voice signal corresponding to a first language through the microphone, convert the received first voice signal into a first text, translate the first text corresponding to the first language into a second text corresponding to a second language, and when a condition for generating a summary related to the translation of the second text is satisfied, generate a first summary based on the second text, generate a second voice signal corresponding to the second language based on the generated first summary, display the second text through the display, and output the generated second voice signal through the speaker.
[0010] A control method for an electronic device according to one embodiment may include: receiving a first voice signal corresponding to a first language through a microphone of the electronic device; converting the received first voice signal into a first text; translating the first text corresponding to the first language into a second text corresponding to a second language; generating a first summary based on the second text when a condition for generating a summary related to the translation of the second text is satisfied; generating a second voice signal corresponding to the second language based on the generated first summary; displaying the second text through a display of the electronic device; and outputting the generated second voice signal through a speaker of the electronic device.
[0011] According to one embodiment, a non-transient computer-readable storage medium (or computer program product) storing one or more programs for performing a method of generating a summary in an electronic device may be described. According to one embodiment, the one or more programs may include instructions that, when executed by a processor of the electronic device, perform the operation of receiving a first voice signal corresponding to a first language through a microphone of the electronic device, the operation of converting the received first voice signal into a first text, the operation of translating the first text corresponding to the first language into a second text corresponding to a second language, the operation of generating a first summary based on the second text when a condition for generating a summary related to the translation of the second text is satisfied, the operation of generating a second voice signal corresponding to the second language based on the generated first summary, the operation of displaying the second text through a display of the electronic device, and the operation of outputting the generated second voice signal through a speaker of the electronic device.
[0012] According to one embodiment, an electronic device can generate a summary based on a text translation based on the second language, in order to provide a first voice signal formed in a first language spoken by a first user to a second user by translating it into a second language. The electronic device can provide the text translation visually to the user through a display and can provide the generated summary audibly to the user through a speaker.
[0013] The electronic device can output a voice signal (e.g., a second voice signal) corresponding to a summary through a speaker while displaying a text translation through a display. The electronic device can provide a more accurate and effective translation service to the user. Since the electronic device outputs a summary of the text translation as a voice signal rather than the entire content of the text translation, the real-time and simultaneity of the translation service can be improved. Even if the speaking time of the first user is prolonged, the simultaneity of the translation service can be improved because the entire translation content is provided to the second user as a summary.
[0014] According to one embodiment, the user can easily identify key content (e.g., keywords included in the summary) while visually checking the entire translation, and can audibly check the summary of the entire translation. The user can receive an effective translation service.
[0015] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.
[0016] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0017] FIG. 1 is a block diagram of an electronic device in a network environment according to various embodiments of the present disclosure.
[0018] FIG. 2 is an embodiment in which the entire translation content according to one embodiment of the present disclosure is visually displayed, and a summary of the entire translation content is audibly output.
[0019] FIG. 3 is a block diagram of an electronic device according to one embodiment of the present disclosure.
[0020] FIG. 4 is a flowchart illustrating a method for visually displaying a translation and audibly outputting a summary of the translation according to one embodiment of the present disclosure.
[0021] FIG. 5a is an exemplary diagram illustrating a method of visually displaying a translation and a summary of the translation according to one embodiment of the present disclosure.
[0022] FIG. 5b is an example diagram illustrating a method in which a summary of a translation according to one embodiment of the present disclosure is not visually displayed but is output only audibly.
[0023] FIG. 6 is a flowchart illustrating a method for generating a summary for a translation when a summary generation condition according to one embodiment of the present disclosure is satisfied.
[0024] FIG. 7 is a flowchart illustrating a method for determining a summary level according to one embodiment of the present disclosure and generating a summary based on the determined summary level.
[0025] FIG. 8 is an example diagram showing that when an external audio device is connected via communication according to one embodiment of the present disclosure, a summary is output through the external audio device.
[0026] FIG. 9 is a time table illustrating the process of generating a summary based on a first voice signal by a first user according to one embodiment of the present disclosure.
[0027] FIG. 10 is an example diagram in which a translation for three languages is displayed according to one embodiment of the present disclosure, and a summary of the translation is audibly output.
[0028] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.
[0029] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or may communicate with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0030] The processor (120) can control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., a program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0031] The auxiliary processor (123) can control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, an auxiliary processor (123) (e.g., an image signal processor or a communication processor (CP)) may be implemented as part of other functionally related components (e.g., a camera module (180) or a communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., a server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0032] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0033] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0034] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0035] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0036] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0037] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0038] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0039] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0040] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0041] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0042] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0043] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0044] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0045] A communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors (CP) that operate independently of a processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0046] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) may support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.
[0047] The antenna module (197) can transmit a signal or power to an external source (e.g., an external electronic device) or receive it from an external source. According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna, a first antenna, a second antenna). For example, the first antenna may generate a first antenna signal according to a first direction based on a linear polarization method, and the second antenna may generate a second antenna signal according to a second direction different from the first direction based on a linear polarization method. For example, the first antenna signal and the second antenna signal may be implemented in directions perpendicular to each other. If the first antenna signal is a communication signal according to the x-axis direction, the second antenna signal may include a communication signal according to the y-axis direction.
[0048] According to one embodiment, at least one antenna suitable for a communication method used in a communication network such as a first network (198) or a second network (199) may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0049] According to one embodiment, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0050] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0051] According to one embodiment, commands or data may be transmitted or received between an electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0052] FIG. 2 is an embodiment in which the entire translation content according to one embodiment of the present disclosure is visually displayed, and a summary of the entire translation content is audibly output.
[0053] Referring to FIG. 2, an electronic device (e.g., the electronic device (101) of FIG. 1) can execute a translation program (e.g., a translation application, an application that performs a translation function based on a language recognition model) to perform a translation function. Referring to FIG. 2, a processor of the electronic device (101) (e.g., the processor (120) of FIG. 1) can recognize a first voice signal (211) formed in a first language (e.g., English) and can translate the recognized first voice signal (211) into a second language (e.g., Korean, Hangul). For example, the processor (120) can generate a first text (201, 203) corresponding to the first voice signal (211) and a first text translation (202, 204) based on the first text (201, 203). The processor (120) can generate a first text translation (202, 204) corresponding to a second language (e.g., Korean) based on the first text (201, 203). According to one embodiment, the electronic device (101) may be in a state where a translation program is running in response to a situation where a translation function is required. For example, the processor (120) may run the translation program in response to the detection of the first language, or run the translation program in response to a user's selection (e.g., selecting the execution icon of the translation program).
[0054] According to one embodiment, the electronic device (101) can acquire a first voice signal (211) of a first language (e.g., English) using a microphone (e.g., input module (150) of FIG. 1) and can generate a first text (201, 203) corresponding to the first voice signal (211).
[0055] According to one embodiment, the electronic device (101) can display a first text (201, 203) and a first text translation (202, 204) through a display (e.g., the display module (160) of FIG. 1). The processor (120) can translate the first text (201, 203) based on a running translation program and can generate a first text translation (202, 204). The translation program can translate the first text (201, 203) in a first language (e.g., English) into a first text translation (202, 204) in a second language (e.g., Korean). According to one embodiment, the first language and the second language are not limited to a specific language and may be changed according to the settings of the translation program.
[0056] According to one embodiment, an electronic device (101) may primarily display a first-1 text (201) and secondarily display a first-1 text translation (202) in which the first-1 text (201) is translated. The first-1 text (201) and the first-1 text translation (202) may be displayed sequentially, and the translation of the first-1 text (201) may be recognized as the first-1 text translation (202). The user may find it convenient to understand the first-1 text translation (202) for the first-1 text (201).
[0057] According to one embodiment, the electronic device (101) can generate a first summary (221) based on a first-1 text translation (202) and a first-2 text translation (204), and can output the first summary (221) using a speaker (e.g., the acoustic output module (155) of FIG. 1). The processor (120) can generate a second voice signal corresponding to a second language (e.g., Korean) based on the first summary (221), and can output the second voice signal externally through a speaker. According to one embodiment, the electronic device (101) can convert a first voice signal corresponding to a first language (e.g., English) into a second voice signal corresponding to a second language (e.g., Korean), and can output the second voice signal through a speaker.
[0058] According to one embodiment, the electronic device (101) can output a first summary (221) generated based on the first text translation (202, 204) through a speaker while displaying together a first text (201, 203) corresponding to a first voice signal and a first text translation (202, 204) in which the first text (201, 203) is translated into a second language. The electronic device (101) can output a summary of the translation sentence audibly while visually displaying the entire translation sentence. According to one embodiment, the user can visually check the first text translation (202, 204) and substantially simultaneously audibly check the first summary (221) generated based on the first text translation (202, 204).
[0059] According to one embodiment, in a situation where the first voice signal is prolonged, the electronic device (101) may provide the first summary (221) summarized based on the first voice signal to the user as an audio signal, and the real-time (e.g., simultaneity) of the translation may be improved.
[0060] According to one embodiment, a translation application (e.g., a translation program) may include an automatic speech recognition module (ASR), a machine translation module (MT), and / or a text-to-speech (TTS) module. For example, the automatic speech recognition module (ASR) may perform the operation of converting a speech signal of a first language (e.g., a first speech signal) into text (e.g., a first text (201)). The processor (120) may convert the speech signal of the first language into text using at least one algorithm among a neural network model, dynamic time warping (DTW), and / or a hidden Markov model (HMM). For example, the machine translation module (MT) may perform the operation of translating a first text (201, 203) corresponding to the first language into a second language (e.g., a first text translation (202, 204)). The processor (120) can translate the first text (201, 203) into a second language using statistical machine learning or deep learning-based neural network modeling. The processor (120) can translate the first language into a second language using at least one algorithm among rule-based machine translation, statistical machine translation, and / or neural machine translation. For example, a speech synthesis module (TTS) can perform the operation of converting text corresponding to the second language (e.g., first text translation (202, 204)) into a voice signal (e.g., second voice signal) and outputting the converted voice signal.The processor (120) can perform the operation of converting the first text translation (202, 204) into a speech signal using at least one algorithm among a neural network model method, a unit selection method, and / or a statistical parametric synthesis method.
[0061] According to one embodiment, a translation application (e.g., a translation program) may include a summary generation module for generating a summary based on text. For example, the summary generation module may generate a summary using various summarization methods, such as extractive summarization and abstract summarization. For example, the summary generation module may generate a summary using an AI model (e.g., a large language model (LM)). The processor (120) may summarize text based on the LLM to match a set summary length. The processor (120) may generate a summary based on the first text translation (202, 204) under the control of the summary generation module. The processor (120) may generate a summary based on the first text translation (202, 204) in response to a situation where a summary generation condition is met.
[0062] According to one embodiment, in the operation of converting a first voice signal into a second voice signal, the electronic device (101) may generate a first summary (221) by at least partially summarizing a first text translation (202, 204) (e.g., a sentence in which a first text (201, 203) of a first language is translated into a second language) generated based on the first voice signal. The electronic device (101) may generate a second voice signal based on the first summary and may output the second voice signal through a speaker. The electronic device (101) may generate a second voice signal corresponding to a second language (e.g., Korean) based on a first voice signal corresponding to a first language (e.g., English) and may output the generated second voice signal through a speaker.
[0063] FIG. 3 is a block diagram of an electronic device according to one embodiment.
[0064] According to one embodiment, the electronic device (101) of FIG. 3 may be at least partially similar to the electronic device (101) of FIG. 1 and FIG. 2, or may include other embodiments of the electronic device (101). The electronic device (101) of FIG. 3 may have a translation program installed to perform a translation function (e.g., a translation application, an application that performs a translation function based on a language recognition model) and may perform a translation function based on the translation program.
[0065] Referring to FIG. 3, the electronic device (101) may include a processor (120) (e.g., processor (120) of FIG. 1), a memory (130) (e.g., memory (130) of FIG. 1), a display module (160) (e.g., display module (160) of FIG. 1), a communication circuit (190) (e.g., communication module (190) of FIG. 1), a microphone (310) (e.g., input module (150) of FIG. 1) and / or a speaker (320) (e.g., sound output module (155) of FIG. 1). Information (311) related to summary generation conditions may be stored in the memory (130) of the electronic device (101). The processor (120) of the electronic device (101) may generate a summary for text under the control of a summary generation unit (330). According to one embodiment, the processor (120) may be operatively, functionally, and / or electrically connected to a memory (130), a display module (160), a microphone (310), a speaker (320), and / or a communication circuit (190).
[0066] According to one embodiment, a processor (120) of an electronic device (101) can execute a program (e.g., program (140) of FIG. 1, translation program, translation application) stored in memory (130) to control at least one other component (e.g., hardware or software component) and perform various data processing or operations. According to one embodiment, the processor (120) may include at least one processor including a processing circuit. The number of processors (120) may be one or more. For example, the processor (120) may have the structure of a multi-core processor such as a dual core, a quad core, or a hexa core. The processor (120) can control the operations of the electronic device (101) by executing instructions stored in memory (130). For example, the processor (120) may include a plurality of processors and may collectively perform a plurality of operations by dividing them among the plurality of processors. According to one embodiment, the memory (130) may store instructions that are executed individually or collectively by the processor (120) (e.g., at least one processor).
[0067] According to one embodiment, a translation application (e.g., a translation program) may include a summary generation module for generating a summary based on text. For example, the summary generation module may generate a summary using various summarization methods, such as extractive summarization and abstract summarization. For example, the summary generation module may generate a summary using an AI model (e.g., a large language model (LM)).
[0068] For example, LLM can refer to a language model based on an artificial neural network that has been trained on a large amount of text data through prior training. LLMs can contain far more parameters (e.g., over 10 billion) than existing general language models. LLMs can use transformer artificial neural network structures based on an attention mechanism.
[0069] For example, the attention mechanism is a technique that helps artificial intelligence models focus on important parts within input data. The attention mechanism can predict the extent to which parts of time-series input data (e.g., input data such as voice or video, or input data for a specific layer of a neural network) contribute to the intermediate or final output of the neural network, and utilize this for the prediction of output data. For instance, while a recurrent neural network (RNN) structure that processes each element of a sequence sequentially may experience degraded prediction performance when there is information dependency over long time-series distances, the attention mechanism can account for information dependency over long time-series distances by controlling the degree of weighted attention within the overall (or partial) context of the input data.
[0070] For example, a Transformer can be composed of an encoder-decoder structure. The encoder can process input data to output compressed information (e.g., contextual representation), and the decoder can process the compressed information to output data in token units. Each of the encoder and decoder may include an independent attention network and a cross-attention network connecting the encoder and decoder.
[0071] For example, LLM learning may include pre-training and / or fine-tuning. Pre-training is a process that enables the LLM to acquire general linguistic knowledge using large amounts of text data, and may include, for example, self-supervised learning that predicts the next word based on the previous word sequence of a text sequence. Fine-tuning is a process of training the LLM to be suitable for a specific domain (e.g., chatbot, translation, summarization, Q&A) or task; based on the pre-trained model, the LLM can perform additional supervised learning (or adaptive learning) using a dataset tailored to the domain purpose. The LLM can perform tasks using text input containing natural language called a prompt.
[0072] For example, fine-tuning can be omitted during LLM training. To improve the performance of the task desired by the user, the prompts input into the LLM can be controlled. In a manner such as in-context learning or zero-shot / few-shot learning, examples of the task and / or guidance for performing the task may be additionally provided in the prompts. For example, the publicly available LLM may include BERT (bidirectional encoder representations from transformer) and GPT (generative pre-trained transformer).
[0073] For example, an LLM can receive additional input such as video information in addition to text. This video information can be converted into text through a separate pre-transformation (e.g., image recognition, scene recognition) and included in the prompt to generate a response. As another example, the input video can be converted into text-aligned image embeddings via an image encoder, and the input text can then use these text embeddings to generate a response using a separately trained model (e.g., a large multimodal model).
[0074] For example, the term 'LLM' can refer to the language neural network model itself, but it can also refer to models for LLM-based applications (e.g., chatbots, translation, summarization, text classification, sentence generation). For instance, an LLM-based chatbot or translator like ChatGPT can also be referred to as 'LLM'.
[0075] For example, 'LLM' may include an inference engine using an LLM neural network model. For example, "inputting an input prompt into the LLM" may mean "inputting an input prompt into an inference engine based on the LLM." For example, "the output of the LLM for the input prompt" may mean the output information of the last neural network layer of the LLM (or output information modified through additional processing) obtained when the input prompt is input into an inference engine based on the LLM.
[0076] According to one embodiment, a summary generation module can generate a summary using an AI model (e.g., a large language model (LM)). Setting information related to the length of the summary (e.g., “Summarize to 10 characters or less,” limit information related to the length of the summary) can be entered into the LLM input prompt. When the processor (120) generates a summary using the LLM, it can check the length of the summary based on the setting information stored in the LLM input prompt and generate the summary according to the length of the summary.
[0077] According to one embodiment, in the memory (130) of the electronic device (101), information (311) related to summary generation conditions for determining whether to generate a summary in the process of generating a summary based on text may be stored. For example, the processor (120) may determine whether to summarize the translated text based on the information (311) related to the summary generation conditions. If it is decided to generate a summary for the translated text, the processor (120) may determine a summary level for the summary based on the information (311) related to the summary generation conditions. For example, the summary level may be a summary standard in which the summary level (e.g., summary strength, summary rate) is divided into stages when generating the summary. A high summary level may mean that the summary strength is strong, and a low summary level may mean that the summary strength is low. For example, when summarizing about 5 sentences, if the summary level is level 3 (e.g., high summary level), about 5 sentences can be summarized into about 1-2 sentences. When summarizing about 5 sentences, if the summary level is level 1 (e.g., low summary level), about 5 sentences can be summarized into about 3-4 sentences.
[0078] According to one embodiment, a summary generation unit (330) included in the processor (120) can determine whether to generate a summary based on information (311) related to the summary generation conditions. When the generation of the summary is determined, the summary generation unit (330) can determine a summary level for the summary based on the information (311) related to the summary generation conditions.
[0079] According to one embodiment, the summary generation condition may include at least one of a first condition in which, while a first text translation (e.g., a first-1 text translation generated based on a first-1 voice signal) is confirmed, there exists a first voice signal corresponding to a first language that is additionally input (e.g., a first-2 voice signal, a first-2 voice signal input within a set time after the first-1 voice signal is input), a second condition in which there exists a second text translation to be converted into a second voice signal, and / or a third condition in which there exists a third voice signal waiting to be output.
[0080] For example, the electronic device (101) may convert a first-1 voice signal corresponding to a first language into first-1 text, translate the first-1 text into a second language, and generate a first-1 text translation corresponding to the second language. The electronic device (101) may determine whether to generate a summary when the first-1 text translation is generated. For example, the processor (120) may check for the existence of an additional first-2 voice signal within a set time after the first-1 voice signal is input in response to the generation of the first-1 text translation, and if the existence of the first-2 voice signal is confirmed, a condition for generating a summary (e.g., a first condition) may be satisfied.
[0081] In another example, the processor (120) can check whether there is a second-1 text translation to be converted into a second-1 voice signal in response to the generation of a first-1 text translation, and if the existence of the second-1 text translation is confirmed, a condition for generating a summary (e.g., a second condition) can be satisfied.
[0082] In another example, the processor (120) may satisfy a summary generation condition (e.g., a third condition) when the existence of a second-second voice signal waiting to be output is confirmed in response to the generation of a first-first text translation.
[0083] According to one embodiment, the electronic device (101) can generate a first summary based on the first text translation when a situation requiring a summary is identified (e.g., a situation where an additional voice signal to be translated is identified, a situation where a voice signal waiting for output is identified, a situation where a condition for generating a summary is satisfied) is identified while the first text translation is generated.
[0084] According to one embodiment, the conditions for determining the summary level may include at least one of the following: a condition in which, with the first text translation confirmed, the length of the additionally input first voice signal (e.g., number of waveform samples, time, number of voice files) exceeds a specified length; a condition in which the length of the first text waiting to be translated (e.g., number of sentences, number of words, number of characters) exceeds a specified length; a condition in which the length of the second text translation to be converted into the second voice signal (e.g., number of sentences, number of words, number of characters) exceeds a specified length; and / or a condition in which the length of the second voice signal waiting to be output (e.g., number of waveform samples, time, number of voice files) exceeds a specified length. For example, the length of the first voice signal that is additionally input may be referred to as “ASR_waiting_voice_length”, the length of the first text waiting for translation may be referred to as “MT_waiting_text_length”, the length of the second text translation to be converted into the second voice signal may be referred to as “TTS_waiting_text_length”, and the length of the second voice signal waiting for output may be referred to as “speaker_output_waiting_voice_length”.
[0085] For example, the electronic device (101) may be in a state of converting a first-1 voice signal corresponding to a first language into first-1 text and translating the first-1 text into a second language to generate a first-1 text translation corresponding to the second language. The electronic device (101) may be in a state of deciding to generate a first summary based on the first-1 text translation. For example, in response to the generation of the first summary, the processor (120) may determine a summary level for the summary if the length of a first-2 voice signal additionally input after the first-1 voice signal exceeds a specified length. As another example, in response to the generation of the first summary, the processor (120) may determine a summary level for the summary if the length of the first-2 text waiting for translation exceeds a specified length. In another example, the processor (120) may determine a summary level for the summary in response to the generation of the first summary if the length of the second-1 text translation to be converted into the second-1 voice signal exceeds a specified length. In another example, the processor (120) may determine a summary level for the summary in response to the generation of the first summary if the length of the second-2 voice signal waiting to be output exceeds a specified length.
[0086] According to one embodiment, the determination of the summary level may be made by selecting one of a plurality of summary levels, and depending on the determination condition of the summary level, the word count condition of the summary result (e.g., maximum word count, minimum word count, average word count, target word count) or the sentence count condition (e.g., maximum sentence count, minimum sentence count, average sentence count, target sentence count) may be directly determined. The electronic device (101) may also predict the word count of the summary result by applying the following (Equation 1).
[0087]
[0088] As another example, the electronic device (101) may predict the number of sentences of the summary result. According to one embodiment, the electronic device (101) may generate a summary based on a word count condition or a sentence count condition.
[0089] The electronic device (101) can determine a summary level based on the determination conditions of the summary level and can generate a summary based on the determined summary level.
[0090] According to one embodiment, the electronic device (101) can determine a summary level for the first summary when the generation of the first summary is determined, and can generate the first summary based on the determined summary level. For example, the processor (120) can determine whether the length of each of the voice signal, text, and text translation to be processed exceeds a specified length, and can determine a summary level (e.g., summary strength, summary rate) based on the degree to which it exceeds the specified length (e.g., difference value). The specified length may be set as a threshold. According to one embodiment, the processor (120) can check the difference value between the length and the set threshold in response to the length of each of the voice signal, text, and text translation to be processed exceeding the set threshold, and can determine a summary level based on the difference value. For example, the larger the difference value, the higher the summary level (e.g., summary strength, summary rate). This may mean that the longer the sentence (e.g., text, translation), the higher the level of summary.
[0091] According to one embodiment, a display module (or display (160)) of an electronic device (101) can visually provide content processed by a processor (120). The display module (160) may include a display. For example, the processor (120) can visually display text translated based on a translation program through the display module (160). A user can check the translation corresponding to the visual information through the display module (160). When displaying text through the display module (160), the processor (120) may display the text by applying a highlight effect to at least a portion of the text. For example, the processor (120) may display the text by changing at least one of the color, size, and font of some words included in the text.
[0092] According to one embodiment, the electronic device (101) may be operationally or functionally connected to an external electronic device using a communication module (190). For example, when the electronic device (101) is connected to an external audio device, the processor (120) may transmit a second voice signal to the external audio device and may output the second voice signal through the speaker of the external audio device.
[0093] According to one embodiment, the microphone (310) of the electronic device (101) can function as an input module (e.g., the input module (150) of FIG. 1) for acquiring an external audio signal (e.g., a voice signal). For example, the processor (120) can use the microphone (310) to acquire a first voice signal corresponding to a first language.
[0094] According to one embodiment, the speaker (320) of the electronic device (101) can function as an acoustic output module (e.g., the acoustic output module (155) of FIG. 1) for outputting an audio signal (e.g., a voice signal). For example, the processor (120) can use the speaker (320) to output a second voice signal corresponding to a second language to an external environment. A user can check the second voice signal corresponding to auditory information output through the speaker (320).
[0095] According to one embodiment, the electronic device (101) may be in a state where a translation program (e.g., a translation application, an application that performs a translation function based on a language recognition model) for performing a translation function is running. A second user (e.g., uses a second language) who is a user of the electronic device (101) may be conversing with a first user (e.g., uses a first language) who uses a different language. The processor (120) of the electronic device (101) may receive a first voice signal corresponding to the first language (e.g., the language of the first user) using a microphone (310). The processor (120) may convert the first voice signal into a first text based on the translation program and may translate the first text of the first language into a second text of the second language (e.g., a first text translation). In response to the generation of the first text translation, the processor (120) may determine whether a condition for generating a summary is satisfied. For example, the processor (120) can determine whether the summary generation condition is satisfied based on information (311) related to the summary generation condition of the memory (130). If the summary generation condition is satisfied, the processor (120) can generate a first summary based on the second text (e.g., first text translation) and can generate a second voice signal corresponding to a second language based on the generated first summary. For example, the first summary can be generated based on at least one word (e.g., keyword) included in the second text (e.g., first text translation). The processor (120) can output the second voice signal through the speaker (320) substantially simultaneously while displaying the second text (e.g., first text translation) through the display module (160).
[0096] According to one embodiment, the electronic device (101) can output a first summary generated based on the first text translation through a speaker (320), while displaying together a first text corresponding to a first language and a first text translation (e.g., a second text) in which the first text is translated into a second language through a display module (160). The electronic device (101) can output a summary of the translation visually while displaying the entire content of the translation sentence audibly. According to one embodiment, the user can visually check the first text translation (e.g., the second text) and substantially simultaneously audibly check the first summary generated based on the first text translation.
[0097] According to one embodiment, the electronic device (101) can provide a first summary based on the first voice signal to the user as an audio signal in a situation where the first voice signal is prolonged. Even if the first voice signal is prolonged, the output time of the audio signal (e.g., the first summary) is reduced, so the delay time in the translation process can be reduced. According to one embodiment, real-time (e.g., simultaneity) according to the translation situation can be improved. For example, a second voice signal corresponding to the summary can be output to the second user in accordance with the situation where the first user speaks the first voice signal. In a situation where the first user and the second user are conversing, real-time (e.g., simultaneity) according to the progress of the conversation can be improved.
[0098] According to one embodiment, an electronic device (101) may include a display (160), a speaker (320), a microphone (310), a processor (120) including a processing circuit, and a memory (130) for storing instructions. When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) can receive a first voice signal corresponding to a first language through the microphone (310), convert the received first voice signal into a first text, translate the first text corresponding to the first language into a second text corresponding to a second language, and when a condition for generating a summary related to the translation of the second text is satisfied, generate a first summary based on the second text, generate a second voice signal corresponding to the second language based on the generated first summary, display the second text through the display (160), and output the generated second voice signal through the speaker (320).
[0099] According to one embodiment, when the instructions are executed individually or collectively by the processor (120), the electronic device (101) checks whether there is an additionally received first voice signal in response to the translation into the second text, and if there is an additionally received first voice signal, the summary generation condition can be satisfied.
[0100] According to one embodiment, when the instructions are executed individually or collectively by the processor (120), the electronic device (101) can determine a summary level related to the generation of a summary based on the length of the first voice signal when the summary generation condition is satisfied, and generate the first summary based on the determined summary level.
[0101] According to one embodiment, when the instructions are executed individually or collectively by the processor (120), the electronic device (101) checks whether there is a second text being converted into a second voice signal in response to the translation into the second text, and if there is a second text being converted into a second voice signal, the summary generation condition can be satisfied.
[0102] According to one embodiment, when the instructions are executed individually or collectively by the processor (120), the electronic device (101) can determine a summary level related to the generation of a summary based on the length of the second text when the summary generation condition is satisfied, and generate the first summary based on the determined summary level.
[0103] According to one embodiment, when the instructions are executed individually or collectively by the processor (120), the electronic device (101) checks whether there is a second voice signal waiting to be output in response to the translation into the second text, and if there is a second voice signal waiting to be output, the summary generation condition can be satisfied.
[0104] According to one embodiment, when the instructions are executed individually or collectively by the processor (120), the electronic device (101) can determine a summary level related to the generation of a summary based on the length of the second voice signal waiting to be output when the summary generation condition is satisfied, and generate the first summary based on the determined summary level.
[0105] According to one embodiment, when the instructions are executed individually or collectively by the processor (120), the electronic device (101) can identify a keyword that matches a word included in the first summary based on the second text, apply a highlight effect to the identified keyword in the second text, and display the second text with the highlight effect applied to the keyword through the display (160).
[0106] According to one embodiment, the electronic device (101) may further include a communication circuit (190). When the instructions are executed individually or collectively by the processor (120), the electronic device (101) may be connected to an external audio device through the communication circuit (190) and may output the generated second voice signal through the external audio device.
[0107] According to one embodiment, when the instructions are executed individually or collectively by the processor (120), the electronic device (101) receives a third voice signal corresponding to a third language through the microphone (310), converts the received third voice signal into a third text, translates the third text corresponding to the third language into a fourth text corresponding to the second language, and when the condition for generating the summary related to the translation of the third text is satisfied, generates a second summary based on the fourth text, generates a fourth voice signal corresponding to the second language based on the generated second summary, displays the second text and the fourth text together through the display (160), and outputs the generated fourth voice signal through the speaker (320).
[0108] According to one embodiment, when the instructions are executed individually or collectively by the processor (120), the electronic device (101) can identify a first keyword that matches a word included in the first summary based on the second text, apply a first highlight effect to the identified first keyword in the second text, identify a second keyword that matches a word included in the second summary based on the third text, apply a second highlight effect to the identified second keyword in the third text, and display the second text with the first highlight effect and the third text with the second highlight effect through the display (160).
[0109] FIG. 4 is a flowchart illustrating a method for visually displaying a translation and audibly outputting a summary of the translation according to one embodiment of the present disclosure.
[0110] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0111] The electronic device (101) of FIG. 4 may be at least partially similar to the electronic device (101) of FIG. 3, or may include other embodiments of the electronic device (101). Operations 401 through 413 may be understood to be performed by a processor (e.g., processor (120) of FIG. 1 and 3) of the electronic device (e.g., electronic device (101) of FIG. 1 and 3).
[0112] According to one embodiment, the electronic device (101) has a translation program (e.g., an application related to a translation function) stored in a memory (e.g., the memory (130) of FIG. 3), and when performing a translation function, it can translate a first language (e.g., the first language, source language, English of FIG. 2) into a second language (e.g., the second language, target language, Korean of FIG. 2) based on the translation program. According to one embodiment, the processor of the electronic device (101) (e.g., the processor (120) of FIG. 3) can translate and convert a first voice signal corresponding to the first language used by a first user into a second voice signal corresponding to the second language used by a second user (e.g., a user of the electronic device (101)).
[0113] Referring to FIG. 4, a second user (e.g., using a second language) who is a user of the electronic device (101) may be in a situation where they are conversing with a first user (e.g., using a first language) who uses a different language. The second user (e.g., user of the electronic device (101)) may be in a situation where they want to obtain the content of the first user's conversation (e.g., a first voice signal) by translating it into a second voice signal corresponding to the second language. The electronic device (101) can translate and convert the first voice signal spoken by the first user into a second voice signal corresponding to the second language, and can provide the second voice signal audibly to the second user.
[0114] According to one embodiment, in operation 401, the processor (120) may receive a first voice signal corresponding to a first language. For example, the electronic device (101) may be in a state where a translation application for translating the first language into a second language is running. To acquire a first voice signal corresponding to a first language, the processor (120) may at least partially activate a microphone (e.g., the microphone (310) of FIG. 3). The processor (120) may acquire a first voice signal corresponding to a first language by using the microphone (310). For example, the processor (120) may perform preprocessing (e.g., voice activity detection (VAD) or end-point detection (EPD)) on the sound input through the microphone (310) and determine a first voice signal corresponding to a voice segment. The first voice signal may include one or more voice segments among the entire input sound.
[0115] According to one embodiment, in operation 403, the processor (120) can convert a first voice signal into a first text. For example, the processor (120) can convert the first voice signal into a first text based on an automatic speech recognition module (ASR) included in a translation application. According to one embodiment, if the first voice signal includes one or more voice segments, the processor (120) can convert the text in units of voice segments. When there are multiple converted texts, the processor (120) can recombine the converted texts to generate at least one sentence. The processor (120) can convert the first voice signal composed of at least one sentence into a first text.
[0116] According to one embodiment, in operation 405, the processor (120) can translate a first text corresponding to a first language into a second text corresponding to a second language. For example, the processor (120) can translate the first text corresponding to a first language into a second text corresponding to a second language based on a machine translation module (MT) included in a translation application. The electronic device (101) can display the first text and the second text, which is the first text translated into a second language, side by side through a display module (e.g., the display module (160) of FIG. 3).
[0117] According to one embodiment, in operation 407, the processor (120) may generate a first summary based on a second text in response to the fulfillment of a summary generation condition. For example, the summary generation condition may include at least one of the following: a condition in which a first voice signal corresponding to an additionally input first language exists while the first text translation is confirmed, a condition in which a second text translation to be converted into a second voice signal exists, and / or a condition in which a second voice signal awaiting output exists. For example, the summary generation condition may include a situation in which a summary is required (e.g., a situation in which an additional voice signal to be translated is confirmed, and / or a situation in which a voice signal awaiting output is confirmed). The processor (120) may extract at least one word (e.g., a keyword) contained in the second text and generate a first summary containing said at least one word. For example, the first summary may consist of fewer sentences than the sentence contained in the second text. The first summary may consist of fewer words than the number of words included in the second text.
[0118] According to one embodiment, in operation 407, the processor (120) may queue a request signal for a second text (e.g., a translation of the first text). For example, queuing may include the operation of lining up the multiple request signals in the order in which they are entered when multiple request signals are entered periodically or non-periodically. The processor (120) may line up the request signals sequentially. For example, the processor (120) may queue the request signal until the second voice signal being output through a speaker (e.g., the speaker (320) of FIG. 3) is terminated. As another example, the processor (120) may queue the request signal until the scheduled termination time of the second voice signal being output through the speaker (320) is less than or equal to a specified time. In operation 407, the processor (120) can check a second text corresponding to the queued request signal in response to the completion of outputting the second voice signal being output, and can generate a first summary based on the checked second text.
[0119] According to one embodiment, in operation 409, the processor (120) can generate a second voice signal corresponding to a second language based on a first summary. For example, the processor (120) can convert a second text corresponding to a second language into a second voice signal based on a speech synthesis module (TTS) included in a translation application.
[0120] According to one embodiment, in operation 411, the processor (120) can display a second text through a display module (160). In operation 407, since the first summary is generated based on at least one word included in the second text, the second text may include at least one word (e.g., keyword) constituting the first summary. The processor (120) can display a second text corresponding to a second language through the display module (160). When displaying the second text, the processor (120) can display it by reflecting an emphasis effect (e.g., highlight effect) in relation to at least one word included in the first summary.
[0121] According to one embodiment, in operation 413, the processor (120) can output a second voice signal through the speaker (320). According to one embodiment, while visually confirming the second text, the user can intuitively recognize at least one word with an emphasis effect and accurately recognize the second voice signal provided audibly.
[0122] According to one embodiment, in operation 413, while the processor (120) is outputting a second voice signal, if a request signal for an additional second voice signal in operation 409 is detected, the request signal may be queued. A first queuing operation related to the generation of a first summary in operation 407 and a second queuing operation related to the output of a second voice signal in operation 413 may be performed only one of the two, or both may be performed.
[0123] According to one embodiment, in operation 413, the processor (120) may or may not visually display the first summary through the display module (160) when outputting the second voice signal through the speaker (320). FIG. 5a below illustrates a situation in which the first summary is displayed through the display module (160), and FIG. 5b illustrates a situation in which the first summary is not displayed through the display module (160).
[0124] According to one embodiment, the electronic device (101) can translate a first voice signal corresponding to a first language into a second voice signal corresponding to a second language, and can provide text (e.g., text translated into the second language) and a voice signal (e.g., a second voice signal) according to the translation result to the user. For example, the text may be provided as visual information, and the voice signal may be provided as auditory information. In response to a situation where a condition for generating a summary is satisfied, the electronic device (101) can generate a summary (e.g., a first summary) based on the text according to the translation result and can generate a second voice signal corresponding to the summary. The electronic device (101) can provide the text according to the translation result to the user as visual information, while simultaneously providing the second voice signal corresponding to the summary to the user as auditory information.
[0125] FIG. 5a is an exemplary diagram illustrating a method in which a translation and a summary of the translation are visually displayed according to one embodiment of the present disclosure. FIG. 5b is an exemplary diagram illustrating a method in which a summary of the translation is not visually displayed but is output only audibly according to one embodiment of the present disclosure.
[0126] The electronic device (101) of FIGS. 5a and 5b may be at least partially similar to the electronic device (101) of FIG. 3, or may further include other embodiments of the electronic device (101). The electronic device (101) of FIGS. 5a and 5b has a translation program installed to perform a translation function (e.g., a translation application, an application that performs a translation function based on a language recognition model), and can perform a translation function based on the translation program.
[0127] Referring to FIGS. 5a and 5b, a second user (e.g., using a second language (e.g., Korean)) who is a user of the electronic device (101) may be conversing with a first user (e.g., using a first language (e.g., English)) who uses a different language. The processor (120) of the electronic device (101) may receive a first voice signal corresponding to the first language using a microphone (e.g., the microphone (310) of FIG. 3) and may convert the first voice signal into a first text (501). For example, the first text (501) may be displayed through a display module (e.g., the display module (160) of FIG. 3). The processor (120) may translate the first text (501) into a second text (502) based on a translation program and may display the second text (502).
[0128] The processor (120) can check the third text (503) after the first text (501) and can display the third text (503) through the display module (160). The processor (120) can translate the third text (503) into a fourth text (504) based on a translation program and can display the fourth text (504).
[0129] According to one embodiment, the processor (120) can generate a first summary (511) based on a second text (502) and a fourth text (504) corresponding to a second language. For example, when the second text (502) and the fourth text (504) are generated, the processor (120) can generate a first summary (511) based on the second text (502) and the fourth text (504) when a situation requiring a summary is identified (e.g., a situation where an additional voice signal to be translated is identified, a situation where a voice signal waiting for output is identified). The first summary (511) can be generated based on at least one word (e.g., a keyword) included in the second text (502) and the fourth text (504). According to one embodiment, the processor (120) can generate a second voice signal (511) corresponding to a second language based on the generated first summary (511). The processor (120) can output the generated second voice signal (511) through a speaker (e.g., the speaker (320) of FIG. 3).
[0130] Referring to FIG. 5a, the processor (120) can display the first summary (511) through the display module (160). The electronic device (101) can provide the first summary (511) (e.g., text) visually through the display module (160) and the first summary audibly through the speaker (320). Referring to FIG. 5a, the electronic device (101) can provide the first summary as an audio signal through the speaker (320) while displaying the entire conversation content and the first summary (511) through the display module (160). At least one word corresponding to the first summary (511) can be displayed with a visual emphasis effect (505, 506) applied. The first summary (511) can be generated based on at least one word with the emphasis effect (505, 506) applied.
[0131] Referring to FIG. 5b, the processor (120) may not display the first summary (521) through the display module (160). The electronic device (101) may provide the first summary (521) audibly through the speaker (320). Referring to FIG. 5b, the electronic device (101) may display the entire conversation through the display module (160) while providing the first summary (521) (e.g., voice signal) as an audio signal. At least one word corresponding to the first summary may be displayed in the entire conversation with a visual emphasis effect (505, 506) applied.
[0132] Referring to FIGS. 5a and 5b, the processor (120) may apply an emphasis effect (e.g., highlight effect) (505, 506) in relation to at least one word (e.g., keyword) included in the first summary when displaying the second text (502) and the fourth text (504). For example, the processor (120) may change at least one of the color, size, and font of the at least one word and may apply a visual effect to emphasize the at least one word.
[0133] FIG. 6 is a flowchart illustrating a method for generating a summary for a translation when a summary generation condition according to one embodiment of the present disclosure is satisfied.
[0134] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0135] The electronic device (101) of FIG. 6 may be at least partially similar to the electronic device (101) of FIG. 3, or may include other embodiments of the electronic device (101). Operations 601 through 617 may be understood as being performed by a processor (e.g., processor (120) of FIG. 1 and 3) of the electronic device (e.g., electronic device (101) of FIG. 1 and 3).
[0136] According to one embodiment, the electronic device (101) has a translation program (e.g., an application related to a translation function) stored in a memory (e.g., the memory (130) of FIG. 3), and when performing a translation function, it can translate a first language (e.g., the first language, source language, English of FIG. 2) into a second language (e.g., the second language, target language, Korean of FIG. 2) based on the translation program. According to one embodiment, the processor of the electronic device (101) (e.g., the processor (120) of FIG. 3) can translate and convert a first voice signal corresponding to the first language used by a first user into a second voice signal corresponding to the second language used by a second user (e.g., a user of the electronic device (101)).
[0137] According to one embodiment, operations 601 to 605 may operate in the same way as operations 401 to 405 of FIG. 4. The description of operations 601 to 605 may be replaced with the description of operations 401 to 405.
[0138] According to one embodiment, in operation 607, the processor (120) may determine whether a summary generation condition is satisfied. For example, the summary generation condition may include at least one of the following: a condition in which, with the second text confirmed, there exists a first voice signal corresponding to an additionally input first language; a condition in which there exists a third text to be converted into a second voice signal (e.g., a third text that was translated prior to the second text, a text waiting to be converted into a voice signal); and / or a condition in which there exists a third voice signal waiting to be output. For example, the processor (120) may be in a state in which it converts a first voice signal corresponding to the first language into a first text, translates the first text into a second language, and generates a second text corresponding to the second language. The electronic device (101) may determine whether to generate a summary when the second text is generated. For example, the processor (120) may check for the existence of a first-1 voice signal additionally input after the first voice signal in response to the generation of the second text, and if the existence of the additional first-1 voice signal is confirmed, the condition for generating a summary may be satisfied. In another example, the processor (120) may check for the existence of a third text to be converted into the second voice signal (e.g., a third text that was translated before the second text, text waiting to be converted into a voice signal) in response to the generation of the second text, and if the existence of the third text is confirmed, the condition for generating a summary may be satisfied. In another example, the processor (120) may check for the existence of a second-1 voice signal waiting to be output in response to the generation of the second text, and if the existence of the third text is confirmed, the condition for generating a summary may be satisfied.
[0139] According to one embodiment, if the condition for generating a summary in operation 607 is satisfied, the processor (120) may generate a first summary based on the second text in operation 609. Operations 609, 611, and 613 may operate identically to operations 407, 409, and 413 of FIG. 4. The description of operations 609, 611, and 613 may be replaced with the description of operations 407, 409, and 413.
[0140] According to one embodiment, if the condition for generating a summary in operation 607 is not satisfied, the processor (120) may generate a third voice signal based on the second text in operation 615. For example, if the condition for generating a summary is not satisfied, the processor (120) may not generate a first summary. In operation 615, the processor (120) may convert the entire conversation content corresponding to the first voice signal into a third voice signal based on the second text.
[0141] According to one embodiment, in operation 617, the processor (120) can output the third voice signal through the speaker (320).
[0142] FIG. 7 is a flowchart illustrating a method for determining a summary level according to one embodiment of the present disclosure and generating a summary based on the determined summary level.
[0143] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0144] The electronic device (101) of FIG. 7 may be at least partially similar to the electronic device (101) of FIG. 3, or may further include other embodiments of the electronic device (101). Operations 701 to 715 may be understood as being performed by a processor (e.g., processor (120) of FIG. 1 and 3) of the electronic device (e.g., electronic device (101) of FIG. 1 and 3).
[0145] According to one embodiment, the electronic device (101) has a translation program (e.g., an application related to a translation function) stored in a memory (e.g., the memory (130) of FIG. 3), and when performing a translation function, it can translate a first language (e.g., the first language, source language, English of FIG. 2) into a second language (e.g., the second language, target language, Korean of FIG. 2) based on the translation program. According to one embodiment, the processor of the electronic device (101) (e.g., the processor (120) of FIG. 3) can translate and convert a first voice signal corresponding to the first language used by a first user into a second voice signal corresponding to the second language used by a second user (e.g., a user of the electronic device (101)).
[0146] According to one embodiment, operations 701 to 705 may operate in the same way as operations 401 to 405 of FIG. 4. The description of operations 701 to 705 may be replaced with the description of operations 401 to 405.
[0147] According to one embodiment, in operation 707, the processor (120) can determine a summary level. For example, the summary level may be a summary criterion in which the level of summary (e.g., summary strength, summary rate) is divided into stages when generating a summary. A high summary level may mean that the summary strength is strong, and a low summary level may mean that the summary strength is low. For example, when summarizing about 5 sentences, if the summary level is level 3 (e.g., high summary level), about 5 sentences may be summarized into about 1-2 sentences. When summarizing about 5 sentences, if the summary level is level 1 (e.g., low summary level), about 5 sentences may be summarized into about 3-4 sentences. The processor (120) may determine the summary level for the summary based on information (311) related to the summary generation conditions stored in memory (e.g., memory (130) of FIG. 3). The processor (120) can determine the conditions for determining the summary level when determining the summary level. According to one embodiment, the electronic device (101) can adjust the summary level to minimize the gap (e.g., time difference) between the input voice signal and the output voice signal. For example, as the time difference between the input voice signal and the output voice signal increases, the summary level can be set higher. The electronic device (101) can set the summary level higher and shorten the summary so that the output voice signal (e.g., summary) is output faster. In another example, as the time difference between the input voice signal and the output voice signal decreases, the summary level can be set lower. The electronic device (101) can set the summary level lower and lengthen the summary so that the output voice signal (e.g., summary) is output slower. The electronic device (101) can determine the summary level in a way that improves real-timeness for the input voice signal and the output voice signal.
[0148] According to one embodiment, the conditions for determining the summary level may include at least one of the following: a condition in which, with the second text confirmed, the length of the additionally input first voice signal exceeds a specified length; a condition in which the length of the first text waiting to be translated exceeds a specified length; a condition in which the length of the third text to be converted into the second voice signal exceeds a specified length; and / or a condition in which the length of the second-1 voice signal waiting to be output exceeds a specified length. For example, the electronic device (101) may be in a state in which it converts a first voice signal corresponding to a first language into a first text, translates the first text into a second language, and generates a second text corresponding to the second language. The electronic device (101) may be in a state in which it has decided to generate a first summary based on the second text. For example, in response to the generation of the first summary, the processor (120) may determine a summary level for the summary if the length of the first-1 voice signal additionally input after the first voice signal exceeds a specified length. In another example, in response to the generation of the first summary, the processor (120) may determine a summary level for the summary if the length of the second-1 text waiting for translation exceeds a specified length. In another example, in response to the generation of the first summary, the processor (120) may determine a summary level for the summary if the length of the third text to be converted into the second voice signal exceeds a specified length. In another example, in response to the generation of the first summary, the processor (120) may determine a summary level for the summary if the length of the second-1 voice signal waiting for output exceeds a specified length.
[0149] According to one embodiment, the processor (120) can determine whether the length of each of the voice signal, text, and text translation to be processed exceeds a specified length, and can determine a summary level (e.g., summary strength, summary rate) based on the degree to which it exceeds the specified length (e.g., difference value). The specified length may be set as a threshold. According to one embodiment, the processor (120) can check the difference value between the length and the set threshold in response to the length of each of the voice signal, text, and text translation to be processed exceeding the set threshold, and can determine a summary level based on the difference value. For example, the greater the difference value, the higher the summary level (e.g., summary strength, summary rate). This may mean that the longer the sentence to be processed (e.g., text, translation), the higher the summary level.
[0150] According to one embodiment, in operation 709, the processor (120) can generate a first summary based on a determined summary level and a second text. For example, the higher the summary level, the fewer words or sentences the first summary may have. The processor (120) can determine the length of the first summary based on the determined summary level.
[0151] According to one embodiment, operations 711 to 715 may operate in the same way as operations 409 to 413 of FIG. 4. The description of operations 711 to 715 may be replaced with the description of operations 409 to 413.
[0152] FIG. 8 is an example diagram showing that when an external audio device is connected via communication according to one embodiment of the present disclosure, a summary is output through the external audio device.
[0153] The electronic device (101) of FIG. 8 may be at least partially similar to the electronic device (101) of FIG. 3, or may further include other embodiments of the electronic device (101). The electronic device (101) of FIG. 8 may have a translation program installed to perform a translation function (e.g., a translation application, an application that performs a translation function based on a language recognition model) and may perform a translation function based on the translation program. The operation performed in FIG. 8 may be understood as being performed by a processor (e.g., the processor (120) of FIG. 1 and FIG. 3) of the electronic device (e.g., the electronic device (101) of FIG. 1 and FIG. 3).
[0154] Referring to FIG. 8, the processor (120) can receive a first voice signal corresponding to a first language using a microphone (e.g., the microphone (310) of FIG. 3) and can convert the first voice signal into a first text (801, 803). For example, the first text (801, 803) can be displayed through a display module (e.g., the display module (160) of FIG. 3). The processor (120) can translate the first text (801, 803) into a second text (802, 804) based on a translation program and can display the second text (802, 804).
[0155] The processor (120) can generate a first summary (821) based on a second text (802, 804) corresponding to a second language. For example, the processor (120) can generate a first summary (821) based on the second text (802, 804) when, with the second text (802, 804) checked, a situation requiring a summary is confirmed (e.g., a situation where an additional voice signal to be translated is confirmed, a situation where a voice signal waiting for output is confirmed, a situation where a condition for generating a summary is satisfied). The first summary (821) can be generated based on at least one word (e.g., a keyword) included in the second text (802, 804). At least one word constituting the first summary (821) can be displayed with a highlighting effect reflected through the display module (160). The processor (120) can generate a second voice signal (821) corresponding to a second language based on the first summary (821) generated above.
[0156] Referring to FIG. 8, the processor (120) can check whether there is a connection with an external electronic device (e.g., an external audio device, wireless earphones) through a communication circuit (e.g., the communication circuit (190) of FIG. 3). If the processor (120) confirms a communication connection with the external audio device, it can transmit the second voice signal (821) to the external audio device so that the generated second voice signal (821) is output through the external audio device. The processor (120) can output the second voice signal (821) through the speaker of the external audio device.
[0157] FIG. 9 is a time table illustrating the process of generating a summary based on a first voice signal by a first user according to one embodiment of the present disclosure.
[0158] The electronic device (101) of FIG. 9 may be at least partially similar to the electronic device (101) of FIG. 3, or may include other embodiments of the electronic device (101). Operations 911 to 920 may be understood as being performed by a processor (e.g., processor (120) of FIG. 1 and 3) of the electronic device (e.g., electronic device (101) of FIG. 1 and 3).
[0159] According to one embodiment, the electronic device (101) has a translation program (e.g., an application related to a translation function) stored in a memory (e.g., the memory (130) of FIG. 3), and when performing a translation function, it can translate a first language (e.g., the first language, source language, English of FIG. 2) into a second language (e.g., the second language, target language, Korean of FIG. 2) based on the translation program. According to one embodiment, the processor of the electronic device (101) (e.g., the processor (120) of FIG. 3) can translate and convert a first voice signal corresponding to the first language used by a first user into a second voice signal corresponding to the second language used by a second user (e.g., a user of the electronic device (101)).
[0160] Referring to FIG. 9, the process performed in the translation program is divided into approximately four operations (901, 902, 903, 904), but is not limited thereto. Operation 901 includes an operation of converting a first voice signal into a first text and can be performed by an automatic speech recognition (ASR) module included in the translation program. Operation 902 includes an operation of translating a first text of a first language into a second text of a second language and can be performed by a machine translation (MT) module included in the translation program. Operation 903 includes an operation of generating a first summary based on the second text and can be performed by a summary generation module included in the translation program. Operation 904 includes an operation of converting the first summary into a second voice signal and outputting the second voice signal and can be performed by a text-to-speech (TTS) module included in the translation program.
[0161] Referring to FIG. 9, the processor (120) may continuously recognize a plurality of voice signals (e.g., sentences) corresponding to a first language. For example, the processor (120) may sequentially recognize a first-1 voice signal (e.g., sentence 1), a first-2 voice signal (e.g., sentence 2), and / or a first-3 voice signal (e.g., sentence 3) corresponding to the first language. The processor (120) may convert (901) the first-1 voice signal (e.g., sentence 1) into a first-1 text (911). The processor (120) may translate (902) the first-1 text (e.g., English) (911) into a second-1 text (e.g., Korean) (912). In response to the confirmation of the second-1 text (912), the processor (120) may determine whether the conditions for generating a summary are satisfied. A situation in which the conditions for generating a summary are not met, and thus a situation in which a summary is not generated. The 2-1 text (912, 913) may be retained. The processor (120) may convert the 2-1 text (913) into a 2-1 voice signal (914) and output the 2-1 voice signal (914) as an audio signal (904). The processor (120) may perform a translation process and a conversion process on the acquired 1st sentence and output the 2-1 voice signal (914) (904).
[0162] Referring to FIG. 9, in a situation where a 2-1 voice signal (914) is output (904), the processor (120) can recognize a 1-2 voice signal (e.g., a 2nd sentence) and a 1-3 voice signal (e.g., a 3rd sentence). The processor (120) can convert (901) the 1-2 voice signal (e.g., a 2nd sentence) into a 1-2 text (915) and translate (902) the 1-2 text (e.g., English) (915) into a 2-2 text (e.g., Korean) (916). Even after the translation (902) of the 2-2 text (916) is completed, the output of the 2-1 voice signal (914) may be in progress. In this case, the processor (120) can convert the first-third voice signal (e.g., the third sentence) into the first-third text (917) (901) and translate the first-third text (e.g., English) (917) into the second-third text (e.g., Korean) (918) (902). When the translation (902) of the second-third text (918) is completed, the output of the second-first voice signal (914) may be completed.
[0163] The processor (120) can determine whether the conditions for generating a summary are met in response to the confirmation of the 2-2 text (916) and the 2-3 text (918). A situation in which the conditions for generating a summary are met may be a situation in which a summary is generated. The processor (120) can generate a first summary (919) based on the 2-2 text (916) and the 2-3 text (918). The processor (120) can convert the first summary (919) into a 2-2 voice signal (920) and output the 2-2 voice signal (920) as an audio signal (904). The processor (120) can perform a conversion process, a translation process, and a summary process on the acquired 2-2 sentence and 3-3 sentence, and output the 2-2 voice signal (920) (904).
[0164] FIG. 10 is an example diagram in which a translation for three languages is displayed according to one embodiment of the present disclosure, and a summary of the translation is audibly output.
[0165] The electronic device (101) of FIG. 10 has a translation program installed to perform a translation function (e.g., a translation application, an application that performs a translation function based on a language recognition model) and can perform a translation function based on the translation program. The operation performed in FIG. 10 can be understood as being performed by a processor (e.g., the processor (120) of FIG. 1 and FIG. 3) of the electronic device (e.g., the electronic device (101) of FIG. 1 and FIG. 3).
[0166] Referring to FIG. 10, there may be a situation in which a first user using a first language (e.g., English), a second user using a second language (e.g., Spanish), and a third user using a third language (e.g., Korean) are conversing. The electronic device (101) may be a portable electronic device used by the third user.
[0167] Referring to FIG. 10, the processor (120) of the electronic device (101) can receive a first-1 voice signal corresponding to a first language (e.g., English) using a microphone (e.g., microphone (310) of FIG. 3) and can convert the first-1 voice signal into a first text (1001). The processor (120) can translate the first text (1001) into a second text (1002) based on a translation program and can display the first text (1001) and the second text (1002) through a display module (e.g., display module (160) of FIG. 3). Since the electronic device (101) is being used by a third user, the first text (1001) can be translated into a second text (1002) corresponding to a third language (e.g., Korean).
[0168] The processor (120) can generate a first summary (1011) based on a second text (1002) corresponding to a third language (e.g., Korean). For example, when the second text (1002) is generated and a situation requiring a summary is identified (e.g., a situation where an additional voice signal to be translated is identified, or a voice signal waiting for output is identified), the processor (120) can generate a first summary (1011) based on the second text (1002). The first summary (1011) can be generated based on at least one word (e.g., keyword) included in the second text (1002). The processor (120) can generate a second-1 voice signal corresponding to a third language (e.g., Korean) based on the generated first summary (1011). The processor (120) can output the generated second-1 voice signal through a speaker (e.g., the speaker (320) of FIG. 3).
[0169] Referring to FIG. 10, the processor (120) can receive a first-second voice signal corresponding to a second language (e.g., Spanish) using a microphone (310) and can convert the first-second voice signal into a third text (1003). The processor (120) can translate the third text (1003) into a fourth text (1004) based on a translation program and can display the third text (1003) and the fourth text (1004) through a display module (160). Since the electronic device (101) is being used by a third user, the third text (1003) can be translated into a fourth text (1004) corresponding to a third language (e.g., Korean).
[0170] The processor (120) can generate a second summary (1021) based on a fourth text (1004) corresponding to a third language (e.g., Korean). For example, when the fourth text (1004) is generated and a situation requiring a summary is identified (e.g., a situation where an additional voice signal to be translated is identified, or a voice signal waiting for output is identified), the processor (120) can generate a second summary (1021) based on the fourth text (1004). The second summary (1021) can be generated based on at least one word (e.g., a keyword) included in the fourth text (1004). The processor (120) can generate a second voice signal corresponding to a third language (e.g., Korean) based on the generated second summary (1021). The processor (120) can output the generated second-second voice signal through the speaker (320).
[0171] According to one embodiment, the electronic device (101) can perform a translation function even in a situation where three or more users speaking different languages are conversing. A first text (1001) spoken by a first user can be displayed with a first visual effect reflected, and a third text (1003) spoken by a second user can be displayed with a second visual effect reflected. The user can intuitively distinguish between the first text (1001) spoken by the first user and the third text (1003) spoken by the second user. The electronic device (101) can display a second text (1002) and a fourth text (1004) corresponding to the third language. For example, the second text (1002) can be displayed with an emphasis effect applied to at least one word corresponding to the first summary (1011). The fourth text (1004) may be displayed with an emphasis effect applied to at least one word corresponding to the second summary (1021). The first language used by the first user and the second language used by the second user may be translated into the third language of the third user. The electronic device (101) may apply different visual effects to each text so that, in a situation where three or more users are conversing, the user can distinguish each speaker.
[0172] A method for generating a summary of an electronic device (101) according to one embodiment may include: receiving a first voice signal corresponding to a first language through a microphone (310) of the electronic device (101); converting the received first voice signal into a first text; translating the first text corresponding to the first language into a second text corresponding to a second language; generating a first summary based on the second text when a summary generation condition related to the translation of the second text is satisfied; generating a second voice signal corresponding to the second language based on the generated first summary; displaying the second text through a display (160) of the electronic device (101); and outputting the generated second voice signal through a speaker (320) of the electronic device (101).
[0173] A method for generating a summary of an electronic device (101) according to one embodiment may further include: an operation to check whether there is an additionally received first voice signal in response to translation into the second text; an operation to satisfy the summary generation condition if the additionally received first voice signal exists; an operation to determine a summary level related to the generation of a summary based on the length of the first voice signal if the summary generation condition is satisfied; and an operation to generate the first summary based on the determined summary level.
[0174] A method for generating a summary of an electronic device (101) according to one embodiment may further include, in response to translation into the second text, an operation to check whether there is a second text being converted into the second voice signal, an operation to satisfy the summary generation condition if there is a second text being converted into the second voice signal, an operation to determine a summary level related to the generation of a summary based on the length of the second text if the summary generation condition is satisfied, and an operation to generate the first summary based on the determined summary level.
[0175] A method for generating a summary of an electronic device (101) according to one embodiment may further include, in response to translation into the second text, an operation to check whether there is a second voice signal waiting to be output, an operation to satisfy the summary generation condition if the second voice signal waiting to be output exists, an operation to determine a summary level related to the generation of a summary based on the length of the second voice signal waiting to be output if the summary generation condition is satisfied, and an operation to generate the first summary based on the determined summary level.
[0176] According to one embodiment, the operation of displaying the second text may include, based on the second text, an operation of identifying a keyword that matches a word included in the first summary, an operation of reflecting a highlight effect on the identified keyword in the second text, and an operation of displaying the second text with the highlight effect reflected on the keyword through the display (160).
[0177] A method for generating a summary of an electronic device (101) according to one embodiment may further include: receiving a third voice signal corresponding to a third language through the microphone (310); converting the received third voice signal into a third text; translating the third text corresponding to the third language into a fourth text corresponding to the second language; generating a second summary based on the fourth text when the summary generation condition related to the translation of the third text is satisfied; generating a fourth voice signal corresponding to the second language based on the generated second summary; displaying the second text and the fourth text together through the display (160); and outputting the generated fourth voice signal through the speaker (320).
[0178] A method for generating a summary of an electronic device (101) according to one embodiment may further include: an operation of identifying a first keyword that matches a word included in the first summary based on the second text; an operation of reflecting a first highlight effect on the identified first keyword in the second text; an operation of identifying a second keyword that matches a word included in the second summary based on the third text; an operation of reflecting a second highlight effect on the identified second keyword in the third text; and an operation of displaying the second text with the first highlight effect reflected and the third text with the second highlight effect reflected through the display (160).
[0179] According to one embodiment, a non-transient computer-readable storage medium (or, computer program product) storing one or more programs for performing a method of generating a summary in an electronic device (101) may be described. According to one embodiment, one or more programs may include instructions that, when executed by a processor (120) of an electronic device (101), perform the following operations: receiving a first voice signal corresponding to a first language through a microphone (310) of the electronic device (101); converting the received first voice signal into a first text; translating the first text corresponding to the first language into a second text corresponding to a second language; generating a first summary based on the second text when a condition for generating a summary related to the translation of the second text is satisfied; generating a second voice signal corresponding to the second language based on the generated first summary; displaying the second text through a display (160) of the electronic device (101); and outputting the generated second voice signal through a speaker (320) of the electronic device (101).
[0180] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.
[0181] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0182] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0183] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.
[0184] According to one embodiment, the method according to the various embodiments disclosed herein may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0185] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device (101), Display (160); Speaker (320); Microphone (310); A processor (120) including a processing circuit; and Includes memory (130) for storing instructions, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, A first voice signal corresponding to a first language is received through the microphone (310), and Converting the received first voice signal into first text, and Translate the first text corresponding to the first language into a second text corresponding to the second language, and If the conditions for generating a summary related to the translation of the above second text are satisfied, a first summary is generated based on the above second text, and Based on the first summary generated above, a second voice signal corresponding to the second language is generated, and The second text is displayed through the display (160), and An electronic device that outputs the generated second voice signal through the speaker (320).
2. In Paragraph 1, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, In response to the translation into the second text above, check whether there is an additionally received first voice signal, and An electronic device that satisfies the summary generation condition if the above additionally received first voice signal exists.
3. In Paragraph 2, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, When the above summary generation condition is satisfied, a summary level related to the generation of the summary is determined based on the length of the first voice signal, and An electronic device that generates the first summary based on the above-determined summary level.
4. In Paragraph 1, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, In response to the translation into the second text above, checking whether there exists a second text being converted into the second voice signal, and An electronic device that satisfies the summary generation condition if there is a second text being converted into the second voice signal.
5. In Paragraph 4, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, If the above summary generation condition is satisfied, the summary level related to the generation of the summary is determined based on the length of the above second text, and An electronic device that generates the first summary based on the above-determined summary level.
6. In Paragraph 1, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, In response to the translation into the second text above, check whether there is a second voice signal waiting to be output, and An electronic device that satisfies the summary generation condition if the second voice signal waiting for output exists.
7. In Paragraph 6, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, When the above summary generation condition is satisfied, a summary level related to the generation of the summary is determined based on the length of the second voice signal awaiting output, and An electronic device that generates the first summary based on the above-determined summary level.
8. In Paragraph 1, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, Based on the above second text, identify keywords that match words included in the above first summary, and Apply a highlight effect to the identified keyword in the second text above, and An electronic device that displays the second text with a highlight effect applied to the keyword through the display (160).
9. In Paragraph 1, It further includes a communication circuit (190); and When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, Through the above communication circuit (190), a communication connection is established with an external audio device, and An electronic device that outputs the generated second voice signal through the above external audio device.
10. In Paragraph 1, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, A third voice signal corresponding to a third language is received through the microphone (310), and Converting the received third voice signal into third text, and Translate the third text corresponding to the third language into the fourth text corresponding to the second language, and If the above summary generation condition related to the translation of the above third text is satisfied, a second summary is generated based on the above fourth text, and Based on the second summary generated above, a fourth voice signal corresponding to the second language is generated, and Through the display (160), the second text and the fourth text are displayed together, An electronic device that outputs the generated fourth voice signal through the speaker (320).
11. In Paragraph 10, When the above instructions are executed individually or collectively by the processor (120), the electronic device (101) is made to, Based on the above second text, identify a first keyword that matches a word included in the above first summary, and Applying the first highlight effect to the identified first keyword in the above second text, and Based on the above third text, identify a second keyword that matches a word included in the above second summary, and Reflecting the second highlight effect on the identified second keyword in the third text above, and An electronic device that displays the second text with the first highlight effect and the third text with the second highlight effect through the display (160).
12. A method for generating a summary of an electronic device (101), The operation of receiving a first voice signal corresponding to a first language through the microphone (310) of the electronic device (101); The operation of converting the received first voice signal into first text; The operation of translating the first text corresponding to the first language into the second text corresponding to the second language; If the condition for generating a summary related to the translation of the second text is satisfied, the operation of generating a first summary based on the second text; An operation to generate a second voice signal corresponding to the second language based on the first summary generated above; The operation of displaying the second text through the display (160) of the electronic device (101); and A method comprising the operation of outputting the generated second voice signal through the speaker (320) of the electronic device (101).
13. In Paragraph 12, An operation to check whether an additionally received first voice signal exists in response to the translation into the second text above; If the above additionally received first voice signal exists, an operation satisfying the above summary generation condition; When the above summary generation condition is satisfied, an operation to determine a summary level related to the generation of a summary based on the length of the first voice signal; and A method further comprising the operation of generating the first summary based on the above-determined summary level.
14. In Paragraph 12, An operation to check whether, in response to the translation into the second text, there exists a second text being converted into the second voice signal, or whether there exists a second voice signal waiting to be output; If there is a second text being converted into the second voice signal, or if there is a second voice signal waiting to be output, an operation satisfying the summary generation condition; When the above summary generation condition is satisfied, an operation to determine a summary level related to the generation of a summary based on the length of the second text, or the length of the second voice signal awaiting output; and A method further comprising the operation of generating the first summary based on the above-determined summary level.
15. A non-transient computer-readable storage medium storing one or more programs for performing a method of generating a summary in an electronic device (101), When one or more of the above programs are executed by the processor (120) of the electronic device (101), The operation of receiving a first voice signal corresponding to a first language through the microphone (310) of the electronic device (101); The operation of converting the received first voice signal into first text; The operation of translating the first text corresponding to the first language into the second text corresponding to the second language; If the condition for generating a summary related to the translation of the second text is satisfied, the operation of generating a first summary based on the second text; An operation to generate a second voice signal corresponding to the second language based on the first summary generated above; The operation of displaying the second text through the display (160) of the electronic device (101); and A computer-readable storage medium comprising instructions for performing the operation of outputting the generated second voice signal through the speaker (320) of the electronic device (101).
Citation Information
Patent Citations
Text summarization method and text summarization system
JP2023034235A
Interpretation system, interpretation method, and interpretation program
JP2024110057A
Heat dissipation composite material and reflector for projection lamp manufactured using same
KR1020220127992A
Artificial Intelligence-Based Security Event Analysis System and Its Method Using Semi-Supervised Machine Learning
KR102089688B1
Method For Manufacturing Gelatin Hydrogel And Gelatin Hydrogel Comprising Drug Which Is Adjustable Release Behavior
KR102876380B1