Electronic device and method for providing call summary

By employing an AI model to map speaker information and generate summaries, the electronic device accurately converts voice data into text and provides real-time summaries, addressing the challenge of distinguishing user and counterparty voices during voice calls.

WO2026155497A1PCT designated stage Publication Date: 2026-07-23SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2026-01-09
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing electronic devices face challenges in accurately distinguishing between user and counterparty voices during voice calls, leading to reduced recognition accuracy when converting recorded voice data into text, especially during simultaneous speaking.

Method used

The electronic device employs an AI model to map speaker information to text data, generating call text data by distinguishing between user and counterparty voices, and uses an on-device AI model to generate a summary of the call content.

Benefits of technology

This approach enhances the accuracy of voice-to-text conversion and provides a real-time summary of the call content, improving user experience and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2026000531_23072026_PF_FP_ABST
    Figure KR2026000531_23072026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to various embodiments described in the present document may comprise: a display; a microphone; a communication circuit; memory; and at least one processor. The memory may store instructions which, when executed by the at least one processor, cause the electronic device to establish a call with an external device, acquire first voice data of a first user of the electronic device through the microphone during the call, and acquire second voice data of a second user corresponding to the external device during the call on the basis of data received from the external device through the communication circuit. The memory may store instructions that cause the electronic device to generate call text data by mapping first speaker information corresponding to the first user to first text data converted from the first voice data and mapping second speaker information corresponding to the second user of the external device to second text data converted from the second voice data. Various other embodiments are possible.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and method for providing a call summary

[0001] This document relates to an electronic device, and, for example, to a method in which an electronic device provides a summary of a voice call.

[0002] Electronic devices (e.g., smartphones, tablet PCs, laptop PCs) can provide calling functions (e.g., voice calls, video calls). For example, the voice calling function is a basic function of electronic devices to transmit and receive voice data from the user of each electronic device via a cellular network (e.g., 4G network, 5G network) or a WLAN (wireless local area network).

[0003] An electronic device may provide a function to record the user's voice and the other party's voice during a call. Based on user input activating the recording function during a call, the electronic device may save the user's voice data acquired on the electronic device and the other party's voice data received through a network as a recording file. Additionally, after the call ends, the electronic device may convert the saved recording file into text information and save it.

[0004] Electronic devices can store the user's voice data and the other party's voice data in a single recording file when recording a call. In this case, it may be difficult to clearly distinguish in the recording file whether the recorded voice belongs to the user or the other party. Additionally, when the user and the other party speak simultaneously, there may be a problem with reduced recognition accuracy when converted to text.

[0005] An electronic device according to various embodiments of this disclosure (or specification, invention) may include a display, a microphone, a communication circuit, a memory, and at least one processor.

[0006] According to one embodiment, the memory may be executed by at least one processor, and may store instructions for the electronic device to perform a call with an external device, to acquire first voice data of a first user of the electronic device through the microphone during the call, and to acquire second voice data of a second user corresponding to the external device based on data received from the external device through the communication circuit during the call.

[0007] According to one embodiment, the memory may store instructions for the electronic device to generate call text data by mapping first speaker information corresponding to the first user to first text data converted from the first voice data, and mapping second speaker information corresponding to the second user of the external device to second text data converted from the second voice data.

[0008] A method performed by an electronic device according to various embodiments of the present document may include: performing a call with an external device; acquiring first voice data of a first user of the electronic device through a microphone of the electronic device during the call; acquiring second voice data of a second user corresponding to the external device based on data received from the external device during the call; mapping first speaker information corresponding to the first user to first text data converted from the first voice data, and mapping second speaker information corresponding to the second user of the external device to second text data converted from the second voice data to generate call text data; generating summary data of the call using the call text data; and providing the summary data in relation to the call.

[0009] According to various embodiments of the present document, an electronic device and a method for providing a summary of a call recording can be provided, which can more accurately convert voice data of a user and a counterparty obtained during a call with an external device into text and provide a summary of the call content using an AI model.

[0010] FIG. 1 is a block diagram of an electronic device in a network environment according to various embodiments.

[0011] FIG. 2 is a block diagram of an electronic device according to various embodiments.

[0012] FIG. 3 is a flowchart of a method for an electronic device according to one embodiment to store call recording data and call text data.

[0013] FIGS. 4a, FIGS. 4b, and FIGS. 4c illustrate a user interface screen that enables a call recording function of an electronic device according to one embodiment.

[0014] FIG. 5a is a flowchart of a method in which an electronic device according to one embodiment provides a summary of a call recording.

[0015] FIG. 5b is a block diagram of modules for audio signal processing of an electronic device according to one embodiment.

[0016] FIGS. 6A and 6B illustrate a user interface screen that provides the content of a call as text after recording a call of an electronic device according to one embodiment.

[0017] FIG. 7 is a flowchart of a method for an electronic device according to one embodiment to obtain summary data of a call using an AI model.

[0018] FIGS. 8A and 8B illustrate a user interface screen providing call content and a summary thereof of an electronic device according to one embodiment.

[0019] FIGS. 9a, FIGS. 9b, and FIGS. 9c illustrate a user interface screen providing call content and a summary thereof of an electronic device according to one embodiment.

[0020] FIG. 10 illustrates a user interface screen providing a list of call records of an electronic device according to one embodiment.

[0021] FIGS. 11a, FIGS. 11b, and FIGS. 11c illustrate a process in which an electronic device according to one embodiment verifies information of a call participant and reflects it in a summary of the call content.

[0022] FIGS. 12a and FIGS. 12b illustrate a user interface screen that provides the call content of each recording file when a plurality of recording files of an electronic device according to one embodiment are stored.

[0023] FIG. 13 illustrates a user interface screen in which an electronic device according to one embodiment provides summary data for a video call.

[0024] FIG. 14 illustrates a user interface screen in which an electronic device according to one embodiment provides a summary of a plurality of calls.

[0025] FIG. 15 illustrates a user interface screen in which an electronic device according to one embodiment provides a summary format corresponding to a conversation topic.

[0026] FIG. 16 is a flowchart of a method in which an electronic device according to one embodiment provides a call summary.

[0027] Hereinafter, embodiments of this document are described in detail with reference to the drawings so that those skilled in the art can easily implement them. However, this document may be implemented in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, identical or similar reference numerals may be used for identical or similar components. Additionally, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.

[0028] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to various embodiments.

[0029] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).

[0030] The processor (120) can control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., a program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.

[0031] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0032] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).

[0033] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0034] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0035] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0036] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0037] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).

[0038] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0039] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0040] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0041] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0042] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0043] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0044] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0045] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).

[0046] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) may support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.

[0047] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).

[0048] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0049] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.

[0050] According to one embodiment, commands or data may be transmitted or received between an electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In one embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0051] FIG. 2 is a block diagram of an electronic device according to various embodiments.

[0052] Referring to FIG. 2, an electronic device (200) according to various embodiments may include a display (230), a communication circuit (240), a microphone (250), a speaker (260), a processor (210), and a memory (220). Various embodiments of this document may be implemented even if some of the illustrated configurations are omitted or replaced with other configurations. In addition to the illustrated configurations, the electronic device (200) may further include at least some of the configurations and / or functions of the electronic device (101) of FIG. 1. At least some of each of the illustrated (or unillustrated) components of the electronic device (200) may be operatively, functionally, and / or electrically connected.

[0053] According to one embodiment, some of the components of the electronic device (200) (e.g., processor (210), memory (220), communication circuit (240)) may be placed inside the housing of the electronic device (200), and some other components (e.g., display (230), microphone (250), speaker (260)) may have at least some of their parts exposed outside the housing.

[0054] According to one embodiment, the electronic device (200) can be implemented as a device of various form factors in which the display area can be expanded, such as a foldable type, a rollable type (or a sliderable type).

[0055] According to one embodiment, the display (230) can display various images provided by the processor (210). For example, the display (230) may be implemented as any one of a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a micro electro mechanical systems (MEMS) display, or an electronic paper display, but is not limited thereto. The display (230) may be configured as a touch screen that detects touch and / or proximity touch (or hovering) input using a part of the user's body (e.g., a finger) or an input device (e.g., a stylus pen). The display (230) may include at least some of the configurations and / or functions of the display module (160) of FIG. 1.

[0056] According to one embodiment, the communication circuit (240) may include various configurations to support wireless communication with an external device. For example, the electronic device (200) may perform cellular wireless communication (e.g., 4G LTE (long term evolution), 5G NR (new radio)) and / or short-range wireless communication (e.g., Wi-Fi, Bluetooth) through the communication circuit (240), and there is no fixed type of wireless communication supported by the electronic device (200). The communication circuit (240) may include at least some of the configurations and / or functions of the communication module (190) of FIG. 1.

[0057] According to one embodiment, the microphone (250) can pick up external sound, convert it into a digital signal, and transmit it to the processor (210). According to one embodiment, the electronic device (200) can acquire a user's voice signal through the microphone (250) placed in the electronic device (200) or through the microphone of an external audio device (e.g., earbuds) connected via short-range wireless communication.

[0058] According to one embodiment, the speaker (260) can output an audio signal. For example, the speaker (260) can convert an electrical signal provided by the processor (210) into audio and output it. The speaker (260) can output audio through a hole formed in one area of ​​the housing. The speaker (260) may include at least some of the configuration and / or functions of the acoustic output module (155) of FIG. 1. The electronic device (200) can output an audio signal through the speaker (260) placed in the electronic device (200) or through the audio output of an external audio device (e.g., earbuds) connected via short-range wireless communication.

[0059] According to one embodiment, the memory (220) may include volatile memory and non-volatile memory and may store various data temporarily or permanently. The memory (220) may include at least some of the configuration and / or functions of the memory (130) of FIG. 1 and may store the program (140) of FIG. 1.

[0060] According to one embodiment, the memory (220) can store various instructions that can be executed by the processor (210). Such instructions may include control commands such as arithmetic and logical operations, data movement, and / or input / output that can be recognized by the processor (210).

[0061] According to one embodiment, the processor (210) may be configured to perform operations or data processing regarding the control and / or communication of each component of the electronic device (200), and may be composed of one or more processors. The processor (210) may include at least some of the configuration and / or functions of the processor (120) of FIG. 1.

[0062] According to one embodiment, the electronic device (200) may further include at least one other type of processor (e.g., a neural processing unit (NPU)) for processing computations of an on-device AI model in addition to the processor (210).

[0063] According to one embodiment, there are no limitations on the computational and data processing functions that the processor (210) can implement on the electronic device (200); however, this document describes various embodiments for converting voice call recording data into text, generating a summary of the call content using an AI model, and providing the generated summary through a user interface screen. The operations of the processor (210) described below can be performed by loading instructions stored in memory (220).

[0064] In this document, the description that a processor (210) can perform a certain operation (or function, task, or operation) may be interpreted substantially as meaning that an instruction (or command, computer program) causing the electronic device (200) (or processor (210)) to perform said operation is stored in memory (220) (e.g., non-volatile memory, storage). Additionally, the description that a processor (210) can perform a certain operation may be interpreted substantially as meaning that at least one processor, without a fixed number, can perform said operation.

[0065] According to one embodiment, the processor (210) can perform a call with an external device. In the following, the call may include a voice call or a video call. For example, the electronic device (200) may be connected to a cellular network (e.g., 4G network, 5G network) or a short-range wireless network (e.g., Wi-Fi) via a communication circuit (240) to perform a voice call (or video call) with an external device. When a voice call is initiated, the processor (210) may acquire a user's voice signal through a microphone (250) placed in the electronic device (200) or a microphone of an external audio device (e.g., earbuds) connected via short-range wireless communication, and receive a voice signal of the call partner via the network. The processor (210) receives a data packet (e.g., a real-time transport protocol (RTP) packet) from a network through a communication circuit and can obtain second voice data of the call partner through decoding, voice signal processing, and / or a digital to analog converter (DAC) process for the data packet.

[0066] According to one embodiment, audio data (e.g., first voice data) for a voice signal acquired by the electronic device (200) through a microphone may be stored at least temporarily in the transmitting audio buffer of the electronic device (200). Audio data (e.g., second voice data) for a voice signal acquired by the electronic device (200) through a network may be stored at least temporarily in the receiving audio buffer of the electronic device (200). For example, the transmitting audio buffer may correspond to a local audio channel, and the receiving audio buffer may correspond to a remote audio channel. For example, the local audio channel and the remote audio channel may be included in an audio framework (e.g., RF call audio framework). For example, the local audio channel and the remote audio channel may be managed by the audio framework. The audio framework and audio buffers will be described in more detail through FIG. 5b.

[0067] According to one embodiment, voice data from a local audio channel or a remote audio channel may be converted into text and stored as call text data. For example, voice data from a local audio channel or a remote audio channel may be stored as call recording data.

[0068] For example, the processor (210) may classify text converted from voice data of a local audio channel as text corresponding to a user of the electronic device (200) (e.g., first text). In relation to the first text, first speaker information corresponding to the user may be recorded (e.g., tagging, assignment, or mapping). For example, the processor (210) may classify text converted from voice data of a remote audio channel as text corresponding to a call partner (e.g., second text). In relation to the second text, second speaker information corresponding to the partner may be recorded (e.g., tagging, assignment, or mapping).

[0069] The electronic device (200) can output a sound corresponding to audio data (e.g., at least one of first voice data or second voice data) through the speaker (260) of the electronic device (200) or the audio output of an external audio device.

[0070] According to one embodiment, the transmitted / received voice signal used for a voice call may be generated through a voice synthesis function. For example, the processor (210) may generate a transmitted voice using text (e.g., a sentence entered by a user, a sentence selected by a user) and transmit it to an external device connected for a voice call. For example, the processor (210) may generate a voice of another language using a voice of a specific language entered by a user and transmit it to an external device connected for a voice call.

[0071] According to one embodiment, when a voice call (or video call) is initiated, the processor (210) may display a user interface through a display (230) that provides various settings related to the voice call (or video call). For example, the processor (210) may provide a user interface through the display (230) that includes at least one item (e.g., a button) for selecting call recording, messaging, Bluetooth connection, speaker activation, sound muting, and / or keypad activation. The user interface providing various settings related to the voice call will be described in more detail with reference to FIG. 4a.

[0072] According to one embodiment, the processor (210) can determine whether a call recording function is enabled or disabled in relation to a voice call. For example, the processor (210) can enable the call recording function during a voice call. For example, the processor (210) can enable or disable the call recording function in response to user input regarding an item that allows call recording to be selected. For example, the processor (210) can automatically perform call recording when a call is initiated in response to a setting that enables the call recording function when a call is initiated. For example, after automatically enabling the call recording function, the processor (210) can disable or re-enable the call recording function in response to user input regarding an item that allows call recording to be selected. According to one embodiment, the call recording function may be set to be automatically enabled when a voice call is initiated under specified conditions (e.g., a voice call with a specified counterparty, a voice call at a specified time, a voice call at a specified location). The processor (210) can detect the occurrence of a call recording based on user input or automatic call recording settings. According to one embodiment, the processor (210) can generate at least one recording data during a call as the call recording is performed by activating the call recording function. For example, if the user turns the call recording item on / off multiple times during a voice call, multiple recording files may be generated and stored for the corresponding voice call.

[0073] According to one embodiment, the processor (210) can generate recording data by recording audio data of a voice call in response to the activation of the recording function. According to one embodiment, the processor (210) can collect voice data of a user obtained through the microphone (250) of the electronic device (200) or the microphone (250) of an external audio device and voice data of the call partner received from a network through the communication circuit (240). For example, the processor (210) can execute a process that performs a recording function and, through said process, record first voice data corresponding to the user's voice of the electronic device, which is temporarily stored in the transmission audio buffer of the audio framework, and second voice data corresponding to the user's voice of the external audio device, which is temporarily stored in the reception audio buffer. The processor (210) can generate recording data in the form of a digital audio file (e.g., m4a) by processing the collected voice data in real time.

[0074] According to one embodiment, the processor (210) can convert audio information of the first voice data and the second voice data of a call into text information. For example, the processor (210) can convert voice data into text data using a speech recognition algorithm (or a speech-to-text (STT) algorithm), or generate text data using an AI model trained to extract text information from voice data. Text data converted from the recording data of a voice call may be referred to as a transcript.

[0075] Table 1 shows a transcript of the conversation 1 between the user of the electronic device (200) and the call partner of an external device, converted into text.

[0076] I have something to do today. What is it all of a sudden? I’m going to study English by watching videos. Are you studying in a study group? No, I’m studying alone. I’m going to a restaurant for lunch today after I finish studying; do you want to come? Are you buying me lunch for free if I study hard? If you study hard, I’ll take you to the department store later and buy you the coat you wanted. Really? I mean it. I really want you to study hard. Okay. Then I’ll study hard! Great. If you get good grades, I’ll buy you something else too. Really? Thank you so much! See you later.

[0077] Table 2 shows a transcript of the conversation 2 between the user of the electronic device (200) and the call partner of the external device converted into text.

[0078] Do you like baseball? Everyone my age likes it. Kim was seriously amazing at the game yesterday. He hit a walk-off home run in his last at-bat to end the game, right? He looked so cool when he hit that home run. I became a fan after watching that. He's a player who performs really well in crucial moments. I saw his interview too, and it was really cool. Did you see the interview too? He's the kind of player that all girls our age would like. I didn't see the interview, but do you happen to have a link to the video? Just wait a moment. Here, I'll share the link with you. Wow, Kim has a really deep voice, unlike how he looks. I guess it's because he's an athlete. Shall we go to the baseball practice range sometime? Okay, let's go together! Today, I'm going to hit a lot of home runs just like Kim.

[0079] According to one embodiment, the processor (210) can convert voice data (or recording data) (e.g., first voice data, second voice data) into text data in real time while the recording of a voice call is in progress. According to one embodiment, the processor (210) can receive voice data (or recording data) in real time and perform text conversion through a process separate from the process performing the recording. For example, the processor (210) can generate recording data by transmitting first voice data containing the voice of an electronic device user stored in a transmission audio buffer (or first audio buffer) and second voice data containing the voice of an external device user stored in a reception audio buffer (or second audio buffer) to a first process that performs a recording function, and at least partially simultaneously with the operation of the first process generating recording data, transmit the first voice data containing the voice of an electronic device user stored in the transmission audio buffer and the second voice data containing the voice of an external device user stored in the reception audio buffer to a second process that performs text conversion, thereby converting the voice data into text data. The second process may be one of the processes of a call application. According to one embodiment, the processor (210) may perform text conversion by dividing voice data (or recording data) into units of a predetermined time (or frame). If text conversion is performed after the recording is finally completed, the conversion time is required (e.g., about 24 seconds when converting a 3-minute call recording), so the call text data may not be provided to the user in real time. According to various embodiments of this document, by performing text conversion in real time during the call recording, the call text data (or transcript) can be provided to the user more quickly after the call ends.

[0080] According to one embodiment, the processor (210) can store voice call recording data and text data. For example, when call recording is disabled (e.g., when call recording is terminated) or when the call is terminated, the processor (210) can create the recording data as a single file and store it in memory (220) (or cloud storage). Additionally, when the creation of text data (or transcript) for the recording data is completed, the processor (210) can create the text data as a single file and store it in memory (220) (or cloud storage). The processor (210) can store the recording data and text data by mapping them to each other.

[0081] According to one embodiment, the processor (210) can generate summary data of a voice call using an artificial intelligence model based on call text data. According to one embodiment, the summary data may include a title, a keyword, and a summary of the call content.

[0082] According to one embodiment, the AI ​​model may be a generative AI model or a large language model (LM) trained to generate titles, keywords, images, and / or summaries for call content based on call text data. For example, the AI ​​model may be designed based on a transformer architecture (e.g., BERT, GPT) to pre-train on a large dataset and fine-tuned to extract summaries, titles, and / or keywords for call content.

[0083] According to one embodiment, an AI model can analyze the content of text data by applying natural language processing technology to analyze sentence structure, semantic analysis, and / or contextual analysis of the text data. For example, the AI ​​model can construct a summary by extracting important sentences from call text data, and / or construct a summary of short sentences by analyzing the context and semantics of the conversation content. Additionally, the AI ​​model can generate a title by composing a brief sentence based on the summary of the call text data. Furthermore, the AI ​​model can extract at least one keyword from the summary of the call text data based on statistics and / or semantics. Additionally, the AI ​​model can generate an image (e.g., a still image or a video) that depicts the topic or context of the call by analyzing the content of the call text data.

[0084] According to one embodiment, an electronic device (200) can generate summary data using an on-device AI model. The on-device AI model may be an AI model designed to perform AI computations using hardware and software within the electronic device (200) without relying on an external server. A processor (210) of the electronic device (200) may handle the training of the AI ​​model, data analysis, and / or result generation, or the electronic device (200) may include a neural processing unit (NPU) for performing the operation of the AI ​​model. As the electronic device (200) uses an on-device AI model, for example, there are advantages in security and privacy, and latency is reduced, enabling real-time processing. As the electronic device (200) uses an on-device AI model, for example, the electronic device (200) can generate call summaries without using network data.

[0085] According to one embodiment, the processor (210) may generate a prompt including call text data and a request to generate a summary of the call content to request a summary of a voice call from the AI ​​model, and transmit it to the AI ​​model.

[0086] According to one embodiment, when generating summary data using an AI model, the processor (210) can record (or tag) speaker information of each text in the text information converted from voice data and transmit it to the AI ​​model.

[0087] According to one embodiment, the processor (210) can generate call text data by mapping first speaker information corresponding to a user of an electronic device to first text data converted from first voice data, and mapping second speaker information corresponding to a user of an external device to second text data converted from second voice data. For example, the processor (210) can record text indicating a user of an electronic device (e.g., me, you) in the first text data and record text indicating a user of an external device (e.g., the other person, the other) in the second text data.

[0088] According to one embodiment, the processor (210) can divide call text data into a plurality of text blocks. For example, when a user and a call partner take turns speaking voices during a voice call, each speech can be composed of a single text block. The processor (210) can tag speaker information for each text block.

[0089] According to one embodiment, the processor (210) can distinguish between the user's first voice data obtained through the microphone (250) and the other party's second voice data received from an external device through the communication circuit (240).

[0090] According to one embodiment, the processor (210) can clearly distinguish between the user's first voice data and the counterpart's second voice data, which are temporarily stored in an audio framework (or RF call audio framework). Here, the audio framework may include an API that provides a function to distinguish audio data transmitted and received in a voice call. In other words, since the first voice data corresponding to the user and the second voice data corresponding to the counterpart are distinguished on the audio framework, the processor (210) can determine that the text converted from the voice data of the audio framework is text corresponding to a specific speaker.

[0091] According to one embodiment, the processor (210) can generate call text data by adding information of the speaker who uttered the voice to text information (e.g., transcript) converted from recording data.

[0092] Table 3 shows call text data recording the speaker in the above conversation 1.

[0093] Me: I have something to do today. Other: What's up all of a sudden? Me: I'm going to study English by watching videos. Other: Are you studying in a study group? Me: No, I'm studying alone. I'm going to a restaurant for lunch today after I finish studying; do you want to come along? Other: If I study hard, you'll buy me lunch for free? Me: If you study hard, I'll take you to the department store later and buy you that coat you wanted. Other: Really? Me: I mean it. I really want you to study hard. Other: Okay. Then I'll study hard! Me: Great. If you get good grades, I'll buy you something else too. Other: Really? Thank you so much! See you later.

[0094] Table 4 shows call text data recording the speaker in the above conversation 2.

[0095] Other: Do you like baseball? Me: Everyone my age likes it. Other: Kim was absolutely amazing in yesterday's game. He hit a walk-off home run in his last at-bat to end the game. He looked so cool when he hit that home run. I became a fan after watching that. Me: He's a player who really performs well in crucial moments. I saw his interview too, and it was really cool. Did you see the interview, too? Other: He's the kind of player that all girls our age would like. I didn't see the interview, but do you happen to have the video link? Me: Just wait a moment. Here, I'll share the link with you. Other: Wow, Kim, your voice is really deep, unlike how you look. I guess it's because you're an athlete. Do you want to go to the baseball practice range sometime? Me: Sure, let's go together! Today, I'm going to hit a lot of home runs just like Kim.

[0096] According to one embodiment, the processor (210) may request the generation of summary data by transmitting call text data, in which speaker information is tagged for each text block as shown in Tables 3 and 4, to an AI model. According to one embodiment, if the current voice call is a multi-party call, the processor (210) may distinguish voice data of multiple counterparts and record it in the call text data. For example, in a multi-party call, the voice data of the first counterpart and the voice data of the second counterpart may be mixed over a network and transmitted to an electronic device (200). According to one embodiment, the processor (210) may distinguish the voice data of the first counterpart and the voice data of the second counterpart based on voice analysis of the received voice data. For example, the processor (210) may distinguish the voice data into the voice data of the first counterpart and the voice data of the second counterpart by analyzing acoustic features (e.g., timbre, pitch, frequency spectrum) that can distinguish the speaker in the voice data received from the network. The electronic device (200) can distinguish between text information converted from the voice data of the first counterpart and text information converted from the voice data of the second counterpart when converting recorded data into text, and tag speaker information in each text information (or text block). According to one embodiment, the processor (210) can check whether the amount of call text data to be summarized can be processed by the AI ​​model. The on-device AI model may be lightweight compared to the server-based AI model in order to operate efficiently on the electronic device (200). The AI ​​model may process input data in token units, and the maximum number of tokens that can be processed for each model may be determined by language, and the size of the tokens that can be processed for the on-device AI model may be relatively smaller compared to the server-based AI model.

[0097] According to one embodiment, the processor (210) can divide call text data into multiple chunks according to the maximum throughput of the AI ​​model. For example, the processor (210) can calculate the sum of the number of tokens of sequential text blocks, form text blocks within a range that does not exceed the maximum number of tokens that the AI ​​model can process into one chunk, and form multiple text blocks from the next text block into another chunk.

[0098] According to another embodiment, the processor (210) can divide text blocks into multiple chunks based on semantic analysis of call text data.

[0099] Table 5 shows an example of dividing the call text data of Conversation 2 into multiple chunks.

[0100] Chunk 1 Other: Do you like baseball? Me: Everyone my age likes it. Other: Kim was absolutely amazing in yesterday's game. He hit a walk-off home run in his last at-bat to end the game. He looked so cool when he hit that home run. I became a fan after watching that. Me: He's a player who performs really well in crucial moments. I saw his interview too, and it was really cool. Did you see the interview, too? Chunk 2 Other: He's the kind of player that all girls our age would like. I didn't see the interview, but do you happen to have the video link? Me: Just wait a moment. Here, I'll share the link with you. Other: Wow, Kim, your voice is really deep, unlike how you look. I guess it's because you're an athlete. Do you want to go to the baseball practice range sometime? Me: Sure, let's go together! Today, I'm going to hit a lot of home runs just like Kim.

[0101] According to one embodiment, when the processor (210) divides the call text data into multiple chunks, it may generate a prompt containing each chunk and transmit it to an AI model, and receive a response from the AI ​​model containing summary data for each chunk. According to one embodiment, the processor (210) may obtain summary data of the call content as a response from the AI ​​model to a prompt for requesting the generation of a summary of the call content. Table 6 shows the summary data generated by the AI ​​model for the call text data of the conversation 1.

[0102] Title: Promise to Buy Lunch and a Coat After Studying Summary: Promising lunch, a coat, and additional rewards to encourage and motivate the other person to study English Keywords: Study, Lunch, Coat, Gift

[0103] Table 7 shows summary data generated by the AI ​​model for the call text data of the above conversation 1.

[0104] Title: Sharing Player Kim's Great Game and Passion for Baseball with the Other Person Summary: Established common ground with the other person while discussing Player Kim's walk-off home run. The other person mentioned that women my age would likely be attracted to an attractive player like Kim and requested an interview video. Afterwards, the other person and I decided to go to the baseball practice range, expressing our determination to hit home runs just like Kim. Keywords: Baseball, Player Kim, Walk-off Home Run, Baseball Practice Range

[0105] According to one embodiment, the processor (210) may include user information of a call participant in a prompt for requesting the generation of a summary of the call content. For example, the electronic device (200) may include user information obtained from an application, or user information obtained through voice analysis, in the prompt and transmit it to an AI model. According to one embodiment, the processor (210) may identify user information of the call counterpart (e.g., name, workplace information, relationship) stored in a contact application and include user information in a prompt for requesting the generation of a summary of the call content. According to one embodiment, the processor (210) may obtain user information of the user or counterpart based on the analysis of the voice data of the user or counterpart of the electronic device (200) from the voice data of the call. For example, the electronic device (200) may obtain user information such as gender, age group, region, emotion, and / or surrounding environment based on the analysis of characteristics of the voice data such as frequency band, intensity, tone, speed, and / or pronunciation. The processor (210) may include user information identified through voice analysis in a prompt for a request to generate a summary of the call content.

[0106] According to one embodiment, an AI model can generate summary data of the call content based on call text data and user information included in a prompt.

[0107] Table 8 shows summary data generated by the AI ​​model considering user information (e.g., Lee, male in his 20s) in Conversation 2.

[0108] Title: Sharing Player Kim's Great Game and Passion for Baseball with Lee Summary: Formed a bond with Lee, a man in his 20s, while discussing Player Kim's walk-off home run. Lee mentioned that women in their 20s would also likely like an attractive player like Kim and requested an interview video. Afterwards, Lee and I decided to go to the baseball practice range, expressing our determination to hit a home run just like Kim. Keywords: Baseball, Player Kim, Female Fan in 20s, Walk-off Home Run, Baseball Practice Range

[0109] An embodiment for generating and displaying summary data including user information will be described in detail through FIG. 11a, FIG. 11b, and FIG. 11c. According to one embodiment, a processor (210) may provide call text data and summary data through a display (230). For example, the processor (210) may distinguish call text data into a plurality of text blocks and display them in a first area of ​​the display (230), and display summary data in a second area of ​​the display (230). Additionally, the processor (210) may display a user interface for controlling the playback of recording data in a third area of ​​the display (230). Here, the first area may include the center area of ​​the display (230), the second area may be the upper area of ​​the first area, and the third area may be the lower area of ​​the first area, but is not limited thereto. According to one embodiment, when the processor (210) displays call text data on the display (230), it may provide a summary item that triggers the creation of summary data. The processor (210) can generate and transmit to the AI ​​model a prompt including a request to generate a summary of call text data and call content, based on user input for a summary item, to request the AI ​​model to generate a summary of a voice call, and can obtain summary data from the AI ​​model.

[0110] According to one embodiment, the processor (210) may display a summary data item containing at least a portion of the summary data through the display (230). For example, the processor (210) may display the summary data item including the title and keywords of the call content, and may display a summary based on user input for a more item. This embodiment will be described in more detail through FIGS. 8a, 8b, 9a, 9b, and 9c.

[0111] According to one embodiment, the processor (210) can output an audio signal of a text block through a speaker (260) (or an external audio device) based on user input for the text block. The processor (210) can provide a user interface including an item for selecting play / pause, backward, and forward when playing the audio of the text block, and an item indicating a playback section in the entire recording data.

[0112] According to one embodiment, the processor (210) may provide at least a portion of summary data corresponding to a call record on a call list screen that includes at least one call record. This embodiment will be described in more detail with reference to FIG. 10.

[0113] According to one embodiment, the processor (210) may provide a user interface that can switch from displaying call text data and summary data corresponding to one of the recording data to displaying call text data and summary data corresponding to another recording data when multiple recording data are generated during a single voice call. This embodiment will be described in more detail through FIGS. 12a and 12b.

[0114] According to one embodiment, the electronic device (200) may provide a video call function. According to one embodiment, the processor (210) may generate summary data based on call text data and video data of the video call (e.g., first video data, second video data). The summary data generated during a video call may include a title, keywords, a summary of the call content, and a summary video. This embodiment will be described in more detail with reference to FIG. 13.

[0115] According to one embodiment, an electronic device (200) may perform multiple voice calls with the same external device (or call partner) and record multiple voice calls. In this case, the processor (210) may analyze the call text data of each voice call to determine whether two calls are related based on the identity or association between conversation topics, the time interval between each call, etc., and may generate summary data based on the voice data of the voice calls of the related conversations. This embodiment will be described in more detail with reference to FIG. 14.

[0116] According to one embodiment, the processor (210) can generate summary data in a format corresponding to the content of the call text data (or voice data) of a voice call. This embodiment will be described in more detail with reference to FIG. 15.

[0117] Instructions for performing the operation of the electronic device (200) (or processor (210)) described above may be stored in a computer-readable recording medium. The recording medium may be tangible and non-transitory. The recording medium may store one or more computer programs containing the instructions.

[0118] FIG. 3 is a flowchart of a method for an electronic device according to one embodiment to store call recording data and call text data.

[0119] According to one embodiment, the illustrated method may be performed by an electronic device (e.g., the electronic device (200) of FIG. 2), and the technical features described above may be omitted from the description below.

[0120] According to one embodiment, in operation 310, the electronic device may initiate a voice call. For example, the electronic device may be connected to a cellular network (e.g., 4G network, 5G network) or a short-range wireless network (e.g., Wi-Fi) via a communication circuit (e.g., communication circuit (240) of FIG. 2) to initiate a voice call (or video call) with an external device. When a voice call is initiated, the electronic device may acquire a voice signal of the user through a microphone placed on the electronic device (e.g., microphone (250) of FIG. 2) or a microphone of an external audio device (e.g., earbuds) connected via short-range wireless communication, and receive a voice signal of the call partner via the network. The electronic device may output the voice of the user and the voice of the call partner through the speaker of the electronic device or the speaker of the external audio device.

[0121] According to one embodiment, in operation 320, the electronic device can check whether user input for call recording is received. According to one embodiment, when a voice call is initiated, the electronic device can display a user interface providing various settings related to the voice call through a display (e.g., the display (230) of FIG. 2). For example, the electronic device may provide a user interface during a voice call that includes at least one item for selecting call recording, messaging, Bluetooth connection, speaker activation, sound muting, and / or keypad activation. According to one embodiment, the electronic device may initiate call recording when the call recording item of the user interface is input.

[0122] According to one embodiment, when a user input for recording a call is received, in operation 325, the electronic device may record audio data of a voice call. According to one embodiment, the electronic device may store recording data in a memory (e.g., memory (220) of FIG. 2) (or cloud storage), including first voice data of a user obtained through a microphone obtained in real time during a call and second voice data of a call partner of an external device obtained from a data packet received from a network through a communication circuit.

[0123] According to one embodiment, in operation 330, the electronic device can convert audio data of a voice call into text information (or transcript) during call recording. For example, the electronic device can convert the first voice data and the second voice data into text data using a speech recognition algorithm (or a speech-to-text (STT) algorithm), or generate text data using an AI model trained to extract text information from voice data.

[0124] According to one embodiment, an electronic device can convert voice data (or recording data) (e.g., first voice data, second voice data) into text data in real time while the recording of a voice call is in progress. For example, the electronic device may receive the voice data (or recording data) in real time and perform text conversion through a process separate from the process of performing the recording. The electronic device may perform text conversion by dividing the voice data into units of a predetermined time (or frame). If text conversion is performed after the recording is finally completed, the conversion time is required (e.g., about 24 seconds when converting a 3-minute call recording), so the call text data may not be provided to the user in real time. According to various embodiments of this document, by performing text conversion in real time during the call recording, the call text data (or transcript) can be provided to the user more quickly after the call ends.

[0125] According to one embodiment, an electronic device can distinguish between first text data generated by converting first voice data of a user acquired through a microphone and second text data generated by converting second voice data of a call partner acquired from a data packet received from an external device through a communication circuit. For example, first voice data for a voice signal acquired by the electronic device through a microphone may be stored at least temporarily in the transmitting audio buffer of the electronic device, and second voice data for a voice signal acquired by the electronic device through a network may be stored at least temporarily in the receiving audio buffer of the electronic device. The first voice data stored in the transmitting audio buffer and the second voice data stored in the receiving audio buffer may be transmitted in real time to a call application or a process that performs text translation, and accordingly, the first voice data and the second voice data may be recognized as being separated into separate channels.

[0126] According to one embodiment, the electronic device can map information that can identify the speaker user (e.g., me, you) to the separated first text data and map information that can identify the speaker call counterpart (e.g., counterpart, the other) to the second text data and record it.

[0127] According to one embodiment, when the current voice call is a multi-party call, the electronic device can distinguish voice data of multiple counterparts and record it in the call text data. For example, in a multi-party call, the voice data of the first counterpart and the voice data of the second counterpart may be mixed over a network and transmitted to the electronic device. According to one embodiment, the electronic device can distinguish the voice data of the first counterpart and the voice data of the second counterpart based on voice analysis of the received voice data. For example, the electronic device can distinguish the voice data into the voice data of the first counterpart and the voice data of the second counterpart by analyzing acoustic features (e.g., timbre, pitch, frequency spectrum) that can distinguish the speaker in the voice data received from the network. When converting the recorded data into text, the electronic device can distinguish the text information converted from the voice data of the first counterpart and the text information converted from the voice data of the second counterpart, and record speaker information in each text information.

[0128] According to one embodiment, an electronic device can divide the texts of call text data into a plurality of text blocks. For example, when a user and a call partner take turns speaking voices in sequence during a voice call, each speech can be composed of a single text block.

[0129] According to one embodiment, in operation 335, the electronic device can determine whether the voice call is terminated. For example, the electronic device may determine that the voice call is terminated if user input for a call termination item is received from a user interface provided during the call, or if a call termination signal is received from a network.

[0130] According to one embodiment, when a voice call ends, in operation 340, the electronic device may store the recording data of the voice call. For example, the electronic device may store the recording data as a single audio file in memory (or cloud storage). According to one embodiment, the electronic device may store the call recording data including information of the call participants and time information.

[0131] According to one embodiment, when a voice call is terminated, in operation 350, the electronic device may store call text data mapped with speaker information of the voice call. The call text data may include first text data converted from first voice data of the electronic device user and second text data converted from second voice data of the call counterpart. Additionally, the electronic device may map (or record) first speaker information corresponding to the electronic device user (e.g., me, you) to the first text data and map (or record) second speaker information corresponding to the call counterpart (e.g., counterpart, the other) to the second text data. The electronic device may store the call text data as a single file in memory (or cloud storage).

[0132] According to one embodiment, operation 340 and operation 350 may be performed by different processes, and at least some of each operation may be performed at least some simultaneously.

[0133] According to one embodiment, instructions for performing each operation constituting the method may be stored on a computer-readable recording medium. The recording medium may be tangible and non-transitory. The recording medium may store one or more computer programs containing said instructions.

[0134] FIGS. 4a, FIGS. 4b, and FIGS. 4c illustrate a user interface screen that enables a call recording function of an electronic device according to one embodiment.

[0135] FIG. 4a illustrates a user interface screen provided on the display (230) of the electronic device (200) when a call is initiated.

[0136] According to one embodiment, an electronic device (200) (e.g., the electronic device (200) of FIG. 2) may display a user interface providing various settings related to the voice call through a display (230) (e.g., the display (230) of FIG. 2) when a voice call (or video call) is initiated. The electronic device (200) may initiate call recording in response to user input to a call recording item (415) of the user interface.

[0137] Referring to FIG. 4a, the electronic device (200) may display on a display (230) a user interface that includes items for selecting voice call-related functions when connected to a voice call via cellular wireless communication (e.g., 4G LTE, 5G NR) or short-range wireless communication (e.g., Wi-Fi) with an external device corresponding to the call partner. For example, the user interface may include an item panel (410) containing a plurality of items, a call assist item (430), and a call recording item (415). The item panel (410) may include an add call item (411), a message item, a Bluetooth connection item, a speaker item, a sound mute item, and a keypad item.

[0138] According to one embodiment, the electronic device may display an item (450) indicating the name of a call partner stored in a contact application (e.g., a call application, a contact application), and in the case of a multi-party call, the names of multiple call partners may be displayed.

[0139] According to one embodiment, the electronic device (200) may provide a call assist item (430) capable of executing additional functions provided during a call. For example, the electronic device (200) may execute a real-time translation function of the call content and / or a text call capable of converting text input into a voice signal and transmitting it based on a touch input to the call assist item (430).

[0140] According to one embodiment, the electronic device (200) can activate (e.g., start) or deactivate (e.g., stop) the recording function based on user input regarding the call recording item (415) while the call is connected. In FIG. 4a, the call recording item (415) is placed outside the item panel (410), but it may also be provided within the item panel (410).

[0141] According to one embodiment, when a call recording item (415) is selected, the electronic device (200) may start a timer (e.g., 3 seconds) for starting call recording and perform call recording when the timer expires. In this case, the electronic device (300) may provide a graphic object representing the timer. According to one embodiment, when call recording is started, the electronic device (200) may display a guide message indicating that call recording is in progress, and a guide message indicating that call recording is in progress may also be provided to an external device that is the call counterpart.

[0142] FIG. 4b illustrates a user interface screen provided on the display (230) of the electronic device (200) when call recording is initiated.

[0143] According to one embodiment, the electronic device (200) may enable or disable the recording function based on user input regarding the call recording item (415) while the call is connected. The call recording item may be displayed within the item panel (410) as in FIG. 4b, or may be displayed in an area separate from the item panel (410) as in FIG. 4a.

[0144] According to one embodiment, the call recording item (412) may provide information indicating that recording is stopped upon touch input when recording is enabled, as shown in FIG. 4b, and information indicating that recording is started upon touch input when recording is disabled.

[0145] According to one embodiment, when multiple segments are recorded in a single call session based on user input for a call recording item (412), the electronic device (200) may store the voice data of each segment as separate files. In this case, the electronic device (200) may store the call text data corresponding to each file as separate files. Alternatively, when multiple segments are recorded in a single call session, the electronic device (200) may merge the voice data of each segment and store them as a single file.

[0146] According to one embodiment, while a call recording is executed and a screen such as FIG. 4b is displayed on a display (230), the electronic device (200) can convert voice data (or recording data) (e.g., first voice data of the electronic device user, second voice data of the external device user) into text data in real time through a background process. The electronic device (200) can distinguish between first text data generated by converting the user's first voice data obtained through a microphone and second text data generated by converting the call counterpart's second voice data obtained from data received from an external device through a communication circuit. The electronic device (200) can map information that can identify the speaker user (e.g., me, you) to the distinguished first text data and map information that can identify the call counterpart (e.g., the other) to the second text data.

[0147] According to one embodiment, the above-described operations may be performed in accordance with user input to a call summary item (e.g., a button that can select the start or end of the call summary function, not shown). In other words, the electronic device (200) may, in response to the activation of the call summary function, generate first text data by converting user voice data obtained through a microphone through a process separate from the process of performing recording, and generate second text data by converting voice data of the call partner received from an external device through a communication circuit. The electronic device (200) may store call text data including the first text data generated by converting the user's voice data and the second text generated by converting the partner's voice data. In the call text data, speaker information capable of identifying the user may be recorded (or mapped) in the first text, and speaker information capable of identifying the partner may be recorded (or mapped) in the second text. The processor (210) may generate a call summary through an AI model based on the call text data containing speaker information and text.

[0148] As the electronic device (200) of the present disclosure can perform recording and summarization as separate processes, the processor (210) can convert voice conversation between the user and the other party into text in near real-time, for example, even while the call recording function is disabled. As the electronic device (200) of the present disclosure can perform recording and summarization as separate processes, the processor (210) can generate call text data from the user's voice data and the other party's voice data, for example, at least partially simultaneously, and generate a recording file from the user's voice data and the other party's voice data.

[0149] According to one embodiment, the call summary function may be activated before a call is performed. The call summary function may be set to automatically start when a call is established with a designated counterparty. The call summary function may automatically terminate in response to the termination of the call. According to one embodiment, while the call is being performed, call text data is generated, and after the call is terminated, a file corresponding to the call text data is saved, and a call summary based on the call text data may be automatically performed.

[0150] FIG. 4c illustrates a user interface screen provided on the display (230) of the electronic device (200) after the call ends.

[0151] According to one embodiment, the electronic device (200) removes the display of the user interface (410) providing call-related settings shown in FIG. 4a or FIG. 4b when the call is completed, and provides a user interface (460) that provides functions executable after the call ends, as in FIG. 4c. Referring to FIG. 4c, the user interface (460) that provides functions executable after the call ends may include at least one item that allows selecting to block the other party, add to contacts, edit tags, reconnect the call, and / or send a message.

[0152] According to one embodiment, the electronic device (200) can store voice call recording data and text data when a voice call ends. For example, the electronic device (200) can store the recording data and text data as separate files in memory (or cloud storage).

[0153] FIG. 5a is a flowchart of a method in which an electronic device according to one embodiment provides a summary of a call.

[0154] According to one embodiment, the illustrated method may be performed by an electronic device (e.g., the electronic device (200) of FIG. 2), and the technical features described above may be omitted from the description below.

[0155] According to one embodiment, in operation 510, the electronic device may store voice call recording data and text data. For example, the electronic device may start a call recording based on user input to a call recording item (e.g., a call recording item (415) in FIG. 4a) of a user interface during a voice call (or video call) with an external device, terminate the call recording upon re-inputting the call recording item or the end of the call, and store the recording data in memory. The electronic device may acquire voice data stored in an audio buffer during call recording and perform text conversion in real time to generate and store call text data.

[0156] According to one embodiment, the electronic device can distinguish and record the speaker (e.g., user of the electronic device, call partner of an external device) for each text included in the call text data.

[0157] According to one embodiment, in operation 515, the electronic device can determine whether a text view of a voice call is selected. For example, if a call recording is completed, the electronic device may provide a notification item related to the call recording on a notification panel. The notification item related to the call recording may include a text view item that triggers the display of call text data.

[0158] According to one embodiment, when a text view of a voice call is selected based on user input, in operation 520, the electronic device can display call text data.

[0159] According to one embodiment, an electronic device can display call text data by dividing it into a plurality of text blocks. For example, when a user and a call partner take turns speaking voices in sequence during a voice call, the electronic device can compose each utterance into a single text block.

[0160] According to one embodiment, in operation 525, the electronic device can check whether a summary of the voice call is selected.

[0161] According to one embodiment, the electronic device may provide a summary item that triggers a summary of call text data. For example, the electronic device may display call text data on a first area including a center area of ​​the display and display a summary item on a second area corresponding to the top of the first area.

[0162] According to one embodiment, when a summary of a voice call is selected based on user input, in operation 530, the electronic device can generate summary data of the voice call using an artificial intelligence model.

[0163] According to one embodiment, an electronic device can generate a prompt including call text data and a request to generate a summary of the call content to request a summary of a voice call from an AI model, transmit the prompt to the AI ​​model, and obtain summary data of the call content as a response from the AI ​​model.

[0164] According to one embodiment, an electronic device may record speaker information for each text block of call text data included in a prompt. For example, the electronic device may record the user as the speaker in the first text data converted from the utterance of the user of the electronic device, and record the counterparty as the speaker in the second text data converted from the utterance of the counterparty. In this way, when speaker information is recorded, the quality of the summary data generated by the AI ​​model may be improved compared to when a transcript, which converts the call content into text without distinguishing the speaker, is transmitted to the AI ​​model.

[0165] According to one embodiment, the AI ​​model may be a generative AI model or a large language model (LM) trained to generate a title, keywords, and / or summary of the call content based on call text data.

[0166] According to one embodiment, an electronic device may generate summary data using an on-device AI model. The on-device AI model may be an AI model designed to perform AI computations using hardware and software within the electronic device without relying on an external server. A processor of the electronic device (e.g., processor (210) of FIG. 2) may process the training of the AI ​​model, data analysis, and / or result generation, or the electronic device may include a neural processing unit (NPU) for performing the operation of the AI ​​model.

[0167] According to one embodiment, the electronic device may generate summary data for a call using call text data in response to the termination of a call or the termination of a call recording, without user input selecting a summary of the voice call. In this case, operation 525 may be omitted.

[0168] According to one embodiment, in operation 535, the electronic device may display summary data along with text data on a display. According to one embodiment, the electronic device may distinguish call text data into a plurality of text blocks and display them in the order in which they were entered. For example, the electronic device may display text blocks in a form such as messages transmitted and received in a message application.

[0169] According to one embodiment, the electronic device can output an audio signal of a text block through a speaker (or external audio device) based on user input to the text block.

[0170] According to one embodiment, instructions for performing each operation constituting the method may be stored on a computer-readable recording medium. The recording medium may be tangible and non-transitory. The recording medium may store one or more computer programs containing said instructions.

[0171] FIG. 5b is a block diagram of modules for audio signal processing of an electronic device according to one embodiment.

[0172] According to one embodiment, an electronic device (e.g., the electronic device (200) of FIG. 2) may include various hardware and / or software modules for processing audio data transmitted and received during a call. For example, an audio framework (550) (or RF call audio framework) may provide various operations such as decoding, filtering, amplification, and auto gain control (AGC) for audio data (e.g., first audio data) input through the microphone (250) of the electronic device and audio data (e.g., second audio data) received from a network or external device through the communication circuit (240). On the audio framework (550), the first audio data and the second audio data may be processed through independent channels and may be stored at least temporarily in the transmitting audio buffer (552) and the receiving audio buffer (554), respectively.

[0173] According to one embodiment, the electronic device may acquire first voice data of a first user of the electronic device in real time through a microphone (250) and store it at least temporarily in a transmitting audio buffer (552) (or a first audio buffer). Additionally, the electronic device may store at least temporarily in a receiving audio buffer (554) (or a second audio buffer) second voice data of a second user of an external device acquired based on data (or data packets) received from a network through a communication circuit (240).

[0174] According to one embodiment, the transmitting audio buffer (552) corresponds to a local audio channel, and the receiving audio buffer (554) corresponds to a remote audio channel. The local audio channel and the remote audio channel may be managed by the audio framework (550).

[0175] According to one embodiment, when a call recording function is activated, the electronic device may store first audio data of a local audio channel (or transmitting audio buffer (552)) and second audio data of a remote audio channel (or receiving audio buffer (554)) as recording data, and at least partially simultaneously, transmit the first audio data and the second audio data to an application (e.g., a call application (560)) to convert them into text. For example, the electronic device may transmit the first voice data stored in the transmitting audio buffer (552) and the second voice data stored in the receiving audio buffer (554) to a first process that performs operations related to the creation and storage of a recording file, and at least partially simultaneously, transmit the first voice data stored in the transmitting audio buffer (552) and the second voice data stored in the receiving audio buffer (554) to a second process that converts the first voice data stored in the transmitting audio buffer (552) and the second voice data stored in the receiving audio buffer (554) into text to generate call text data.

[0176] According to one embodiment, the electronic device may store call text data generated in the second process and recording data (570) generated in the first process in memory based on the termination of the recording function. For example, the electronic device may generate the generated recording data as a file having a specified extension (e.g., m4a) and store it in a specified folder in memory. Additionally, the electronic device may store the generated call text data in a database (562) corresponding to a call application.

[0177] FIGS. 6A and 6B illustrate a user interface screen that provides the content of a call as text after recording a call of an electronic device according to one embodiment.

[0178] According to one embodiment, an electronic device (200) (e.g., the electronic device (200) of FIG. 2) may provide a notification item (610) related to a call recording on a notification panel when a call recording is completed. The notification item (610) related to a call recording may include a text view item (612) that triggers the display of call text data. According to one embodiment, if multiple call recordings are performed in a single call session, the electronic device (200) may provide a notification item corresponding to each call recording.

[0179] According to one embodiment, the electronic device (200) may provide information related to call recording at the end of the call or at the end of the recording during the call.

[0180] According to one embodiment, the electronic device (200) can expand and display a notification panel based on touch and swipe inputs on a status bar displayed at the top of the screen. For example, the electronic device (200) can provide information such as application notifications, system notifications, communication connections, and media controls through the notification panel.

[0181] Referring to FIG. 6a, the electronic device (200) may display a notification item (610) for a call recording on a notification panel. The notification item (610) for a call recording includes information indicating that the call recording is complete and may include a text view item (612) and a delete item (614) selectable by the user. The electronic device (200) may delete the stored recording data and / or call text data based on a touch input to the delete item (614). The electronic device (200) may provide call text data (620), such as FIG. 6b, on the screen based on a touch input to the text view item (612).

[0182] Figure 6b illustrates a screen providing call text data for conversation 1.

[0183] According to one embodiment, the electronic device (200) can display call text data by dividing it into a plurality of text blocks. Referring to FIG. 6b, text blocks (621, 622, 623, 624) of first text data converted from voice spoken by the user of the electronic device (200) during a voice call can be displayed on the right side of the screen, and text blocks (626, 627, 628) of second text data converted from voice spoken by the call partner can be displayed on the left side of the screen. The electronic device (200) can distinguish visual effects such as color, shading, and font of each text block according to the speaker.

[0184] According to one embodiment, the electronic device (200) may display an indication (e.g., text or image) representing speaker information within or around a corresponding text block. For example, in the case of a multi-party call, indications indicating a corresponding speaker (opponent) may be displayed for each of the text blocks of the second text data.

[0185] According to one embodiment, the electronic device (200) may display an indication indicating utterance time information within or around a corresponding text block.

[0186] According to one embodiment, the electronic device (200) can distinguish call text data into individual text blocks and display them in a first area of ​​the display (230). For example, the first area may include the center area of ​​the display (230).

[0187] According to one embodiment, the electronic device (200) may display a summary item (690) for creating summary data of a call. Referring to FIG. 6b, the electronic device (200) may display a summary item (690) selectable by user input on a second area which is the top of a first area.

[0188] According to one embodiment, the electronic device (200) may generate a prompt including call text data and a request to generate a summary of the call content to request the AI ​​model to generate a summary of the voice call based on user input for a summary item (690), and transmit it to the AI ​​model. The AI ​​model may generate summary data by analyzing the contents of the prompt. According to one embodiment, the summary data may include a title, keywords, and / or a summary of the call content.

[0189] According to one embodiment, the electronic device (200) can output an audio signal of a text block through a speaker (or external audio device) based on user input to the text block.

[0190] According to one embodiment, the electronic device (200) may display an audio playback item (670) for audio playback of call text data. Referring to FIG. 6b, the electronic device (200) may display the audio playback item (670) on a third area which is the bottom of a first area. The audio playback item (670) may include an item that allows selecting play / pause, backward, and forward, and an item (672) that indicates a playback section in the entire recording data. According to one embodiment, when a specific call text block is selected and audio of that section is played, the item (672) indicating the playback section may be changed to indicate that section.

[0191] According to one embodiment, the audio playback item (670) is not provided when the audio signal is not playing, and may be provided when the audio signal is playing depending on the selection of the text block.

[0192] FIG. 7 is a flowchart of a method for an electronic device according to one embodiment to obtain summary data of a call using an AI model.

[0193] According to one embodiment, the illustrated method may be performed by an electronic device (e.g., the electronic device (200) of FIG. 2), and the technical features described above may be omitted from the description below.

[0194] According to one embodiment, in operation 710, the electronic device may receive user input selecting a summary of a voice call. For example, the electronic device may display a summary item (e.g., summary item (690) of FIG. 6b) for creating summary data of the call content along with call text data, and detect user input for the summary item.

[0195] According to one embodiment, the electronic device may perform an operation to generate summary data when recording ends without user input selecting a summary. In this case, operation 710 may be omitted.

[0196] According to one embodiment, in operation 715, the electronic device can determine whether the stored voice call is a multi-party call. Here, a multi-party call may mean a voice call performed between three or more users of the device, including the user of the electronic device. The electronic device can determine whether it is a multi-party call and the number of call participants based on voice call session information received through a call application and / or a network. When a voice call is initiated, the electronic device can set a flag indicating whether it is a multi-party call or a one-to-one call.

[0197] According to one embodiment, when a voice call is a multi-party call, the electronic device can distinguish voice data of multiple call partners received through a network based on voice analysis. In a multi-party call, the voice data of the first partner and the voice data of the second partner may be mixed on the network and transmitted to the electronic device. The electronic device can distinguish the voice data of each call partner by analyzing acoustic features (e.g., timbre, pitch, frequency spectrum) capable of distinguishing the speaker in the voice data received from the network. According to one embodiment, the electronic device can distinguish the voice data of multiple call partners by acquiring voice data stored in a received audio buffer in real time during a voice call, or by analyzing the recorded data after the recording is completed.

[0198] According to one embodiment, in operation 725, the electronic device may include text information indicating a user (e.g., first speaker information) corresponding to each text block corresponding to the voice of the electronic device user, and text information indicating each speaker corresponding to each text block corresponding to the voices of the call partners. For example, when the electronic device distinguishes the voices of call partners during a multi-party voice call, it may record data capable of distinguishing the speaker in the voice data or text data of each distinguished speaker (e.g., recorded in the header information of the text block), and in operation 725, based on the recorded data, it may include text information corresponding to each speaker (e.g., partner 1 and partner 2, or the name of each speaker) in the call text data.

[0199] According to one embodiment, when the voice call is a one-to-one call rather than a multi-party call, in operation 730, the electronic device may include text information indicating the user (e.g., first speaker information) corresponding to each text block corresponding to the voice of the electronic device user, and text information indicating the other party (e.g., second speaker information) corresponding to each text block corresponding to the voice of the call counterpart. In the case of a one-to-one call, since the voice data transmitted by the electronic device and the voice data received from the network can be distinguished by an audio framework (e.g., RF call audio framework), a voice analysis operation to distinguish the speaker may not be performed. The electronic device may record data capable of distinguishing the speaker in the voice data or text data of the user's call counterpart during the voice call (e.g., recorded in the header information of the text block), and in operation 730, based on the recorded data, may include text information corresponding to the user and the counterpart (e.g., me and the counterpart) in the call text data.

[0200] According to one embodiment, in operation 740, the electronic device can determine whether the amount of call text data is processable by the AI ​​model.

[0201] According to one embodiment, an electronic device can generate summary data using an on-device AI model operated by the electronic device. The on-device AI model may be lightweight compared to a server-based AI model in order to operate efficiently on the electronic device. The AI ​​model can process input data in token units, and the maximum number of tokens that can be processed for each model may be determined by language, and the size of the tokens that can be processed for the on-device AI model may be relatively smaller compared to a server-based AI model.

[0202] According to one embodiment, in operation 745, the electronic device may divide call text data into multiple chunks according to the maximum throughput of the AI ​​model. For example, the electronic device may calculate the sum of the number of tokens of sequential text blocks, form a text block within a range that does not exceed the maximum number of tokens that can be processed by the on-device AI model of the electronic device into one chunk, and form multiple text blocks from the next text block into another chunk.

[0203] According to another embodiment, the electronic device can divide text blocks into multiple chunks based on semantic analysis of call text data.

[0204] According to one embodiment, if the number of tokens in the call text data is less than the maximum throughput of the AI ​​model, the electronic device can determine the entire call text data as one chunk.

[0205] According to one embodiment, in operation 750, the electronic device may generate a prompt requesting a summary of the call content and transmit it to an AI model. For example, the electronic device may generate a prompt including call text data and a request to generate a summary of the call content, transmit it to an AI model, and obtain the summary data of the call content as a response from the AI ​​model.

[0206] According to one embodiment, when an electronic device divides call text data into multiple chunks, it can generate a prompt for each chunk.

[0207] According to one embodiment, the electronic device may record speaker information for each text block of call text data included in the prompt. For example, the electronic device may record the user (or I) as text information as the speaker in the first text data converted from the utterance of the electronic device's user, and record the counterparty as text information as the speaker in the second text data converted from the utterance of the counterparty. Additionally, in the case of a multi-party call, the electronic device may record information of the counterparties identified through voice analysis in each text block. In this way, when speaker information is recorded, the quality of the summary data generated by the AI ​​model can be improved compared to when a transcript, which converts the call content into text without distinguishing the speaker, is transmitted to the AI ​​model.

[0208] According to one embodiment, in operation 755, the electronic device can obtain summary data from the AI ​​model.

[0209] According to one embodiment, the summary data may include a title, keywords, and / or a summary of the call content.

[0210] According to one embodiment, an electronic device may display summary data together with call text data on a display. For example, the electronic device may display call text data in a first area including the center area of ​​the display and display summary data in a second area above the first area.

[0211] According to one embodiment, instructions for performing each operation constituting the method may be stored on a computer-readable recording medium. The recording medium may be tangible and non-transitory. The recording medium may store one or more computer programs containing said instructions.

[0212] FIGS. 8A and 8B illustrate a user interface screen providing call content and a summary thereof of an electronic device according to one embodiment.

[0213] According to one embodiment, an electronic device (200) (e.g., the electronic device (200) of FIG. 2) can convert voice data (or recording data) (e.g., first voice data of the electronic device user, second voice data of the external device user) into text in real time through a separate process when recording of a voice call begins.

[0214] The transcript of the conversation 1 between the user of the electronic device (200) and the call partner of the external device, converted into text, was previously explained through Table 1, and the transcript of the conversation 2, converted into text, was previously explained through Table 2.

[0215] According to one embodiment, an electronic device (200) can generate summary data of a call using an AI model based on user input for a summary item (e.g., summary item (690) of FIG. 6b). The electronic device (200) can generate a prompt including call text data and a request to generate a summary of the call content, transmit it to an AI model (e.g., an on-device AI model), and obtain summary data of the call content as a response from the AI ​​model.

[0216] According to one embodiment, the electronic device (200) can record speaker information for each text block of call text data included in the prompt. For example, the electronic device (200) can record the user as the speaker in the first text data converted from the utterance of the user of the electronic device (200), and record the other party as the speaker in the second text data converted from the utterance of the call partner.

[0217] The call text data recording the speaker in the above conversation 1 has been explained in Table 3 above, and the call text data recording the speaker in the above conversation 2 has been explained in Table 4 above.

[0218] Referring to Tables 3 and 4, the electronic device (200) can record the speaker, either me or the other party, for each text block. The electronic device (200) can generate a prompt containing call text data with speaker information recorded as in Table 3 or Table 4 and transmit it to an AI model.

[0219] According to one embodiment, an AI model can generate summary data by analyzing the content of a prompt. According to one embodiment, the summary data may include a title, keywords, a summary of the call content, and / or an image (representative image) representing the topic of the conversation.

[0220] The summary data generated for the call text data of the above conversation 1 was explained earlier through Table 6, and the summary data generated for the call text data of the above conversation 2 was explained earlier through Table 7.

[0221] According to one embodiment, the electronic device (200) can provide summary data obtained using an AI model through a display (230). For example, the electronic device (200) can display call text data (820, 840) in a first area including the center area of ​​the display (230), and display summary data items (830, 850) including summary data in a second area at the top of the first area.

[0222] FIG. 8a illustrates a screen displaying call text data (820) and summary data (830) converted from a conversation 1 between a user of an electronic device (200) and a call partner of an external device.

[0223] Referring to FIG. 8a, the electronic device (200) can display the call text data of conversation 1 by dividing it into individual text blocks. For example, text blocks (821, 822, 823, 824) generated from the utterance of the user of the electronic device (200) may be displayed on the right side of the first area, and text blocks (826, 827, 828) generated from the utterance of the call partner may be displayed on the left side of the first area.

[0224] According to one embodiment, the electronic device (200) may display a summary data item (830) in an area different from the first area where call text data (820) is displayed (e.g., a second area). Referring to FIG. 8a, the electronic device (200) may display the title "Promise to buy lunch and a coat after studying" (832) and the keywords "study, lunch, coat, gift" (834) of conversation 1 obtained using an AI model on the summary data item (830). The summary data item (830) may further include a view more item (836). If the user selects the view more item (836), the electronic device (200) may display a summary of conversation 1. According to another embodiment, the electronic device (200) may display the title, keywords, and summary all on the summary data item (830) when displaying the summary data without a selection process for the view more item (836).

[0225] According to one embodiment, the electronic device (200) may display an audio playback item (870) for audio playback of call text data in a third area, which is the bottom of the first area. The audio playback item (870) may include an item that allows selecting play / pause, backward, and forward, and an item that indicates a playback section in the entire recording data. According to one embodiment, when any one of the text blocks is selected, the electronic device (200) may output the audio data of the corresponding text block through a speaker (or external audio device) and control audio playback based on user input for the audio playback item (870).

[0226] According to one embodiment, an electronic device (200) may provide a visual effect to a text block corresponding to the text selected in response to user input among text blocks of call text data (820) when a title (832), keyword (834), or part of text of a summary (e.g., word) displayed within a summary data item (830) is selected based on user input. For example, if "study" among the keywords is selected based on user input, the electronic device (200) may highlight a text block containing "study" (e.g., 822, 828). In this case, the electronic device may move to a playback section of recording data corresponding to the selected text. For example, if "study" among the keywords is selected based on user input, the electronic device may play audio of a playback section corresponding to text block 822.

[0227] According to one embodiment, the electronic device (200) can display text of high importance (e.g., words) among the title, keywords, and summary of the summary data so as to be visually distinguished from other texts. For example, if an AI model conveys "lunch" and "coat" as important words in the title among the information included in the summary data, the electronic device (200) can highlight "lunch" and "coat" in the title and keywords of the summary data area (834) and highlight and display "lunch" and "coat" in the text block.

[0228] FIG. 8b illustrates a screen displaying converted call text data and summary data of a conversation 2 between a user of an electronic device (200) and a call partner of an external device.

[0229] Referring to FIG. 8b, the electronic device (200) can display the call text data (840) of conversation 2 by dividing it into individual text blocks. For example, text blocks (841, 842, 843) generated from the utterance of the user of the electronic device (200) may be displayed on the right side of the first area, and text blocks (846, 847, 848) generated from the utterance of the call partner may be displayed on the left side of the first area.

[0230] According to one embodiment, the electronic device (200) may display a summary data item (850) in a different area (e.g., a second area) from the first area where call text data (840) is displayed. Referring to FIG. 8b, the electronic device (200) may display the title of Conversation 2, "Sharing Player Kim's Great Game and Baseball Passion with the Other Party" (852), and the keywords "Baseball, Player Kim, Walk-off Home Run, Baseball Practice Field" (854), obtained using an AI model, in the summary data item. The summary data item (850) may further include a view more item (856). If the user selects the view more item (856), the electronic device (200) may display a summary of Conversation 2. According to another embodiment, the electronic device (200) may display the title, keywords, and summary all on the summary data item when displaying the summary data without a selection process for the view more item (856).

[0231] According to one embodiment, the electronic device (200) can display an audio playback item (870) for audio playback of call text data in a third area at the bottom of a first area.

[0232] FIGS. 9a, FIGS. 9b, and FIGS. 9c illustrate a user interface screen providing call content and a summary thereof of an electronic device according to one embodiment.

[0233] According to one embodiment, an electronic device (200) (e.g., the electronic device (200) of FIG. 2) can provide a summary data item containing the contents of summary data obtained using an AI model through a display (230).

[0234] According to one embodiment, the electronic device (200) may display the title (832, 852) and keywords (834, 854) of the call content in summary data items (830, 850) as in FIG. 8a or FIG. 8b, and may display a more item (836, 856). The electronic device (200) may display a summary of the call content based on user input for the more item.

[0235] Referring to FIG. 9a, the electronic device (200) may display the call text data (920) of Conversation 1 in a first area and the summary data item (930) in a second area. The summary data item (930) may include the title of Conversation 1, "Promise to buy lunch and a coat after studying" (932), keywords, "study, lunch, coat, gift" (934), and the summary, "Promise lunch, a coat gift, and additional rewards to encourage and motivate the other person to study English" (938).

[0236] According to another embodiment, the electronic device (200) can display a title (932), a keyword (934), and a summary (938) directly on a summary data item (930) as in FIG. 9a, without selecting a more item (836, 856) from a summary data item as in FIG. 8a or 8b.

[0237] According to one embodiment, the electronic device (200) may display an audio playback item (970) for audio playback of call text data (920) in a third area at the bottom of a first area. The audio playback item (970) may include an item that allows selecting play / pause, backward, and forward, and an item that indicates a playback section in the entire recording data. According to one embodiment, when any one of the text blocks is selected, the electronic device (200) may output the audio data of the corresponding text block through a speaker (or external audio device) and control audio playback based on user input for the audio playback item (970).

[0238] According to one embodiment, the electronic device (200) may provide items (941, 942, 943) related to the control of call text data or summary data while summary data is displayed in a summary data area (930) displayed on a display (230). For example, the electronic device (200) may display items capable of performing control functions, such as deleting, editing, or regenerating call text data and summary data, in response to user input for a more item (e.g., the more item (910) of FIG. 9a) while recording data is displayed.

[0239] Referring to FIG. 9b, the electronic device (200) may provide a text conversion delete item (941), a summary delete item (942), and a re-summary item (943) when receiving user input for a more item (910). When the text conversion delete item (941) is selected according to user input, the electronic device (200) may delete the call text data converted from the first audio data and the second audio data of the voice call. In this case, the display of the call text data (920) of the corresponding call displayed on the display (230) may be removed. Additionally, when the summary delete item (942) is selected according to user input, the electronic device (200) may delete the summary data displayed in the summary data area (930). Additionally, when the re-summary item (943) is selected according to user input, the electronic device (200) may transmit the call text data to an AI model to generate new summary data.

[0240] According to one embodiment, the electronic device (200) may provide an action item related to the summary content of the summary data. For example, the action item may be provided to perform actions that can be provided on the application of the electronic device (200) in relation to the keywords and / or content of the summary of the summary data, such as adding a schedule. If the summary data includes a date related to a schedule, the electronic device (200) may provide an action item that can add a schedule related to call content on that date.

[0241] Referring to FIG. 9c, the call content may include information such as a review meeting at 9:00 AM tomorrow and a report at 1:00 PM the day after tomorrow, and the electronic device (200) may display an action item (955) on the summary data area (950) that can add these schedules to a schedule application. When the corresponding action item (955) is selected, the electronic device (200) may add the schedule of the summary data to the schedule application.

[0242] FIG. 10 illustrates a user interface screen providing a list of call records of an electronic device according to one embodiment.

[0243] According to one embodiment, an electronic device (200) (e.g., the electronic device (200) of FIG. 2) may provide at least a portion of summary data corresponding to a call record on a call list screen that includes at least one call record. For example, the electronic device (200) may display a title of the summary data for a call record for which summary data was generated on the call list screen.

[0244] Referring to FIG. 10, four call records may be displayed on the call list screen. Among these, call recordings were performed on the first and third calls, and summary data of the call content may be generated from the first call.

[0245] According to one embodiment, when a specific call record is selected on a call list screen, the electronic device may provide detailed information of the corresponding call record. Referring to FIG. 10, when a first call record (1010) is selected, the electronic device may display call information (1012) (e.g., other party's phone number, whether the call was made or received, call duration, call time), a call recording item (1025), and summary data (1020). In FIG. 10, only the title of the summary data is displayed, but is not limited thereto, and at least some of the keywords or summary text may also be displayed.

[0246] Referring to FIG. 10, in the case of a third call record (1030) for which summary data was not generated, only the call recording item (1035) may be displayed without summary data.

[0247] According to one embodiment, the electronic device may display information (1060) that guides a spam call blocking function on a call list screen. For example, regarding a call or text message among incoming calls or text messages that is determined to be spam, the electronic device may display information (1060) that guides a function to automatically block spam along with displaying caller information.

[0248] FIGS. 11a, FIGS. 11b, and FIGS. 11c illustrate a process in which an electronic device according to one embodiment verifies information of a call participant and reflects it in a summary of the call content.

[0249] According to one embodiment, an electronic device (e.g., the electronic device (200) of FIG. 2) may include user information of a call participant in a prompt requesting the generation of summary data of the call content. For example, the electronic device may include user information obtained from an application or user information obtained through voice analysis in the prompt and transmit it to an AI model.

[0250] FIG. 11a illustrates a screen showing information about a specific user in a contact application (1100).

[0251] Referring to FIG. 11a, the name (1110) (e.g., lee), workplace information (1120) (e.g., employee, Department A, Samsung), and relationship (1130) (e.g., family) of a specific user may be stored in the contact application (1100). When generating summary data after recording a voice call, the electronic device (200) may check the user information (e.g., name, workplace information, relationship) of the call partner stored in the contact application (1100) and include the user information in a prompt for a summary request and transmit it to an AI model.

[0252] Figure 11b illustrates a screen displaying the call text data and summary data of conversation 2.

[0253] Referring to FIG. 11b, the electronic device (200) can display text blocks (1141, 11442, 1143) of first text data obtained from voice data of the user of the electronic device (200) and text blocks (1146, 1147, 1148) of second text data obtained from voice data of the call counterpart. Additionally, the electronic device (200) can display summary data obtained from an AI model. The summary data item (1150) includes a title (1152) and keywords (1154) of the call content, and can display a more item (1156) that allows the summary text to be displayed when selected.

[0254] According to one embodiment, the electronic device (200) can obtain user information of the user or the counterpart based on the analysis of voice data of the user or the counterpart of the electronic device (200) in call recording data. For example, the electronic device (200) can obtain user information such as gender, age group, region, emotion, and / or surrounding environment based on the analysis of characteristics of the voice data such as frequency band, intensity, tone, speed, and / or pronunciation.

[0255] Referring to FIG. 11b, the electronic device (200) can confirm that the call partner is a male in his 20s based on the recognition of the call partner's voice corresponding to the data blocks (1146, 1147, 1148) of the second text data.

[0256] According to one embodiment, the electronic device (200) may include the user's name "Lee" obtained from the contact application (1100) and information obtained based on voice recognition, such as "male in his 20s," in a prompt for a call summary request and transmit it to an AI model. According to one embodiment, the electronic device (200) may record user information (e.g., name) in information (e.g., me, other party) for identifying the speaker in the call text data.

[0257] According to one embodiment, an AI model can generate summary data of the call content based on call text data and user information included in a prompt.

[0258] Table 9 shows summary data generated by the AI ​​model considering user information (e.g., Lee, male in his 20s) in Conversation 2.

[0259] Title: Sharing Player Kim's Great Game and Passion for Baseball with Lee Summary: Formed a bond with Lee, a man in his 20s, while discussing Player Kim's walk-off home run. Lee mentioned that women in their 20s would also likely like an attractive player like Kim and requested an interview video. Afterwards, Lee and I decided to go to the baseball practice range, expressing our determination to hit a home run just like Kim. Keywords: Baseball, Player Kim, Female Fan in 20s, Walk-off Home Run, Baseball Practice Range

[0260] When comparing summary data generated without considering user information (e.g., Table 6) and summary data generated with considering user information (e.g., Table 9) for the content of the same conversation 2, the other party in the title (1152) is changed to the other party's name, Lee, and in the summary (1158), my peers are changed to "male in their 20s" and "female in their 20s" to reflect the age group, and "female fan in their 20s" can be added to the keyword (1154). Referring to FIG. 11c, the electronic device (200) can display the title (1152) and keyword (1154) in the summary data item (1150), and display the summary (1158) generated by reflecting user information. FIG. 12a and FIG. 12b illustrate a user interface screen that provides the call content of each recording file when a plurality of recording files of the electronic device according to one embodiment are stored.

[0261] According to one embodiment, the electronic device (200) can enable or disable call recording based on user input regarding a call recording item during a voice call. Accordingly, multiple recording files may be stored for a single voice call.

[0262] According to one embodiment, when the electronic device (200) provides call text data and summary data corresponding to a plurality of recording files generated in one call, it may provide a user interface that allows easy switching to a display of call text data and summary data of another recording file of the call.

[0263] Referring to FIG. 12a, the electronic device (200) may display a call list screen (1210) containing three call records (1212, 1214, 1220) on a display (230). Among these, if the user selects the third call record (1220), detailed information of the call record may be displayed. The third call may have multiple recording files stored. The electronic device (200) may provide call information (1222) (e.g., call time, whether it was outgoing / incoming, call duration), a call recording item (1224), summary data (1230), a recording data playback item (1226), and information (1228) indicating that multiple recording files are stored for the third call record.

[0264] Referring to FIG. 12b, the electronic device (200) can display call text data (1240) and summary data (1250). For example, the summary data item (1250) may display a title (1252), keywords (1254), and a more item (1256) of the call content, and the summary may be displayed when the more item (1256) is selected.

[0265] According to one embodiment, the electronic device (200) may display an indicator (1260) that can switch to a display of call text data and summary data of another recording file of the call. Referring to FIG. 12b, the indicator (1260) may include information indicating the order of the current recording file and a button that can switch to another recording file. For example, if the left button is pressed while the call text data and summary data of the third recording file among the current three recording files are displayed, the call text data and summary data of the second recording file may be displayed. According to another embodiment, the electronic device (200) may be configured to switch to a display of call text data and summary data of another recording file based on a swipe input in the left or right direction.

[0266] FIG. 13 illustrates a user interface screen in which an electronic device according to one embodiment provides summary data for a video call.

[0267] According to one embodiment, the electronic device (200) can perform a video call with an external device using a call application or another application that provides a video call function (e.g., messenger, SNS). When the recording function is activated during a video call, the electronic device (200) can store video call data including video data of the video call and voice data synchronized with the video data. For example, when the recording function is activated during a video call, the electronic device (200) can generate and store in memory video call data including first voice data of the user of the electronic device (200) obtained through a microphone, first video data obtained through a camera, and second voice data and second video data of the user of the external device obtained based on data received from a network through a communication circuit. The electronic device (200) can convert the first voice data and the second voice data into text and generate call text data by mapping the first speaker information of the user of the electronic device (200) and the second speaker information of the user of the external device.

[0268] According to one embodiment, an electronic device (200) can generate summary data based on call text data and video data of a video call (e.g., first video data, second video data). The summary data generated during a video call may include a title, keywords, a summary, and a summary video of the call content. The electronic device (200) can generate a prompt including a request to generate call text data, video call data, and a summary of the call content, and transmit it to an AI model, and receive summary data from the AI ​​model.

[0269] According to one embodiment, the AI ​​model may be a generative AI model (or a multimodal language model (MMLM)) trained to generate a title, keywords, a summary, and / or a summary video of a call based on video data (e.g., first video data, second video data) and call text data.

[0270] According to one embodiment, the AI ​​model may generate summary data (1330) based further on user information of the electronic device (200) and / or external device obtained from an application (e.g., a contact application) of the electronic device (200), and / or user information of the electronic device (200) and / or external device obtained based on the analysis of video data and / or voice data of video call data. For example, the AI ​​model may generate summary data (1330) using speaker-to-speaker relationship information obtained based on contact information of the call counterpart, voice or video data.

[0271] According to one embodiment, an AI model can extract at least one image frame from image data (e.g., first image data, second image data) based on image analysis and generate summary data (1330) including a caption for the image frame.

[0272] According to one embodiment, the summary video may include at least one image frame of a specified length time interval before or after an image frame related to the topic (or keyword) of the summary data (1330). The summary video may include a highlight section within the entire duration of the video call. Here, the highlight section may be a significant moment extracted from the entire frames of the video data based on an analysis of voice data mapped to the image frame (e.g., speech content, tone of voice, speaker's emotion).

[0273] FIG. 13 illustrates a user interface that displays call text data, summary video, and summary data (1330) on an electronic device (200) for a video call between a grandmother, who is a user of the electronic device (200), and a grandson, who is a user of an external device.

[0274] Referring to FIG. 13, the electronic device (200) can confirm that the call partner is a grandson and that the relationship between the speakers is that of a grandmother and a grandson, based on the analysis of contact information (e.g., grandson) or the second video data and / or second voice data of an external device of video call data. The electronic device can generate and display a title (1332) displayed on the summary data area (1330) as "Grandson and Grandmother's Greetings," including the relationship between the speakers.

[0275] According to one embodiment, the electronic device (200) may display a summary image (1325) together with a text block of call text data (1320) in a first area of ​​the display (230). The summary image (1325) may include at least one image frame of a highlight section generated based on video analysis of a video call, and may be a single image frame or a video.

[0276] According to one embodiment, the summary data (1330) may include "the grandson makes a heart gesture above his head" as a description (1335) of the summary video (1325).

[0277] FIG. 14 illustrates a user interface screen in which an electronic device (200) according to one embodiment provides a summary of a plurality of calls.

[0278] According to one embodiment, the electronic device (200) can perform multiple voice calls with the same external device (or call partner) and record multiple voice calls. For example, a user can have a follow-up conversation related to a previous conversation with a specific partner through a voice call. According to one embodiment, the electronic device (200) can analyze the call text data of each voice call to determine whether multiple calls are related to each other based on the identity or association between conversation topics, the time interval between each call, etc.

[0279] Referring to FIG. 14, Conversation 3 (1403) may correspond to a voice call performed after the voice call of Conversation 1 (1401) with the same counterpart as Conversation 1 (1401). Referring to FIG. 14, Conversation 3 (1403) may include a voice in which a user says to a counterpart, "I heard rumors that your test scores were good this time; your scores went up a lot, right? Congratulations," a voice in which a counterpart says, "Yeah, fortunately, my scores went up quite a bit this time. It really paid off to study hard," and a voice in which the user says, "I should keep my promise from last time. I'll buy you a gift. Let's go to the department store this weekend." According to one embodiment, the electronic device (200) may determine that Conversation 3 (1403) is a conversation associated with Conversation 1 (1401) based on at least one of the voice, speaker information, text, summary, keywords, or topics of Conversation 1 (1401) and Conversation 3 (1403).

[0280] According to one embodiment, the electronic device (200) can convert speaker-specific voice data (e.g., first voice data of a transmitting audio buffer, second voice data of a receiving audio buffer) separated on an audio framework into text information and generate call text data by mapping each speaker information.

[0281] According to one embodiment, the electronic device (200) can determine the relationship between conversations of each voice call based on the participants of the voice call, keywords in the conversation, and / or the time interval from the previous conversation. For example, the electronic device (200) can determine that conversation 3 (1403) is a conversation associated with conversation 1 (1401) through information such as the counterpart of conversation 3 (1403) and keywords such as grades, studying, gift, and appointment.

[0282] According to one embodiment, the electronic device (200) may further utilize at least a portion of conversation 1 (1401) when generating summary data of conversation 3 (1403) based on the determination that conversation 3 (1403) is a conversation associated with conversation 1 (1401). For example, the electronic device (200) may pass at least a portion of the call text data of conversation 1 (1401) and / or summary data together with the call text data of conversation 3 (1403) to an AI model and obtain summary data (1430) of conversation 3 (1403) generated by referencing conversation 1 (1401) from the AI ​​model. Referring to FIG. 14, the summary for Conversation 3 (1403) may include content such as the summary corresponding to the content of Conversation 1 (1401), “Promised to buy A a coat at a department store if he studies hard during the phone call on December 1, 2024” (1432), and the summary corresponding to the content of Conversation 3 (1403), “A’s grades have improved a lot, so we decided to meet this weekend to buy the gift promised last time” (1434). According to one embodiment, the summary may include content such as the summary corresponding to the combined content of Conversation 1 and Conversation 3, “A’s grades have improved a lot, so we decided to meet A this weekend to buy the gift promised during the previous phone call on December 1.”

[0283] According to one embodiment, the electronic device (200) may display a button associated with the record of conversation 3 (e.g., summary data (1430)) that can be moved to the record of the previous conversation (e.g., conversation 1) (e.g., recording file, call text, or summary).

[0284] According to one embodiment, the electronic device (200) can generate summary data of the current voice call based on the conversation of the previous voice call as well as messages sent and received with the same counterpart in other applications (e.g., messages, messenger, SNS).

[0285] FIG. 15 illustrates a user interface screen in which an electronic device according to one embodiment provides a summary format corresponding to a conversation topic.

[0286] According to one embodiment, the electronic device (200) can generate summary data in a format corresponding to the content of the call text data (or voice data) of a voice call. For example, the electronic device (200) can identify the content or topic of the conversation from the call text data and select a format of summary based on the identified content or topic. Here, the format of the summary may include a format consisting of one or more consecutive sentences, a format consisting of bullet points, a format using a written style, or a format using a short and simple style. Additionally, the format of the summary may include a format including numbers to indicate order, a format including bullet points to indicate item-by-item separation, a format including different bullet points to indicate item-by-item hierarchy, a tree structure to indicate item-by-item association, a graph to indicate changes, and a map image to indicate paths.

[0287] Referring to FIG. 15 (a), the user can have a conversation with the other party about a meeting schedule (e.g., conversation 4) via voice call. If the electronic device (200) determines through the analysis of the call text data (or voice data) that the topic of conversation 4 is about a meeting schedule, it can generate and display a summary (1510) summarized as an itemized list of time, place, attendees, and preparations.

[0288] Referring to FIG. 15 (b), the user can have a conversation about a cooking recipe (e.g., conversation 5) via voice call. If the electronic device (200) determines through the analysis of the call text data (or voice data) that the topic of conversation 5 is about a cooking recipe, it can generate and display a summary (1520) in the form of a step-by-step sequence, such as ingredients, tools, and step-by-step precautions (time, quantity, temperature).

[0289] Referring to Fig. 15 (c), the user can have a conversation (e.g., conversation 6) with the other party about a task (e.g., errand) through a voice call. If the electronic device (200) determines through the analysis of the call text data (or voice data) that the topic of conversation 6 is about a task, it can generate and display a summary (1530) that summarizes the call text data based on guidelines on what, when, and how to perform the task.

[0290] According to one embodiment, an electronic device (200) may include candidate categories corresponding to the conversation of a voice call and candidate formats mapped to each category in a prompt for generating a summary of the call content. When an AI model receives the prompt from the electronic device (200), it may generate summary data based on call text data using at least one of the categories and / or formats included in the prompt.

[0291] FIG. 16 is a flowchart of a method in which an electronic device according to one embodiment provides a call summary.

[0292] According to one embodiment, the illustrated method may be performed by an electronic device (e.g., the electronic device (200) of FIG. 2), and the technical features described above may be omitted from the description below.

[0293] According to one embodiment, in operation 1610, the electronic device may perform a call with an external device. For example, the call may include a voice call or a video call.

[0294] According to one embodiment, in operation 1620, the electronic device can acquire first voice data of a first user of the electronic device through a microphone. For example, the first voice data acquired in real time through the microphone may be temporarily stored in a transmission audio buffer of an audio framework and may be transmitted from the transmission audio buffer to a process for performing call recording and a process for performing text translation and summary generation.

[0295] According to one embodiment, in operation 1630, the electronic device may acquire second voice data of a second user corresponding to an external device based on data received from an external device through a communication circuit. For example, the electronic device may receive a data packet (e.g., a real-time transport protocol (RTP) packet) from a network via cellular wireless communication and acquire second voice data of a call partner through decoding, voice signal processing, and / or a digital-to-analog converter (DAC) process for the data packet. The second voice data may be temporarily stored in a receiving audio buffer of an audio framework and may be transferred from the receiving audio buffer to a process for performing call recording and a process for performing text translation and summary generation.

[0296] According to one embodiment, operation 1620 and operation 1630 can be performed at least partially simultaneously during a call between the electronic device and an external device.

[0297] According to one embodiment, the electronic device can perform the recording and summarization of voice data through a separate process.

[0298] According to one embodiment, in operation 1640, the electronic device may generate call text data by mapping first speaker information corresponding to a first user to first text data converted from first voice data, and mapping second speaker information corresponding to a second user to second text data converted from second voice data. For example, the processor (210) may record text indicating a user of the electronic device (e.g., me, you) in the first text data and record text indicating a user of an external device (e.g., the other person, the other) in the second text data.

[0299] According to one embodiment, the generation of call text data may be performed based at least in part on the activation of a call recording function. For example, when call recording is initiated, call text data may be generated at least in part simultaneously with the generation of call recording data. For example, at least part of the call text data may be generated during the call. For example, call text data may be generated before the call ends.

[0300] According to one embodiment, in operation 1650, the electronic device can generate summary data using call text data.

[0301] According to one embodiment, an electronic device can generate summary data using an AI model based on call text data. The AI ​​model may be a generative AI model or a large language model (LM) trained to generate titles, keywords, images, and / or summaries of call content based on call text data. The electronic device can transmit a prompt to the AI ​​model that includes call text data and a request to generate a summary of the call content, and obtain summary data as a response from the AI ​​model.

[0302] According to one embodiment, the summary data may include a title, keyword, and summary of the call content. If the call currently being performed is a video call, the summary data may further include a summary video of the video call and a description thereof.

[0303] According to one embodiment, in operation 1660, the electronic device may provide summary data in relation to a call through a display.

[0304] According to one embodiment, instructions for performing each operation constituting the method may be stored on a computer-readable recording medium. The recording medium may be tangible and non-transitory. The recording medium may store one or more computer programs containing said instructions.

[0305] An electronic device according to various embodiments of the present document may include a display, a microphone, a communication circuit, a memory, and at least one processor.

[0306] According to one embodiment, the memory may be executed by at least one processor, and may store instructions for the electronic device to perform a call with an external device, to acquire first voice data of a first user of the electronic device through the microphone during the call, and to acquire second voice data of a second user corresponding to the external device based on data received from the external device through the communication circuit during the call.

[0307] According to one embodiment, the memory may store instructions for the electronic device to generate call text data by mapping first speaker information corresponding to the first user to first text data converted from the first voice data, and mapping second speaker information corresponding to the second user of the external device to second text data converted from the second voice data.

[0308] According to one embodiment, the memory may store instructions for the electronic device to generate summary data of the call using the call text data and to provide the summary data through the display in relation to the call.

[0309] According to one embodiment, the memory may store instructions such that the electronic device stores first voice data acquired in real time through the microphone in a first audio buffer, generates second voice data using data received in real time from the external device through the communication circuit and stores it in a second audio buffer, and, when a recording function for the call is activated, generates recording data based on the first voice data stored in the first audio buffer and the second voice data stored in the second audio buffer.

[0310] According to one embodiment, the memory may store instructions that cause the electronic device to generate the call text data using the first voice data stored in the first audio buffer and the second voice data stored in the second audio buffer at least partially simultaneously with the generation of the recording data.

[0311] According to one embodiment, the memory may store instructions for the electronic device to transmit the first voice data stored in the first audio buffer and the second voice data stored in the second audio buffer to a first process to generate the recording data, and at least partially simultaneously with the generation of the recording data, transmit the first voice data stored in the first audio buffer and the second voice data stored in the second audio buffer to a second process to generate the call text data.

[0312] According to one embodiment, the memory may store instructions that cause the electronic device to store the call text data generated in the second process and the recording data generated in the first process in the memory based on the termination of the recording function.

[0313] According to one embodiment, the memory may store instructions that cause the electronic device to generate at least a portion of the call text data before the call is terminated.

[0314] According to one embodiment, the memory may store instructions that cause the electronic device to generate the summary data using an AI model based on the call text data.

[0315] According to one embodiment, the memory may store instructions that cause the electronic device to transmit a prompt to the AI ​​model, including a request to generate the call text data and a summary of the call content, and to obtain the summary data as a response from the AI ​​model.

[0316] According to one embodiment, the memory may store instructions that cause the electronic device to divide the call text data into a plurality of chunks based on the number of tokens supported by the AI ​​model, and to generate a plurality of prompts including each divided chunk to request the AI ​​model to generate a summary.

[0317] According to one embodiment, the memory may store instructions for the electronic device to transmit to the AI ​​model information related to the first user and / or information related to the second user obtained based on the application or voice analysis of the recorded data, and to obtain the summary data generated from the AI ​​model based on the information related to the first user and / or information related to the second user.

[0318] According to one embodiment, the AI ​​model may be an on-device AI model.

[0319] According to one embodiment, the memory may store instructions that cause the electronic device to identify information related to the second user based on information stored in the memory, and to determine the second speaker information based on the identified information related to the second user.

[0320] According to one embodiment, the summary data may include at least one of a title, keyword, summary, or image regarding the content of the call.

[0321] According to one embodiment, the memory may store instructions for the electronic device to distinguish voice data of multiple call partners of the multi-party call based on voice analysis of voice data received from a network through the communication circuit when the call is a multi-party call, and to generate summary data of the multi-party call using text data converted from each of the distinguished voice data and speaker information of each user.

[0322] According to one embodiment, the memory may store instructions for the electronic device to distinguish the call text data into a plurality of text blocks and display them in a first area of ​​the display, to display the summary data in a second area of ​​the display, and to display a user interface for controlling the playback of the recording data in a third area of ​​the display.

[0323] According to one embodiment, the memory may store instructions that, when any of the texts of the summary data displayed on the display are selected based on user input, provide a visual effect for a text block among the text blocks corresponding to the selected text, and / or move to a playback section of the recording data corresponding to the selected text.

[0324] According to one embodiment, the memory may store instructions that, when the call is terminated, the electronic device checks a previous call and / or message associated with the terminated call, and generates call text data of the terminated call, call text data or summary data of the previous call, and / or summary data corresponding to the terminated call using the message.

[0325] According to one embodiment, the memory may store instructions that cause the electronic device to classify the content of the call based on the call text data and, based on the classification, determine the format of the summary data.

[0326] According to one embodiment, the memory may store instructions that cause the electronic device to generate summary data including a summary image based on the call text data and the call video data of the video call when the call is a video call.

[0327] A method performed by an electronic device according to various embodiments of the present document may include: performing a call with an external device; acquiring first voice data of a first user of the electronic device through a microphone of the electronic device during the call; acquiring second voice data of a second user corresponding to the external device based on data received from the external device during the call; mapping first speaker information corresponding to the first user to first text data converted from the first voice data, and mapping second speaker information corresponding to the second user of the external device to second text data converted from the second voice data to generate call text data; generating summary data of the call using the call text data; and providing the summary data in relation to the call.

[0328] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.

[0329] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0330] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0331] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0332] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0333] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device, display; microphone; Communication circuit; Memory; and It includes at least one processor, The above memory can be executed by at least one processor, and at the time of execution, the electronic device, Performing calls with external devices, During the above call, first voice data of a first user corresponding to the electronic device is obtained through the microphone, and During the above call, based on data received from the external device through the communication circuit, second voice data of the second user corresponding to the external device is obtained, and A first speaker information corresponding to the first user is mapped to the first text data converted from the first voice data, and a second speaker information corresponding to the second user of the external device is mapped to the second text data converted from the second voice data to generate call text data. Using the above call text data, generate summary data of the above call, and An electronic device that stores instructions for providing the above summary data through the display in relation to the above call.

2. In Paragraph 1, The above memory is, the electronic device, First voice data acquired in real time through the above microphone is stored in a first audio buffer, and Using data received in real time from the external device through the communication circuit, the second voice data is generated and stored in the second audio buffer, and An electronic device that stores instructions for generating recording data based on the first voice data stored in the first audio buffer and the second voice data stored in the second audio buffer when the recording function for the above call is activated.

3. In Paragraph 2, The above memory is, the electronic device, At least partially simultaneously with the generation of the above recording data, the call text data is generated using the first voice data stored in the first audio buffer and the second voice data stored in the second audio buffer, and The first voice data stored in the first audio buffer and the second voice data stored in the second audio buffer are transferred to the first process to generate the recording data, and An electronic device storing instructions for generating the call text data by transmitting the first voice data stored in the first audio buffer and the second voice data stored in the second audio buffer to a second process at least partially simultaneously with the generation of the above recording data.

4. In any one of paragraphs 1 to 3, The above memory is, the electronic device, An electronic device that stores instructions for generating the summary data using an AI model based on the above call text data.

5. In Paragraph 4, The above memory is, the electronic device, An electronic device that stores instructions for transmitting a prompt to an AI model, including a request to generate the above call text data and a summary of the call content, and for obtaining the summary data as a response from the AI ​​model.

6. In Paragraph 5, The above memory is, the electronic device, An electronic device that stores instructions for requesting the AI ​​model to generate a summary by dividing the call text data into multiple chunks based on the number of tokens supported by the AI ​​model and generating multiple prompts containing each divided chunk.

7. In Paragraph 5, The above memory is, the electronic device, Information related to the first user and / or information related to the second user obtained based on the application or voice analysis of the recording data is transmitted to the AI ​​model, and An electronic device storing instructions for obtaining the summary data generated from the AI ​​model based on information related to the first user and / or information related to the second user.

8. In any one of Paragraphs 4 through 7, The above AI model is an electronic device that is an on-device AI model.

9. In any one of paragraphs 1 through 8, The above summary data is an electronic device comprising at least one of a title, keyword, summary, or image regarding the content of the above call.

10. In any one of paragraphs 1 through 9, The above memory is, the electronic device, The above call text data is distinguished into multiple text blocks and displayed in the first area of ​​the display, and Display the above summary data in the second area of ​​the above display, and An electronic device that stores instructions for displaying a user interface for controlling the playback of the above-mentioned recording data in a third area of ​​the display.

11. In Paragraph 10, The above memory is, the electronic device, An electronic device that stores instructions for providing a visual effect for a text block among the text blocks corresponding to the selected text, and / or moving to a playback section of the recording data corresponding to the selected text, when any of the texts of the summary data displayed on the display are selected based on user input.

12. In any one of paragraphs 1 through 11, The above memory is, the electronic device, When the above call is terminated, check the previous call and / or message associated with the terminated call, and An electronic device storing call text data of the above-mentioned terminated call, and call text data or summary data of the above-mentioned previous call, and / or instructions for generating summary data corresponding to the above-mentioned terminated call using the above-mentioned message.

13. In any one of paragraphs 1 through 12, The above memory is, the electronic device, Based on the above call text data, classify the content of the above call, and An electronic device that stores instructions for determining the format of the summary data based on the above classification.

14. In any one of paragraphs 1 through 13, The above memory is, the electronic device, An electronic device that stores instructions for generating summary data including a summary image based on the call text data and the call video data of the video call, when the call is a video call.

15. In a method performed by an electronic device, The operation of performing a call with an external device; During the above call, the operation of acquiring first voice data of a first user of the electronic device through the microphone of the electronic device; During the above call, based on data received from the external device, an operation of obtaining second voice data of a second user corresponding to the external device; The operation of generating call text data by mapping first speaker information corresponding to the first user to first text data converted from the first voice data, and mapping second speaker information corresponding to the second user of the external device to second text data converted from the second voice data; The operation of generating summary data of the call using the above call text data; and A method comprising the operation of providing the above summary data in relation to the above call.