Electronic device for summarizing call content and control method thereof
The electronic device uses neural network models to summarize call content and identify call types, offering efficient post-call summaries and alerts, improving user productivity.
Patent Information
- Application Number
- PCT/KR2025/009981
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-30
- Filing Date
- 2025-07-09
- Publication Date
- 2026-01-29
AI Technical Summary
Existing technologies lack efficient methods to summarize the content of phone calls using voice data, particularly in identifying call types and extracting relevant information post-call.
An electronic device equipped with neural network models processes voice data during and after a call to identify call types and generate summary information, including agenda, schedule, or inquiry details, and displays caution information when necessary.
The device effectively summarizes call content, enabling users to access key information quickly and accurately, enhancing productivity by providing detailed summaries and alerts.
Smart Images

Figure KR2025009981_29012026_PF_FP_ABST
Abstract
Description
Electronic device for summarizing call contents and method for controlling the same
[0001] The present disclosure relates to an electronic device for summarizing the contents of a call and a method for controlling the same, and more particularly, to an electronic device for obtaining summary information through a neural network model stored in the electronic device using voice data of the call and a method for controlling the same.
[0002] With the advancement of electronic technology, the use of electronic products that provide phone call functions and can record calls is increasing.
[0003] In particular, a technology has recently been developed to provide brief information by summarizing the contents of a phone call from voice data obtained by recording the call.
[0004] According to one embodiment of the present disclosure, an electronic device includes a memory storing at least one instruction and at least one neural network model, a communication device, a microphone, and at least one processor connected to the communication device, the microphone, and the memory to control the electronic device, wherein the at least one neural network model includes a model learned to output summary information based on voice data acquired by the electronic device, and the at least one processor acquires first summary information through the at least one neural network model based on first voice data acquired from a portion of a call through at least one of the communication device and the microphone, and, after the call is terminated, when the state of the electronic device is identified as at least one of a plurality of activation states capable of summarizing the content of the call corresponding to second voice data generated by recording the call until the end of the call, updates the first summary information through the at least one neural network model based on the second voice data to acquire second summary information.
[0005] The at least one processor may obtain the first summary information through the neural network model based on first voice data corresponding to voices continuously input for a set period of time after the call is initiated among the voice data.
[0006] The at least one processor may obtain the first summary information based on first voice data corresponding to a plurality of different partial periods during the period during which the call was performed after the call is terminated.
[0007] The above plurality of activation states may include at least one of a first activation state in which a call is made with a designated contact as the call party according to a user operation input and the call is terminated, and a second activation state in which the electronic device is charging and the user operation input is not detected for a set period of time after the call is terminated.
[0008] The at least one neural network model includes a classification model trained to identify the content of the call as one of a plurality of types based on the second voice data, and a plurality of type-specific summary models trained to output second summary information corresponding to each of the plurality of types, and the at least one processor can identify a type corresponding to the content of the call among the plurality of types through the classification model based on the acquired second voice data, and obtain the second summary information using a summary model corresponding to the identified type among the plurality of type-specific summary models.
[0009] The above-mentioned plurality of types includes at least one of a business call, a reservation call, an inquiry call, and a daily call, and the at least one processor, when the call type is identified as a business call, obtains the second summary information including at least one of a topic (Agenda), a to-do, and a schedule corresponding to the call content, when the call type is identified as a reservation call, obtains the second summary information including at least one of a reservation schedule, a number of people, and a reservation location, and when the call type is identified as an inquiry call, obtains the second summary information including an inquiry and an answer corresponding to the inquiry.
[0010] The at least one neural network model includes a model trained to output summary information based on voice data acquired by the electronic device and a time point at which a spoken voice of at least one of a plurality of participants of the call is received through the microphone, and the at least one processor records the call through the microphone until the end of the call, and records a time point at which a spoken voice of at least one of at least one of the plurality of participants of the call, excluding a user of the electronic device, is received to obtain time point data, and after the call is ended, when the state of the electronic device is identified as at least one of the plurality of active states, the second summary information can be obtained through the at least one neural network model based on the acquired time point data and second voice data corresponding to the spoken voice of the user acquired through the microphone.
[0011] The at least one processor may record the time at which a voice stronger than a predetermined intensity is received when the intensity of the speech of at least one of the at least one counterparty received through the microphone is identified as being stronger than a predetermined intensity.
[0012] The electronic device further includes a display, and the at least one processor controls the display to display caution information together with the second summary information when information included in at least one of the first summary information and the second summary information is identified as corresponding to at least one of a plurality of predetermined main information types, wherein the plurality of predetermined main information types include at least one of a time type and a place type, and the caution information may include information on whether the second voice data includes voice data corresponding to the plurality of counterpart's spoken voices.
[0013] The electronic device further includes a display, and the neural network model includes a model trained to convert input voice data into text and output it, and the at least one processor can control the display to input the second voice data into the neural network model and display the text together with the second summary information separately for each of the plurality of participants making the call.
[0014] According to one embodiment of the present disclosure, a method for controlling an electronic device includes the steps of: obtaining first summary information through at least one neural network model based on first voice data obtained from a portion of a call; and, after the call is terminated, when the state of the electronic device is identified as at least one of a plurality of activation states capable of summarizing the content of the call corresponding to second voice data generated by recording the call until the end of the call, updating the first summary information through the at least one neural network model based on the second voice data to obtain second summary information, wherein the at least one neural network model includes a model learned to output summary information based on the voice data obtained by the electronic device.
[0015] The step of obtaining the first summary information may include a step of obtaining the first summary information through the neural network model based on first voice data corresponding to voice continuously input for a set period of time after the call is initiated among the voice data.
[0016] The step of obtaining the first summary information may include a step of obtaining the first summary information based on first voice data corresponding to a plurality of different partial periods during the period during which the call was performed after the call is terminated.
[0017] The above plurality of activation states may include at least one of a first activation state in which a call is made with a designated contact as the call party according to a user operation input and the call is terminated, and a second activation state in which the electronic device is charging and the user operation input is not detected for a set period of time after the call is terminated.
[0018] The at least one neural network model may include a classification model trained to identify the content of the call as one of a plurality of types based on the second voice data and a plurality of type-specific summary models trained to output second summary information corresponding to each of the plurality of types, and the step of obtaining the second summary information may include a step of identifying a type corresponding to the content of the call among the plurality of types through a classification model based on the obtained second voice data, and a step of obtaining the second summary information using a summary model corresponding to the identified type among the plurality of type-specific summary models.
[0019] The above-mentioned multiple types may include at least one of a business call, a reservation call, an inquiry call, and a daily call, and the step of obtaining the second summary information may include a step of obtaining the second summary information including at least one of an agenda, a to-do, and a schedule corresponding to the call content when the call type is identified as a business call, a step of obtaining the second summary information including at least one of a reservation schedule, a number of people, and a reservation location when the call type is identified as a reservation call, and a step of obtaining the second summary information including an inquiry and an answer corresponding to the inquiry when the call type is identified as an inquiry call.
[0020] The at least one neural network model may include a model trained to output summary information based on voice data acquired by the electronic device and a time point at which a spoken voice of at least one of a plurality of participants of the call is received, and the step of acquiring the second summary information may include a step of recording the call until the end of the call and recording a time point at which a spoken voice of at least one of at least one of the plurality of participants of the call, excluding a user of the electronic device, is received to acquire time point data, and after the call is ended, if the state of the electronic device is identified as at least one of the plurality of activation states, a step of acquiring the second summary information through the at least one neural network model based on the acquired time point data and second voice data corresponding to the spoken voice of the user.
[0021] The step of acquiring the above point-in-time data may include a step of recording the point in time at which a voice having an intensity higher than a predetermined intensity is received when the intensity of the speech voice of at least one of the received at least one counterparty is identified as being higher than a predetermined intensity.
[0022] The method further includes a step of displaying caution information together with the second summary information when information included in at least one of the first summary information and the second summary information is identified as corresponding to at least one of a plurality of predetermined main information types, wherein the plurality of predetermined main information types include at least one of a time type and a place type, and the caution information may include information on whether the second voice data includes voice data corresponding to the plurality of counterpart's spoken voices.
[0023] In accordance with one embodiment of the present disclosure, a non-transitory computer-readable recording medium storing computer instructions that, when executed by a processor of an electronic device, cause the electronic device to perform an operation, the operation includes: obtaining first summary information through at least one neural network model based on first voice data obtained from a portion of a call; and, after the call is terminated, when the state of the electronic device is identified as at least one of a plurality of activation states capable of summarizing the content of the call corresponding to second voice data generated by recording the call until the end of the call, updating the first summary information through the at least one neural network model based on the second voice data to obtain second summary information, wherein the at least one neural network model may include a model learned to output summary information based on the voice data obtained by the electronic device.
[0024] FIG. 1 is a diagram illustrating the operation of an electronic device according to one or more embodiments of the present disclosure.
[0025] FIG. 2 is a block diagram illustrating a configuration of an electronic device according to one or more embodiments of the present disclosure.
[0026] FIG. 3 is a detailed block diagram illustrating a detailed configuration of an electronic device according to one or more embodiments of the present disclosure.
[0027] FIG. 4 is a diagram illustrating the operation of an electronic device according to one or more embodiments of the present disclosure.
[0028] FIG. 5 is a diagram illustrating an activation state according to one or more embodiments of the present disclosure.
[0029] FIG. 6 is a diagram illustrating an activation state according to one or more embodiments of the present disclosure.
[0030] FIG. 7 is a diagram illustrating an activation state according to one or more embodiments of the present disclosure.
[0031] FIG. 8 is a diagram illustrating an activation state according to one or more embodiments of the present disclosure.
[0032] FIG. 9 is a diagram illustrating second summary information according to one or more embodiments of the present disclosure.
[0033] FIG. 10 is a diagram illustrating a subsequent screen after obtaining second summary information according to one or more embodiments of the present disclosure.
[0034] FIG. 11 is a diagram illustrating a call type according to one or more embodiments of the present disclosure.
[0035] FIG. 12 is a diagram illustrating a call type according to one or more embodiments of the present disclosure.
[0036] FIG. 13 is a flowchart illustrating an operation of obtaining summary information based on the time of utterance according to one or more embodiments of the present disclosure.
[0037] FIG. 14 is a drawing for explaining whether the other party's speech is recorded according to one or more embodiments of the present disclosure.
[0038] FIG. 15 is a diagram illustrating a conversation record according to one or more embodiments of the present disclosure.
[0039] FIG. 16 is a drawing for explaining caution information according to one or more embodiments of the present disclosure.
[0040] FIG. 17 is a diagram illustrating a conversation record of a video conference according to one or more embodiments of the present disclosure.
[0041] FIG. 18 is a flowchart illustrating a method for controlling an electronic device according to one or more embodiments of the present disclosure.
[0042] The present embodiments may be modified and have various embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope to specific embodiments, but should be understood to encompass various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In connection with the description of the drawings, similar reference numerals may be used for similar components.
[0043] In describing the present disclosure, if it is determined that a specific description of a related known function or configuration may unnecessarily obscure the gist of the present disclosure, a detailed description thereof will be omitted.
[0044] Additionally, the following embodiments may be modified in various other forms, and the scope of the technical concepts of the present disclosure is not limited to the following embodiments. Rather, these embodiments are provided to further faithfully and completely convey the technical concepts of the present disclosure to those skilled in the art.
[0045] The terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the scope of the rights. Singular expressions include plural expressions unless the context clearly dictates otherwise.
[0046] In this disclosure, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a corresponding feature (e.g., a component such as a number, function, operation, or part), and do not exclude the presence of additional features.
[0047] In this disclosure, expressions such as “A or B,” “at least one of A and / or B,” or “one or more of A or / and B” can include all possible combinations of the listed items. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” can all refer to (1) including at least one A, (2) including at least one B, or (3) including both at least one A and at least one B.
[0048] The expressions “first,” “second,” “first,” or “second,” etc., used in this disclosure can describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.
[0049] When it is said that a component (e.g., a first component) is “(operatively or communicatively) coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that said component may be directly coupled to said other component, or may be coupled via another component (e.g., a third component).
[0050] On the other hand, when it is said that a component (e.g., a first component) is "directly connected" or "directly connected" to another component (e.g., a second component), it can be understood that no other component (e.g., a third component) exists between said component and said other component.
[0051] The expression "configured to" as used in the present disclosure may be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" may not necessarily mean only "specifically designed to" in terms of hardware.
[0052] Instead, in some contexts, the phrase "a device configured to" may mean that the device, in conjunction with other devices or components, is "capable of" performing A, B, and C. For example, the phrase "a processor configured (or set) to perform A, B, and C" may refer to a dedicated processor (e.g., an embedded processor) for performing those operations, or a general-purpose processor (e.g., a CPU or application processor) that can perform those operations by executing one or more software programs stored in a memory device.
[0053] In the embodiments, a 'module' or 'part' performs at least one function or operation, and may be implemented as hardware or software, or as a combination of hardware and software. Furthermore, a plurality of 'modules' or 'parts' may be integrated into at least one module and implemented as at least one processor, except for a 'module' or 'part' that needs to be implemented as a specific hardware.
[0054] Meanwhile, the various elements and areas in the drawings are schematically drawn. Therefore, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.
[0055] Hereinafter, with reference to the attached drawings, embodiments according to the present disclosure will be described in detail so that a person having ordinary knowledge in the technical field to which the present disclosure pertains can easily implement the present disclosure.
[0056] FIG. 1 is a diagram illustrating the operation of an electronic device according to one or more embodiments of the present disclosure.
[0057] Referring to FIG. 1, an electronic device (100) and a user (10) of the electronic device (100) are illustrated. The electronic device (100) can display a first screen (100-1) and a second screen (100-2).
[0058] Here, the electronic device (100) is a smartphone, a tablet PC (personal computer), a desktop PC, a laptop PC, a PC, a set-top box, an OTT service (over-the-top media service) server, a console (video game console), a Blu-ray player, a DVD (Digital Video Disc or Digital Versatile Disc) player, a home automation control panel, a security control panel, a media box (e.g., Samsung HomeSync) TM , Apple TV TM , or Google TV TM), game consoles (e.g. Xbox TM , PlayStation TM ) can be implemented as at least one of the following, but is not limited thereto.
[0059] A user (10) can make a call through an electronic device (100). The user (10) can exchange conversations while making a call with another user of an electronic device other than the electronic device (100). Here, the electronic device (100) can perform a call function while exchanging voice signals with the user (10) and another user.
[0060] For example, the electronic device (100) may be implemented as a device that provides the above-described calling function, such as a smartphone. However, the present invention is not limited thereto, and the electronic device (100) may also be implemented as a tablet PC, desktop PC, laptop PC, etc. equipped with a calling function.
[0061] Here, the call may correspond to a voice call through the users' speech, as illustrated in FIG. 1. However, the present invention is not limited thereto, and the call may correspond to a video call in which the user can exchange voices while viewing the other party's appearance through the screen of the electronic device (100). In this case, the electronic device (100) may exchange video signals with other electronic devices, along with voice signals from the user (10) and other users.
[0062] At this time, when the user (10) ends the call, the electronic device (100) can display information related to the ended call. For example, the electronic device (100) can display a first screen (100-1) showing recent call records including records of the ended call. For example, the electronic device (100) can display a second screen (100-2) which is a summary screen showing call content related to the ended call.
[0063] Here, termination may mean a state in which the electronic device (100) no longer transmits or receives voice signals from the other party and the user (10).
[0064] Here, the user (10) can end the call by directly manipulating the electronic device (100) to end the call. Alternatively, the other party to the call (10) can end the call by manipulating the other party's electronic device to end the call. However, this is not limited to this, and the call can also be ended if the communication environment for conducting the call is unstable and the transmission and reception of voice signals are blocked.
[0065] On the first screen (100-1), multiple recent call records are sorted in chronological order by date and time. Here, the first screen (100-1) may include information about the other party of each call, as well as the date and time. Here, the date and time may refer to the date and time of the call.
[0066] Although the first screen (100-1) is illustrated as containing only records of calls made by the user, this is not limited to this, and the first screen (100-1) may also contain records of received calls. In this case, the first screen (100-1) may include the date and time the call was received.
[0067] Here, the other party's information may correspond to the name and phone number of the other party with whom the user (10) spoke. If the user (10) has previously stored the other party's phone number along with the other party's name in the electronic device (100), the electronic device (100) may display the other party's name instead of the other party's phone number. Here, the other party's name does not necessarily mean the other party's actual name (as recorded in the resident registration), but may also mean a nickname or alias (e.g., 'Yujeong', 'Mom') set by the user (10) instead of the user's name.
[0068] Here, the first screen (100-1) can briefly display the call details for each of the multiple call records. For example, in the case of a recently ended call with "Jinu," the electronic device (100) can display the caller "Jinu" along with the call details, "#wedding #dress shop." The call details can be summarized as keywords, representing the conversation with the caller.
[0069] Here, the electronic device (100) can record the call audio while the user (10) is on a call with the other party, and then obtain summary information about the call content. At this time, the electronic device (100) can utilize a neural network model to obtain summary information about the call audio. This will be described in detail below in FIG. 2.
[0070] The user (10) can select the first call record (recently ended call record) section on the first screen (100-1). Here, the user can touch the section where the first call record is displayed via the electronic device (100). However, this is not limited thereto.
[0071] The electronic device (100) can display a second screen (100-2). Here, the second screen (100-2) can include a detailed summary of the content of a recently ended call (a call with 'Jinu').
[0072] For example, the electronic device (100) can generate summary information based on the recorded voice until the user (10) ends the call. Here, the electronic device (100) can utilize a neural network model to obtain summary information about the call voice. This will be described in detail below in FIG. 2.
[0073] Meanwhile, the second screen (100-2) may include generated summary information. Specifically, the summary information may include call topics, keywords, and key information.
[0074] Here, the second screen (100-2) may include a call topic, such as "Discussing an appointment time for a dress shop visit." The second screen (100-2) may include keywords related to the call, such as "#dress shop," "#appointment time," and "#dermatologist consultation." Furthermore, the summary screen (100-2) may include key content, such as "Greetings," "Preparing for a dress shop tour," "Meeting time and place," and "End of call." The key content may be information that briefly summarizes the call's content in 2-3 words, chronologically.
[0075] Meanwhile, the second screen (100-2) may include a conversation record generated by converting the call voice into text. Specifically, the electronic device (100) may perform STT (Speech-to-Text) on the call voice, thereby performing voice recognition on the user's voice. The electronic device (100) may then convert the recognized voice into text and display it as a conversation record. This will be described in detail later in FIGS. 4 and 15.
[0076] Accordingly, the electronic device (100) can provide the user (10) with not only a call function, but also a function to record calls and provide a summary of the recorded calls. This allows the user (10) to check recent call records at a glance using simple keywords after a call has ended. Furthermore, the user (10) can select one of the call records and receive more detailed summary information related to the selected call.
[0077] FIG. 2 is a block diagram illustrating a configuration of an electronic device according to one or more embodiments of the present disclosure.
[0078] According to FIG. 2, the electronic device (100) may include a memory (110), a communication device (120), a microphone (130), and at least one processor (140).
[0079] The memory (110) is electrically connected to at least one processor (140) and can store data required for various embodiments of the present disclosure. For example, the memory (110) may be implemented as an internal memory such as a ROM (Read-Only Memory) (e.g., an electrically erasable programmable read-only memory (EEPROM)), a RAM (Random Access Memory)) included in the processor (140), or may be implemented as a separate memory from at least one processor (140).
[0080] The memory (110) may be implemented in the form of memory embedded in the electronic device (100) or may be implemented in the form of memory that can be attached or detached from the electronic device (100) depending on the purpose of data storage. For example, data for driving the electronic device (100) may be stored in a memory embedded in the electronic device (100), and data for expanding the functions of the electronic device (100) may be stored in a memory that can be attached or detached from the electronic device (100). When implemented as a memory embedded in an electronic device (100), the memory (110) may be at least one of volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM), or synchronous dynamic RAM (SDRAM)), non-volatile memory (e.g., one time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard drive, or solid state drive (SSD)).
[0081] Meanwhile, in the illustrated example, the electronic device (100) is depicted as being composed of one memory, but when referring to volatile memory and non-volatile memory separately, the electronic device (100) may be referred to as including multiple memories.
[0082] The memory (110) according to one or more embodiments may store at least one instruction. Here, the at least one instruction may correspond to at least one command for the electronic device (100) to obtain summary information. The memory (110) may also store information necessary for the operation of the electronic device (100).
[0083] A memory (110) according to one or more embodiments may store at least one neural network model.
[0084] Here, the neural network model is a computer system or software module that implements human-level intelligence, and has the characteristics of a machine learning and making judgments on its own, and its recognition rate improving with use.
[0085] A neural network model is composed of machine learning (deep learning) technology that uses an algorithm that classifies / learns the characteristics of input data on its own, and element technologies that use machine learning algorithms to simulate the cognitive and judgment functions of the human brain.
[0086] Here, the neural network model may also be referred to as a learning model, an AI (Artificial Intelligence) model, or a deep learning model.
[0087] The element technologies may include at least one of, for example, a linguistic understanding technology that recognizes human language / characters, a visual understanding technology that recognizes objects as if they were human vision, an inference / prediction technology that judges information and logically infers and predicts, and a knowledge representation technology that processes human experience information into knowledge data.
[0088] For example, a neural network model may correspond to a model trained to convert input speech data into text and output it. Here, the neural network model may be implemented as a speech recognition model. In this case, the speech recognition model may correspond to a model trained to output text corresponding to the input speech through Automatic Speech Recognition (ASR).
[0089] Here, ASR is a technology that automatically converts speech into text. It can be defined as a technology that analyzes speech data and accurately transcribes human speech into text. In addition to ASR, a speech recognition model can be implemented as a model trained to output text corresponding to speech based on various speech recognition technologies.
[0090] Meanwhile, the neural network model may correspond to a model trained to output summary information based on voice data acquired by the electronic device (100).
[0091] Here, a neural network model can be trained to obtain summary information through the Natural Language Processing (NLP) process. NLP can be defined as a technology that enables computers to understand and process human language, and can be implemented through a large language model (LLM). This can also be applied to the various types of neural network models described below.
[0092] Here, voice data may correspond to voice data obtained by the electronic device (100) by recording a call. Here, recording may refer to the process of recording voices exchanged during a phone call and storing them in the form of a digital file.
[0093] For example, voice data may refer to data converted into digital format by recording voice during a call. Alternatively, voice data may refer to data converted from voice rather than recorded voice.
[0094] For example, voice data may correspond to data converted into digital format from voice input through a microphone or communication device, etc. Here, voice input through a microphone, etc. may correspond to voice input through a microphone, etc. in real time.
[0095] Here, the voice input in real time may mean voice that is immediately collected and processed through a microphone or the like during a call. In this case, the electronic device (100) can receive voice input in real time and immediately convert it into voice data.
[0096] The voice data converted here can be used to generate a summary of the call content. This will be explained in detail in the following section.
[0097] Here, voice data may include information about the user and the other party's voice. This will be described in more detail in the sections below.
[0098] Meanwhile, the neural network model may correspond to a classification model trained to identify the content of a call as one of multiple types based on voice data. Here, the classification model can identify the type of call based on the estimated content of the call based on the voice data.
[0099] The types of calls here may include business calls, reservation calls, etc. These are described in detail in Figures 11 and 12.
[0100] Meanwhile, the neural network model may correspond to multiple type-specific summary models trained to output summary information corresponding to each of the multiple types. Here, the type-specific summary model may correspond to a model trained to output appropriate summary information based on the call type.
[0101] Meanwhile, a neural network model may be trained to output summary information based on voice data and the timing of the participants' speech. In other words, the neural network model can be trained to output summary information based not only on the recorded voice data of the call, but also on the timing of the participants' speech (the other parties).
[0102] For example, an electronic device (100) can record only the user's speech and obtain information about the time of the other party's speech. The electronic device (100) can input voice data corresponding to the recorded user's speech and information about the time of speech into a neural network model to obtain summary information. This will be described in detail later in FIG. 13.
[0103] Meanwhile, neural network models can be implemented as E2E multimodal AI. Here, E2E can refer to a method in which a single model or system directly handles the entire process, from data input to final output.
[0104] Here, multimodal AI can refer to an AI model capable of processing various types of data simultaneously. In other words, in the case of end-to-end multimodal AI, it could refer to a model trained to simultaneously input data such as text, images, voice, and video and output results.
[0105] The neural network model is not limited to the examples described above, and the neural network model can be implemented as a variety of models trained to perform the operations necessary for the electronic device (100) to obtain summary information.
[0106] The communication device (120) is a configuration that performs communication with various types of external devices according to various types of communication methods. The communication device (120) may include a Wi-Fi module, a Bluetooth module, an infrared communication module, a wireless communication module, etc. Here, each communication module may be implemented in the form of at least one hardware chip.
[0107] Wi-Fi and Bluetooth modules can communicate via Wi-Fi and Bluetooth, respectively. When using a Wi-Fi or Bluetooth module, connection information, such as the SSID and session key, is first transmitted and received. This information is then used to establish a connection before various other information can be transmitted and received.
[0108] Infrared communication modules perform communication based on infrared communication (IrDA, infrared Data Association) technology, which transmits data wirelessly over short distances using infrared light, which lies between visible light and millimeter waves.
[0109] In addition to the above-described communication method, the wireless communication module may include at least one communication chip that performs communication according to various wireless communication standards such as zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), LTE-A (LTE Advanced), 4G (4th Generation), 5G (5th Generation), etc.
[0110] In addition, the communication device (120) may include at least one of wired communication modules that perform communication using a LAN (Local Area Network) module, an Ethernet module, a pair cable, a coaxial cable, an optical fiber cable, or a UWB (Ultra Wide-Band) module. Such a communication device (120) may also be referred to as a transceiver.
[0111] According to one or more embodiments, the electronic device (100) may receive voice information (or voice data). Specifically, the electronic device (100) may receive a digital signal corresponding to the voice of the other party from an external electronic device, etc., via a communication device (120). Here, the external electronic device may correspond to an electronic device (such as a smartphone, tablet PC, or PC) used by the other party to conduct a call.
[0112] The electronic device (100) can decode a digital signal received through the communication device (120). Thereafter, the electronic device (100) can record the voice of the other party during a voice call by recording the decoded signal.
[0113] A microphone (130) provided in an electronic device (100) may receive a user's voice and transmit the received user's voice to at least one processor (140). Subsequently, the at least one processor (140) may input the received user's voice into a voice recognition model to perform voice recognition. For example, the at least one processor (140) may perform voice recognition on the user's voice in order to perform STT (Speech to Text) on the user's voice. Here, the voice recognition may be performed by the above-described ASR.
[0114] According to one or more embodiments, the electronic device (100) can receive a user's voice during a call via a microphone (130). Specifically, the electronic device (100) can convert the voice input via the microphone (130) into a digital signal. The electronic device (100) can record the user's voice during a call by recording the digital signal.
[0115] However, the present invention is not limited thereto, and the electronic device (100) can obtain summary information using a converted digital signal (voice data). That is, the electronic device (100) can convert voice input in real time during a call into voice data and obtain summary information using the converted voice data. Since voice input in real time, etc. have been previously described, a duplicate description will be omitted.
[0116] Here, the user's voice may correspond to a voice spoken by the user through vocalization. However, this is not limited to this, and the user's voice may correspond to a voice generated in the user's surroundings and input through a microphone (130).
[0117] At least one processor (140) can perform overall control operations of the electronic device (100).
[0118] At least one processor (140) may be implemented as a digital signal processor (DSP), a microprocessor, a time controller (TCON) for processing a digital signal. However, the present invention is not limited thereto, and may include one or more of a central processing unit (CPU), a micro controller unit (MCU), a micro processing unit (MPU), a controller, an application processor (AP), a graphics-processing unit (GPU), a communication processor (CP), an ARM processor, or may be defined by the relevant term. In addition, at least one processor (140) may be implemented as a system on chip (SoC) having a built-in processing algorithm, a large scale integration (LSI), or may be implemented in the form of a field programmable gate array (FPGA). In addition, at least one processor (140) may perform various functions by executing computer executable instructions stored in a memory (110). Meanwhile, in FIG. 2, an electronic device (100) Although illustrated as containing only one processor, the implementation may include multiple processors (e.g., CPU + GPU, CPU + DSP).
[0119] According to one or more embodiments, at least one processor (140) may obtain first summary information through at least one neural network model based on first voice data obtained from a portion of a call through at least one of a communication device (120) and a microphone (130).
[0120] Here, a portion of a call may refer to audio data for a portion of the entire call duration. The first audio data obtained from a portion of a call may correspond to audio data obtained from the call for a portion of the entire call duration.
[0121] Here, at least one processor (140) can receive voice input for a portion of a call in real time and use the input voice to obtain first voice data. Here, the voice data may correspond to data in which the input voice is converted into a digital format in real time.
[0122] Additionally, at least one processor (140) may record audio for a portion of a call. For example, at least one processor (140) may convert audio for a portion of a call into a digital format and store it as first audio data in the memory (110). However, this is not a limitation.
[0123] For example, at least one processor (140) may obtain first summary information through a neural network model based on first voice data corresponding to voices continuously input for a predetermined period of time after a call is initiated among voice data. At this time, at least one processor (140) may continuously receive voices from a call for a predetermined period of time starting immediately after the call is initiated. However, the present invention is not limited thereto, and calls may be continuously received for a predetermined period of time starting after a predetermined period of time has elapsed since the call is initiated.
[0124] Meanwhile, the user can configure the electronic device (100) to initiate a call summary operation based on a specific greeting. For example, the user can set a word commonly spoken at the start of a call, such as "Hello," "Good morning," or "Hello," as the summary initiation word. If the electronic device (100) recognizes the above-mentioned summary initiation word during a call, the electronic device (100) can record the call for a set period of time.
[0125] The time set here may be less than the total call time. In this case, the set time may correspond to a user-defined time (hours, minutes, seconds). Alternatively, the set time may correspond to a percentage of the total call time, as determined by the user.
[0126] Here, the neural network model may correspond to a model trained to summarize a call based on voice data, as described above. However, the neural network model may correspond to a model trained to summarize a call based on voice data corresponding to a portion of the call. Here, the neural network model may be implemented as a model that operates in the E2E multimodal manner described above.
[0127] For example, a neural network model may have as input data voice data obtained from a portion of a call, text converted from the voice data, and information related to the call (such as call records with the same caller, text messages, and call times). The neural network model may correspond to a model trained to output first summary information using such data as input data.
[0128] Here, the first summary information may be information that briefly summarizes the content of the call using keywords. The first summary information is distinct from the second summary information described below, and may be simpler than the second summary information. This will be described in detail later in Figure 4.
[0129] For example, at least one processor (140) may obtain first summary information based on first voice data corresponding to different sub-periods during the call period after the call is terminated. However, this is not limited to the first processor (140), and at least one processor (140) may obtain first summary information based on first voice data even before the call is terminated.
[0130] The period during which the call was made here may correspond to the total call time described above.
[0131] Here, the different sub-periods may refer to multiple time intervals within the total call time. Each of these time intervals may correspond to non-contiguous, non-connected segments. In other words, the sum of the times corresponding to each of these time intervals may be less than the total call time.
[0132] Here, each of the multiple time intervals can correspond to a user-defined interval. For example, a user may set the system to receive voice input for 10 seconds every two minutes. For example, a user may set the system to receive voice input for one minute when a specific word is recognized. However, this is not a limitation.
[0133] According to one or more embodiments, at least one processor (140) may, after a call is terminated, identify that the state of the electronic device is at least one of a plurality of active states capable of summarizing the content of the call corresponding to second voice data generated by recording the call until the end of the call, update the first summary information through at least one neural network model based on the second voice data to obtain second summary information.
[0134] Here, the second voice data may correspond to voice data obtained from the call during the entire call duration. In other words, the second voice data may correspond to data containing the voice of the entire call.
[0135] Here, the operation of updating the first summary information by the electronic device (100) may mean an operation of generating second summary information that includes more information than the first summary information that has been previously acquired.
[0136] For example, the second summary information includes the first summary information and may include more diverse types of information (e.g., keywords, topics, main content) than the types of the first summary information (e.g., keywords).
[0137] Meanwhile, the second summary information may include information that modifies the first summary information. For example, since the first summary information includes information generated based on a portion of the call audio, it may contain inaccurate information (content different from the call content). In this case, the incorrect first summary information can be corrected using the second summary information, which is generated by summarizing the entire call audio.
[0138] Here, the active state may mean a state in which the electronic device (100) can summarize the content of a call corresponding to the second voice data. Here, the content of a call corresponding to the second voice data may mean the content of the entire call.
[0139] Specifically, the active state may mean a state in which the entire call content corresponding to the second voice data can be summarized according to the battery status of the electronic device (100), user settings, etc.
[0140] Here, the battery status may mean the amount of power available to the electronic device (100).
[0141] For example, a state in which the entire call content can be summarized based on the battery status may mean a state in which there is sufficient power to summarize the entire call content based on the second voice data capacity.
[0142] Here, "user settings" may refer to settings that allow users to manually summarize the entire call. Furthermore, "user settings" may refer to settings that allow users to summarize the entire call when making a call with a specific party.
[0143] At least one processor (140) can summarize the content of the entire call when the electronic device (100) is identified as being in at least one of a plurality of active states.
[0144] For example, the plurality of activation states may include at least one of a first activation state and a second activation state. However, this is not limited to the first activation state, and the plurality of activation states may also include a third activation state. The third activation state may refer to a state in which a user operation input for obtaining second summary information has been received.
[0145] Here, the first active state may refer to a state in which a call is made to a designated contact based on user input and the call is terminated. Here, the second active state may refer to a state in which the electronic device is charging and no user input is detected for a specified period of time after the call is terminated.
[0146] The states in which the electronic device (100) summarizes the entire call content are not limited to the first activation state, the second activation state, and the third activation state. The first activation state, the second activation state, and the third activation state will be described in detail in FIG. 5 and below.
[0147] Meanwhile, at least one processor (140) can update the first summary information based on the second voice data when the electronic device (100) is in an activated state.
[0148] Here, at least one processor (140) can obtain second summary information including first summary information.
[0149] Here, the second summary information may be a summary of the entire call content. The second summary information is distinct from the first summary information described above and may include more detailed information than the first summary information.
[0150] Here, the second summary information, like the first summary information, may include multiple keywords. Furthermore, the second summary information may further include a topic and key information, as illustrated in Figure 1.
[0151] For example, at least one processor (140) can identify a type corresponding to the content of a call among a plurality of types through a classification model based on the acquired second voice data.
[0152] Thereafter, at least one processor (140) can obtain second summary information using a summary model corresponding to the identified type among a plurality of type-specific summary models.
[0153] For example, the plurality of types may include at least one of business calls, reservation calls, inquiry calls, and routine calls.
[0154] For example, if the call type is identified as a business call, second summary information can be obtained that includes at least one of the topic (Agenda), to-do, and schedule corresponding to the call content.
[0155] For example, if the call type is identified as a reservation call, second summary information including at least one of a reservation schedule, number of people, and reservation location can be obtained.
[0156] For example, if the call type is identified as an inquiry call, secondary summary information including the inquiry and the corresponding response to the inquiry can be obtained.
[0157] The multiple currency types are not limited to the examples described above, and may include various types depending on the content of the call. The multiple currency types and the summary information corresponding to each type are described in detail in Figures 11 and 12.
[0158] According to one or more embodiments, at least one processor (140) may record a call through the microphone (130) until the end of the call.
[0159] Here, at least one processor (140) can record a call through the microphone (130) until the call ends. For example, at least one processor (140) can record the user's voice until the call ends. In this case, at least one processor (140) may not record the other party's voice through the communication device.
[0160] According to one or more embodiments, at least one processor (140) may obtain point-in-time data by recording a point in time at which at least one spoken voice is received from at least one of the plurality of participants in the call, excluding the user of the electronic device.
[0161] Here, multiple participants may mean multiple users conducting a call together with the user.
[0162] For example, at least one processor (140) here can receive the spoken voice of at least one of the counterparts through the microphone (130). That is, at least one processor (140) can receive a signal corresponding to the spoken voice of the counterpart through the microphone (130) without receiving it through the communication device.
[0163] At this time, at least one processor (140) may directly receive the user's spoken voice output from the electronic device (100) through the microphone (130). Alternatively, at least one processor (140) may receive the user's spoken voice output through the speaker described below through the microphone (130). This will be described in detail later in FIG. 3.
[0164] Meanwhile, at least one processor (140) may utilize a separate voice output device that receives and outputs the other party's spoken voice. That is, when a separate voice output device located outside the electronic device (100) receives a voice signal for the other party's spoken voice and outputs the user's spoken voice, the electronic device (100) can receive the outputted voice through the microphone (130).
[0165] Here, the voice output device can be implemented as an external speaker capable of communicating with an external device, an AI speaker, etc.
[0166] Meanwhile, at least one processor (140) can record the time at which the other party's spoken voice is received to obtain time data. The time data may include information about the elapsed time since the call began. Alternatively, the time data may include information about the time at which the spoken voice was received. However, this is not a limitation.
[0167] For example, at least one processor (140) may record the time at which a voice of at least one of the parties received through the microphone (130) is received when the intensity of the voice is identified as being greater than a predetermined intensity.
[0168] Here, the intensity of the spoken voice can refer to the loudness of the received spoken voice. Here, the sound intensity can be expressed in units of dB (Decibel). However, this is not limited to this, and the sound intensity (loudness) can be expressed in various ways, such as in phon, son, or sound pressure.
[0169] The intensity set here may correspond to an intensity set by the user. Alternatively, the intensity may correspond to an intensity automatically set by the electronic device (100) based on the surrounding environment (e.g., noise environment) of the electronic device (100). Here, the electronic device (100) may set the set intensity higher as the intensity of the surrounding noise increases.
[0170] For example, at least one processor (140) may record a time when a voice above a predetermined intensity is received when the other party's voice, rather than the user's voice, is identified as being above a predetermined intensity while recording a call through a microphone (130).
[0171] Thereafter, at least one processor (140) can record the point in time when a voice of a predetermined intensity or higher is received to obtain point-in-time data.
[0172] According to one or more embodiments, at least one processor (140) may obtain second summary information through at least one neural network model based on the acquired point-in-time data and second voice data corresponding to the user's spoken voice obtained through the microphone (130) when the state of the electronic device is identified as at least one of a plurality of active states after the call is terminated.
[0173] For example, at least one processor (140) can obtain second summary information based on the point-of-view data as well as second voice data including the user's spoken voice.
[0174] Meanwhile, at least one processor (140) can predict the time of the other party's speech based on the time at which the spoken voice was received. Here, at least one processor (140) can predict the time of speech by calculating a specific time backward from the time at which the spoken voice was received.
[0175] Here, the specific time may refer to the time it takes for the other party's spoken voice to be transmitted to the electronic device (100) as a voice signal. In other words, the specific time may refer to the time it takes for the other party's electronic device to receive the other party's spoken voice, and for a signal corresponding to the voice to be transmitted to the electronic device (100) and received by the microphone (130).
[0176] The time required here may also mean the estimated time required until the signal is received by the microphone (130) of the electronic device (100) through the above process.
[0177] Thereafter, at least one processor (140) can obtain second summary information based on the predicted ignition timing. A detailed description thereof will be provided later in FIG. 13.
[0178] Although FIG. 2 illustrates the electronic device (100) as including only basic components (i.e., memory, communication device, microphone, and processor), the electronic device (100) may further include various components in addition to the aforementioned components. Such examples are described below with reference to FIG. 3.
[0179] FIG. 3 is a detailed block diagram illustrating a detailed configuration of an electronic device according to one or more embodiments of the present disclosure.
[0180] According to FIG. 3, the electronic device (100) may include a memory (110), a communication device (120), a microphone (130), at least one processor (140), a display (150), and a speaker (160).
[0181] The memory (110), communication device (120), microphone (130) and at least one processor (140) have been previously described in FIG. 2, and thus, a duplicate description thereof will be omitted.
[0182] The electronic device (100) displays video data through a display (150). The display (150) may be implemented as a TV, but is not limited thereto, and may be applied to any device with a display function, such as a video wall, a large format display (LFD), a digital signage, a digital information display (DID), a projector display, etc.
[0183] In addition, the display (150) can be implemented as various types of displays such as a liquid crystal display (LCD), an organic light-emitting diode (OLED), a liquid crystal on silicon (LCoS), a digital light processing (DLP), a quantum dot (QD) display panel, a quantum dot light-emitting diodes (QLED), a micro light-emitting diodes (μLED), a mini LED, etc. / Meanwhile, the display (150) can also be implemented as a touch screen combined with a touch sensor, a flexible display, a rollable display, a 3D display, a display in which a plurality of display modules are physically connected, etc.
[0184] According to one or more embodiments, at least one processor (140) may control the display (150) to display caution information together with the second summary information when information included in at least one of the first summary information and the second summary information is identified as corresponding to at least one of a predetermined plurality of primary information types.
[0185] Meanwhile, if at least one processor (140) records only the user's voice through the microphone (130), at least one processor (140) may not be able to obtain voice data for the other party's voice.
[0186] In this case, since at least one processor (140) obtains the second summary information without audio data regarding the other party's voice, the second summary information may be inaccurate. That is, the second summary information obtained based on the user's voice may contain errors compared to the second summary information based on the voices of both the user and the other party. Here, an error may mean that the second summary information is not factual.
[0187] Here, "critical information" can refer to information that, if inaccurate, could cause serious problems for the user. Specifically, "critical information" can refer to information requiring relatively high accuracy, such as appointment times and locations. Accuracy here can refer to how close the information (summary information) is to the truth.
[0188] Here, the key information type can be a type that distinguishes key information by type based on its content. For example, key information types can correspond to appointment time types and appointment information types, but this is not limited to these.
[0189] Here, the primary information type can be set based on user interaction. For example, a user may set information about appointment times as the primary information type. However, this is not a limitation.
[0190] For example, a given plurality of primary information types may include at least one of a time type and a location type.
[0191] Here, the information included in at least one of the first summary information and the second summary information may refer to at least one word included in at least one of the first summary information and the second summary information. For example, the summary information (the first summary information or the second summary information) may include multiple words such as "meeting reservation time 8:30."
[0192] Here, since the 'meeting reservation time' corresponds to information corresponding to the main information type such as 'reservation time', at least one processor (140) can identify that the first summary information or the second summary information includes information corresponding to the main information type.
[0193] However, this is not limited to this, and multiple information types may include other types in addition to time and location types. For example, multiple information types may include a call type for "key decision items." Here, the call type for key decision items may be referred to as a "decision item type."
[0194] Here, the decision type may correspond to a call type, rather than the time and location types described above. For example, if the first and second summary information contain words related to "business meetings" (e.g., project, report, presentation, materials, etc.), the call type may be identified as "key decision." However, this is not limited to this.
[0195] Accordingly, at least one processor (140) can control the display (150) to display the attention information together with the second summary information.
[0196] Here, the caution information may mean information to alert the user that the second summary information may not be accurate.
[0197] For example, the attention information may include information regarding whether the second voice data includes voice data corresponding to multiple counterparts' spoken voices. This will be described in detail later in FIG. 16.
[0198] According to one or more embodiments, at least one processor (140) may control the display (150) to display text obtained by inputting second voice data into a neural network model, together with second summary information, separately for each of the plurality of participants making the call.
[0199] Here, the neural network model can be implemented using the aforementioned speech recognition model. The speech recognition model can be trained to output text corresponding to the input speech data. The text can correspond to a sentence obtained by converting the speech of at least one of the user and the other party.
[0200] For example, the text may include a sentence corresponding to a spoken voice and information about the participant who uttered the spoken voice. The multiple participants may include the user. The multiple participants may refer to anyone who participated in the call, including the user. Furthermore, the multiple participants may refer to anyone who spoke at least once during the call.
[0201] At least one processor (140) can obtain text from the second voice data and display the text along with the second summary information. The text may be displayed in the form of a chat conversation between multiple participants. This will be described in detail later in FIGS. 4 and 17.
[0202] Meanwhile, the electronic device (100) can output the rendered sound signal through the speaker (160). In this case, the speaker (160) can be implemented with at least one speaker (160) unit. For example, the speaker (160) can include a plurality of speakers (160) for multi-channel reproduction. For example, the speaker (160) can include a plurality of speakers (160) responsible for each channel that is mixed and output. In some cases, the speaker (160) responsible for at least one channel can also be implemented with a speaker (160) array that includes a plurality of speaker (160) units for reproducing different frequency bands.
[0203] For example, at least one processor (140) can output an acoustic signal (voice signal) received through a communication device (120) through a speaker (160).
[0204] For example, in a case where at least one processor (140) records only the user's voice through the microphone (130), at least one processor (140) can receive, through the microphone (130), a voice output through the speaker (160) in addition to the user's voice.
[0205] Here, when the other party's spoken voice output through the speaker (160) is received by the microphone (130), at least one processor (140) can record the time at which the voice was received. Since this has been described above in FIG. 2, a redundant description will be omitted.
[0206] Meanwhile, at least one processor (140) may output the above-described caution information through the speaker (160). For example, if at least one of the first summary information and the second summary information includes information corresponding to a primary information type, at least one processor (140) may output a warning sound or a voice such as "Information about the appointment time may be inaccurate" through the speaker (160).
[0207] FIG. 4 is a diagram illustrating the operation of an electronic device according to one or more embodiments of the present disclosure.
[0208] Referring to FIG. 4, a first screen (410) and a second screen (420) are respectively illustrated. The first screen (410) may include first summary information (411), and the second screen (420) may include second summary information (421) and dialogue text (422).
[0209] When a user ends a call through the electronic device (100), the electronic device (100) can obtain first summary information. The electronic device (100) can output a first screen (410) indicating the first summary information.
[0210] Here, the first summary information (411) can be displayed as brief keywords, such as "#wedding" or "#dress shop." The first summary information (411) can also display an update icon next to the keywords, indicating an update. The update icon can indicate that the first summary information is updatable. However, this is not a limitation.
[0211] Here, the multiple keywords included in the first summary information (411) may correspond to keywords generated during the call. For example, the electronic device (100) may record the call for a set period of time after the call is initiated, and generate keywords based on the recorded voice after the set period of time has elapsed. Thereafter, the electronic device (100) may display the generated keywords as the first summary information (411) during the call.
[0212] However, the multiple keywords are not limited to the examples described above, and may correspond to keywords generated after the call ends. For example, the electronic device (100) may generate keywords based on voices during multiple partial periods during the call. In this case, the electronic device (100) may display the keywords generated after the call ends as first summary information (411).
[0213] Thereafter, if the electronic device (100) is identified as being in at least one of multiple active states, the electronic device (100) may display a second screen (420). The second screen (420) may include the subject matter, keywords, and key information of the call. As this has been described in FIG. 1, a redundant description will be omitted.
[0214] The electronic device (100) can display dialogue text (422) together with second summary information (421) on a second screen (420). Here, the dialogue text (422) may correspond to text displayed in a dialogue format. Here, the text may correspond to a sentence obtained by converting second voice data.
[0215] Meanwhile, text can be displayed in a conversational format, differentiated by the call participants. For example, if a user and a caller are on a call, the second voice data may include the spoken voices of both the user and the caller. Furthermore, the second voice data may include information about the subject of the spoken voice, i.e., the participant who uttered the spoken voice. Furthermore, the second voice data may include information about the timing of each participant's speech (or the order of their speech).
[0216] At this time, in the dialogue text (422), sentences converted from the user's spoken voice may be displayed on the right, and sentences converted from the other party's spoken voice may be displayed on the left. Additionally, in the dialogue text (422), sentences uttered by each participant may be displayed from top to bottom according to the speaking time (or order).
[0217] Accordingly, the user can check brief first summary information (411) through the first screen (410), and if the electronic device (100) is in an activated state, the user can check second summary information (421) that is more detailed than the first summary information (411) through the second screen (420).
[0218] At this time, the electronic device (100) can display the conversation content together with conversation text (422), thereby enabling the user to easily search for necessary information in addition to the second summary information (421).
[0219] FIG. 5 is a diagram illustrating an activation state according to one or more embodiments of the present disclosure.
[0220] According to FIG. 5, after a call is terminated, the electronic device (100) may display a summary screen (500) according to an activation condition. Here, the summary screen (500) may include a first screen (510) and a second screen (520).
[0221] Here, the first screen (510) may correspond to a screen displaying keywords of the second summary information. Meanwhile, the second screen (520) may correspond to a screen displaying the entire contents of the second summary information.
[0222] For example, if the electronic device (100) is implemented as a smartphone that provides a call function and a call summary function, after the call is ended, the smartphone can identify whether the current state of the smartphone is active.
[0223] At this time, the smartphone can detect that the smartphone is in a second active state when the smartphone is charging or not in use.
[0224] For example, the smartphone may initiate an action to obtain the second summary information if it identifies that the smartphone is charging and no user input has been detected for a set period of time after a call has ended.
[0225] The time set here may refer to the time at which the electronic device (100) determines that the user is not currently using the electronic device (100). The time set here may be set by user operation and may be automatically adjusted by the electronic device (100) according to the current time zone (e.g., dawn, daytime, evening time). However, the present invention is not limited thereto.
[0226] That is, the smartphone has sufficient power to obtain second summary information based on the entire call audio, and if the smartphone is identified as not being used for a certain period of time, the second summary information can be generated based on the second voice data for the entire call audio.
[0227] Additionally, if the user does not use the smartphone for a certain period of time, such as while sleeping or engaging in other activities, secondary summary information can be generated.
[0228] An electronic device (100) such as the above-described smartphone can utilize a neural network model (summary model) installed within the electronic device (100).
[0229] Here, the neural network model installed within the electronic device (100) may be referred to as an on-device AI model. Here, the on-device AI model may refer to an AI model within a device such as a smartphone. The on-device summary model may correspond to a model that processes voice data within the electronic device (100), converts the voice into text, and obtains summary information.
[0230] The on-device summary model is advantageous in protecting personal information because all data is processed within the device, and it has the advantage of enabling real-time processing even without an Internet connection.
[0231] However, due to limitations in device performance (battery and memory performance), processing speeds are somewhat slower than when using AI models installed on servers or the cloud. Furthermore, when an electronic device (100) uses an on-device summary model, excessive power consumption may occur relative to the battery capacity of the electronic device (100).
[0232] As a result, if the electronic device (100) is not charging, generating the second summary information through the summary model built into the electronic device (100) may result in the electronic device (100) being powered off due to excessive battery consumption. If the electronic device (100) is powered off while generating the second summary information, it may have a fatal impact on the electronic device (100).
[0233] Additionally, while the electronic device (100) is generating second summary information using a summary model built into the electronic device (100), the electronic device (100) may not be able to provide functions other than the summary function.
[0234] If the electronic device (100) is the first activation condition, and the electronic device (100) initiates an operation to obtain the second summary information, the second summary information can be obtained by automatically recording the call without battery problems or inconvenience to the user.
[0235] FIG. 6 is a diagram illustrating an activation state according to one or more embodiments of the present disclosure.
[0236] According to FIG. 6, the electronic device (100) can display a first screen (610) and a second screen (620) after the call ends.
[0237] The first screen (610) may include a UI (612) for setting the name (611) of a user-specified contact and whether to automatically summarize it.
[0238] The second screen (620) may include the name of the designated contact (621) and an icon (622) indicating that a call with the designated contact has been summarized.
[0239] For example, if the electronic device (100) is implemented as a smartphone that provides a call summary function, etc., after the call is ended, the smartphone can identify whether the current state of the smartphone is active.
[0240] Meanwhile, users can set up a preset contact list to automatically summarize the entire call when making a call to that contact. In this case, the user can manipulate the UI (612) displayed as "Call summarize" by setting individual contacts.
[0241] Here, when the user manipulates the UI (612) to display ON, when a call is made to 'John Brian Adams', the owner of the contact, an action may be set to be initiated to generate the second summary information after the call is ended.
[0242] Meanwhile, the smartphone can detect that it is in the first active state when a call is made to a contact specified as the callee according to user operation input and the call is ended.
[0243] The smartphone can record the entire call and generate second summary information based on the acquired second voice data, and display a summary completion icon (622). Here, the icon (622) can be displayed with content such as 'Call ended' indicating the end of the call.
[0244] Accordingly, users can conveniently receive second summary information automatically when making a call to important contacts, such as work-related contacts, by pre-designating such contacts.
[0245] Meanwhile, in addition to the case where the electronic device (100) is in the first activation state and the second activation state described above, the user may directly input an operation after ending a call to cause the electronic device (100) to generate second summary information.
[0246] In other words, the active state may refer to a state in which a user input requesting activation for summarizing call content has been entered. For example, the active state may include a state in which a user inputs an operation to directly generate summary information.
[0247] FIG. 7 is a diagram illustrating an activation state according to one or more embodiments of the present disclosure.
[0248] Referring to FIG. 7, when a call is ended, the electronic device (100) may display a call end screen (700). The call end screen (700) may include a UI (710) that summarizes the entire call.
[0249] After the call ends, there may be cases where the call content is not automatically summarized because the electronic device (100) is not in at least one of the first activation state and the second activation state.
[0250] For example, if all of the following conditions are satisfied: the electronic device (100) is not charging, no user operation is input for a specified period of time, and the most recently ended call partner is not the designated call partner, the electronic device (100) may not initiate an operation to summarize the entire call.
[0251] A user can input an operation to select a UI (710) displayed on an electronic device (100) to control the electronic device (100) to obtain second summary information by summarizing the entire call content.
[0252] FIG. 8 is a diagram illustrating an activation state according to one or more embodiments of the present disclosure.
[0253] According to FIG. 8, after a call is ended, the electronic device (100) may display a call end screen (800). Here, the call end screen (800) may include information about the recently ended call (recipient information, first summary information, phone number, etc.) and a UI (810) that summarizes the entire call.
[0254] The user can input an operation to select the UI (810) to control the electronic device (100) to obtain second summary information by summarizing the entire call content, as in FIG. 7.
[0255] FIG. 9 is a diagram illustrating second summary information according to one or more embodiments of the present disclosure.
[0256] According to FIG. 9, after a call is terminated, the electronic device (100) may display a summary screen (900). Here, the summary screen (900) may include a first screen (910) and a second screen (920).
[0257] Here, each of the first screen (910) and the second screen (920) may include second summary information (911, 921).
[0258] The second summary information (911) included in the first screen (910) may include a plurality of keywords, and the second summary information (921) included in the second screen (920) may include a call topic, keywords, main information, etc.
[0259] For example, as described in FIGS. 7 and 8, if second summary information is generated according to a user operation input after the call is terminated, the electronic device (100) may display the second summary information on the first screen (910) or the second screen (920).
[0260] For example, a user can check the second summary information (911) organized by brief keywords through the recent call history with 'Jinu' on the first screen (910). The user can select the recent call history with 'Jinu' and check more detailed second summary information (921) through the second screen (920).
[0261] FIG. 10 is a diagram illustrating a subsequent screen after obtaining second summary information according to one or more embodiments of the present disclosure.
[0262] Referring to FIG. 10, the electronic device (100) may display a subsequent screen (1000) after the second summary information is obtained. The subsequent screen (1000) may include a title (1010), schedule information (1020), and location information (1030).
[0263] Here, the subsequent screen (1000) may correspond to a screen displayed in relation to the second summary information after the electronic device (100) displays the second summary information.
[0264] For example, if the second summary information includes information about the appointment location and schedule, the electronic device (100) may execute a schedule-related application. The schedule-related application may be an application that allows the user to input and save schedules so that they do not forget them. However, the present invention is not limited thereto.
[0265] At this time, the electronic device (100) can utilize a trained neural network model to output application information to be executed using the second summary information as input data. Here, the neural network model can be implemented through LLM. The neural network model can output information about an application predicted to be executed by the user based on keywords in the second summary information.
[0266] Thereafter, if the second summary information includes information about the appointment time and appointment location, the electronic device (100) may execute a calendar application, etc. At this time, the electronic device (100) may display the appointment time and location, etc. included in the second summary information, on the corresponding application screen.
[0267] For example, the electronic device (100) may display the content corresponding to the 'topic' among the second summary information in the title (1010) section and display the date and time in the schedule information (1020). Here, the electronic device (100) may display a UI that allows setting a notification for the corresponding schedule, such as 'all day', in the section where the schedule information (1020) is displayed.
[0268] Additionally, the electronic device (100) can display the content corresponding to the appointment location among the second summary information in the location information (1030) section.
[0269] The subsequent screen (1000) may correspond to the screen where the calendar application is executed, as described above, but is not necessarily limited thereto. For example, if the second summary information includes information related to "content delivery" after a work-related call, the electronic device (100) may execute a messaging application or an email application, etc.
[0270] For example, after a business-related call is made, if the second summary information includes information about the 'mail recipient' and 'delivery content', the electronic device (100) can execute an email application for sending the email.
[0271] The electronic device (100) can display an email address and ‘delivery contents’, etc. on the email application screen based on the contact information for the ‘mail recipient’ in the executed email application.
[0272] Through this, the electronic device (100) can predict the user's subsequent actions based on the second summary information. The electronic device (100) can automatically execute an application predicted to be executed by the user and display necessary information thereon, thereby increasing user convenience.
[0273] FIG. 11 is a diagram illustrating a call type according to one or more embodiments of the present disclosure.
[0274] According to FIG. 11, the electronic device (100) can display a first screen (1110) and a second screen (1120) after a call is terminated. Here, each of the first screen (1110) and the second screen (1120) can display second summary information (1111, 1121).
[0275] Here, the first screen (1110) may correspond to a screen displayed by the electronic device (100) when the electronic device (100) identifies the type of a recently ended call as a business call.
[0276] For example, an electronic device (100) can identify the call type as a business call through a classification model based on voice data (second voice data) for the entire call. In this case, the electronic device (100) can display second summary information according to a template suitable for the business call type. This also applies to FIGS. 11 and 12 .
[0277] Here, a template can refer to a framework with a specific format or structure for displaying information. Templates can be used to efficiently perform specific tasks or document creation. Here, a template suitable for a specific currency type can refer to a framework that can be used to efficiently convey information according to the characteristics of that currency type. This also applies to Figures 11 and 12.
[0278] For example, a business call might include information about the call's purpose (topic), instructions, and schedule (e.g., deadlines). In this case, a template appropriate for this type of call might consist of the "topic," "to-do," and "schedule" structures.
[0279] For example, the second summary information (1111) displayed on the first screen (1110) can be structurally displayed according to an agenda (topic) such as 'Ondevice project', a to do (task) such as 'send email', and a schedule such as 'revise report on May 30'.
[0280] Meanwhile, the second screen (1120) may correspond to a screen displayed by the electronic device (100) when the electronic device (100) identifies the type of a recently ended call as a reservation call.
[0281] Here, the second summary information (1121) displayed on the second screen (1120) may be displayed as a template consisting of the reservation location name, reservation schedule, number of people, and location structure.
[0282] For example, the second summary information (1121) may be displayed as the name of the reservation location such as 'Gopchangwang', the reservation schedule such as '2024. 05. 12 (Fri) 7:00 PM', the number of reservations such as '3 people', and the reservation location such as 'Map for reservation location'.
[0283] In this way, the electronic device (100) can identify one of a plurality of call types based on the voice data of the call and provide second summary information (1111, 1121) as a template suitable for the identified type.
[0284] Meanwhile, the multiple call types are not limited to the business call type and reservation call type as described above, and may include various other call types depending on the call content.
[0285] FIG. 12 is a diagram illustrating a call type according to one or more embodiments of the present disclosure.
[0286] According to FIG. 12, the electronic device (100) can display a third screen (1210) and a fourth screen (1220) after the call is terminated. Here, each of the third screen (1210) and the fourth screen (1220) can display second summary information (1211, 1221).
[0287] Here, the third screen (1210) may correspond to a screen displayed by the electronic device (100) when the electronic device (100) identifies the type of the recently ended call as an inquiry call.
[0288] Here, the second summary information (1221) displayed on the third screen (1210) may be displayed as a template consisting of ‘inquiry’ and ‘answer corresponding to inquiry’.
[0289] For example, the second summary information (1211) may be displayed as inquiries such as ‘Inquiry about Sudme makeup price’, ‘Inquiry about Sudme makeup schedule’, and replies such as ‘Estimated at 2.5 to 3 million won’, ‘May 6 schedule available’.
[0290] Meanwhile, the fourth screen (1220) may correspond to a screen displayed by the electronic device (100) when the electronic device (100) identifies the type of the recently ended call as a daily call.
[0291] For example, the electronic device (100) may identify a call type as not falling into any of the various types described above (business call, reservation call, inquiry call). In this case, the electronic device (100) may classify the call type as a daily call. However, this is not a limitation.
[0292] Here, the second summary information (1221) displayed on the fourth screen (1220) may be displayed as a template consisting of a subject and key information structure. The content included in the subject and key information has been described in Figure 1, so any redundant explanation will be omitted.
[0293] As illustrated in FIGS. 11 and 12, the electronic device (100) can identify one of a plurality of call types (business call, reservation call, inquiry call, daily call) based on voice data of a call, and provide second summary information (1111, 1121, 1211, 1221) as a template suitable for the identified type.
[0294] Meanwhile, the multiple currency types are not limited to the examples described above, and may include various other currency types depending on the content of the call.
[0295] FIG. 13 is a flowchart illustrating an operation of obtaining summary information based on the time of utterance according to one or more embodiments of the present disclosure.
[0296] Fig. 13 illustrates an operation in which an electronic device (100) receives a user's voice through a microphone and obtains second summary information.
[0297] An electronic device (100) can typically receive a user's voice through a microphone and record a voice signal corresponding to the other party's voice through a communication device (120) to record the voice during a call. However, there may be cases where it is impossible to record the other party's voice through the communication device (120).
[0298] For example, call recording may be prohibited in areas where users make calls via electronic devices (100). For example, the laws of a particular country (or state) may prohibit recording the voice of a caller without the other party's consent.
[0299] Even if call recording isn't prohibited in the region where the user is making the call, the other party may not consent to the recording. However, this isn't the only case; there may be various circumstances where recording the other party's voice is prohibited.
[0300] In this case, the electronic device (100) can receive the user's voice through the microphone and summarize the call content. Alternatively, the electronic device (100) can record the user's voice as well as the other party's voice output from the electronic device (100) and an external voice output device through the microphone.
[0301] The electronic device (100) can obtain summary information based on the time at which the other party's voice was received through the microphone, along with the user's voice data recorded through the microphone. The operation of the electronic device (100) receiving the other party's voice has been described in FIG. 3, so a redundant description will be omitted.
[0302] The electronic device (100) can obtain second summary information through the information collection step (S1310 to S1340) and the summary information acquisition step (S1350 to 1380).
[0303] In the information collection step, the electronic device (100) can identify whether all participants (all call participants) have consented to recording (S1310). If all participants have consented to recording, the electronic device (100) can record the entire call (S1320). On the other hand, if all participants have not consented to recording, i.e., if even one participant has not consented to recording, the electronic device (100) can record microphone input sound and collect the time of sound generation of the other participants (the remaining participants excluding the user) (S1330). The electronic device (100) can end the call (S1340).
[0304] Here, the time of sound generation may mean the time when the electronic device (100) receives the voice of the other participant through the microphone.
[0305] The electronic device (100) can record the entire voice of a call until the call ends. Alternatively, the electronic device (100) can record microphone input sound and collect the time at which the received sound occurs.
[0306] In the summary information acquisition step, the electronic device (100) can convert the recorded voice into text (S1350). Here, the electronic device (100) can convert the recorded voice into text using the voice recognition model described above. Here, the voice may refer to the user's voice.
[0307] Thereafter, the electronic device (100) can infer the other party's speech timing (S1360). Here, the electronic device (100) can utilize a neural network model implemented as an LLM. Here, the electronic device (100) can infer the other party's speech timing using information about the time of occurrence of the received sound.
[0308] Here, the electronic device (100) can infer the time of utterance by calculating back a certain amount of time from the time of occurrence of the received sound. Since this has been described in FIG. 1, a redundant description will be omitted.
[0309] Meanwhile, the electronic device (100) can also record the duration of the other party's voice through the period during which the other party's voice is continuously received. In other words, the electronic device (100) can record the duration of the other party's speech (hereinafter, "speech duration") through the period during which the user's voice is continuously received. This will be described later in FIG. 17.
[0310] Thereafter, the electronic device (100) can integrate the entire conversation data (S1370). The integrated conversation data may include text obtained by converting the user's voice and information about the other party's speech timing. The information about the speech timing may include information about the other party's inferred speech timing and information about the duration of the speech.
[0311] Here, the integrated conversation data may also include information about the chronological relationship between the user's utterance and the other party's utterance. In other words, the integrated conversation data may correspond to data in which the user's utterance content (text) and the utterance timing (and duration) are aligned according to the chronological order of the utterance timing.
[0312] Thereafter, the electronic device (100) can interpret the conversation context and obtain summary information (S1380). Here, the electronic device (100) can analyze the conversation context (or context) and obtain summary information through a neural network model implemented as an LLM.
[0313] Here, LLM can be a model trained to analyze context and obtain summary information using integrated conversation data as input data.
[0314] Here, context can refer to background information, including previous utterances, topics, and participant information, that occur during a conversation. Context can be utilized by LLM to accurately understand the user's intent and generate consistent responses. In other words, context can encompass all linguistic and situational elements necessary to understand the flow of the conversation and provide meaningful responses.
[0315] For example, context may include information about the time and place of each utterance by the other party or the user. Context may also include the duration of the user's utterance.
[0316] Through this, the electronic device (100) can obtain summary information by analyzing the conversation context through the above neural network model based on the user's speech content (text) and the user's speech time (and speech duration).
[0317] Afterwards, when summary information is acquired, the electronic device (100) can display the conversation history and summary (S1390). Here, the conversation history can be displayed as the conversation text described above in FIG. 4. Here, the summary can refer to second summary information, and can be displayed in various templates depending on the call type, as shown in FIGS. 11 and 12.
[0318] Accordingly, even if the electronic device (100) cannot record a call including the other party's voice, it can obtain summary information through a neural network model (LLM) using only the user's spoken voice and the user's speaking time.
[0319] That is, the electronic device (100) can obtain summary information with improved accuracy by using the user's speech timing compared to when using only one person's speech voice.
[0320] FIG. 14 is a drawing for explaining whether the other party's speech is recorded according to one or more embodiments of the present disclosure.
[0321] According to FIG. 14, the screen (1400) may include a first icon (1410) and a second icon (1420).
[0322] The screen (1400) may include recent call records, including recently ended call records (call records with 'Jinu'). Here, each call record may include an icon in the form of a first icon (1410) or a second icon (1420).
[0323] Here, the first icon (1410) may indicate that the other party's speech is not being recorded. For example, this may apply when the other party has not consented to call recording. In this case, the electronic device (100) may record only the user's voice based on user actions, such as not recording the other party's speech.
[0324] Here, the second icon (1420) may indicate that the other party's speech has been recorded. For example, this may indicate that the user is in an area where call recording is permitted, or that the other party has consented to call recording.
[0325] That is, the electronic device (100) displays an icon in the form of a first icon (1410) or a second icon (1420) for each recent call, thereby allowing the user to check whether the other party's speech has been recorded.
[0326] In general, summary information in cases where the voice of a single speaker (user) is recorded may be less accurate than in cases where the voice of a single speaker (user) is not recorded. Accordingly, the user (10) can use the first icon (1410) and the second icon (1420) to check whether the other party's speech has been recorded and determine whether the summary information is reliable.
[0327] For example, in a call recording where the voice of a single speaker is recorded, the user can select the call recording to view the detailed conversation (conversation text) to avoid the possibility that the summary information may be inaccurate. The detailed conversation text may include information about the user's speech (text) and the time of the other party's speech.
[0328] Through this, the electronic device (100) can allow the user to check whether the other party's speech has been recorded and induce subsequent actions (such as checking detailed conversation content).
[0329] FIG. 15 is a diagram illustrating a conversation record according to one or more embodiments of the present disclosure.
[0330] According to FIG. 15, the electronic device (100) can display a first screen (1510) or a second screen (1520).
[0331] Here, the first screen (1510) may correspond to a screen in which the other party's speech is not recorded. On the other hand, the second screen (1520) may correspond to a screen in which the other party's speech is recorded.
[0332] Here, the first screen (1510) may include a first icon (1511) indicating whether the other party's speech is recorded. Here, the first icon (1511) may mean that the other party's speech is not recorded.
[0333] On the other hand, the second screen (1520) may include a second icon (1521) indicating whether the other party's speech has been recorded. Here, the second icon (1521) may mean that the other party's speech has not been recorded.
[0334] Meanwhile, the first screen (1510) and the second screen (1520) may each include a first dialogue text and a second dialogue text. Here, the first dialogue text and the second dialogue text may include a record of the other party's speech (1521, 1522).
[0335] The first dialogue text may include the user's utterances and the time and duration of the other party's utterances. Conversely, the second dialogue text may include both the user's and the other party's utterances.
[0336] That is, if only the user's speech is recorded during a call, the electronic device (100) can display only the content of the user's speech, and for the other party's speech, the time of speech can be displayed as the other party's speech record (1512). Here, the time of speech may correspond to the time inferred based on the time when the other party's voice was received by the microphone of the electronic device (100). Since this has been described in FIG. 1, a duplicate explanation will be omitted.
[0337] For example, the electronic device (100) may display a first conversation text and a second conversation text based on the timing of the other party's speech and the timing of the user's speech. Here, each of the first conversation text and the second conversation text may include a speech bubble differentiated for each of the call participants (the user and the other party).
[0338] Here, the other party's speech records (1512, 1522) included in each of the first screen (1510) and the second screen (1520) may be displayed as speech bubbles. In addition, the user's speech records included in each of the first screen (1510) and the second screen (1520) may also be displayed as speech bubbles.
[0339] The speech bubbles included in each of the first dialogue text (1512) and the second dialogue text (1522) are only examples and may be displayed in various forms.
[0340] In the first dialogue text (1512), the multiple speech bubbles on the right may contain the user's speech content. On the other hand, the other party's speech records (1512) (multiple speech bubbles on the left) may each contain information about the time and duration of the other party's speech.
[0341] Here, the position of each of the multiple speech bubbles on the left can indicate the time of utterance. For example, between the user saying "Hello" and "I ate. How about you?", the other party's speech may be received through the microphone. In this case, the speech bubble corresponding to the other party's utterance may be located between the speech bubble labeled "Hello" and the speech bubble labeled "I ate. How about you?"
[0342] At this time, the electronic device (100) only records the time and duration of the other party's utterance, so it may not be displayed as a sentence that expresses the exact meaning, such as 'Hello' or 'I ate. How about you?' That is, the electronic device (100) may display the other party's utterance as a speech bubble with characters such as '....'.
[0343] Meanwhile, the size (or length) of each of the multiple speech bubbles on the left or the number of '.' in '....' may indicate the duration of the speech. That is, if the conversation data includes information on the duration of the other party's speech, the electronic device (100) may display the speech bubble by adjusting the size of the speech bubble or the number of '.' according to the duration of the speech.
[0344] Through this, the electronic device (100) can display the timing and length of the other party's speech, rather than displaying the exact content of the other party's speech. Accordingly, the electronic device (100) can enable the user to intuitively understand the flow of the conversation.
[0345] FIG. 16 is a drawing for explaining caution information according to one or more embodiments of the present disclosure.
[0346] According to FIG. 16, the electronic device (100) can display a first screen (1610) and a second screen (1620).
[0347] Here, the first screen (1610) may correspond to a screen on which the electronic device (100) displays second summary information after the call ends. Here, the second summary information may include key content such as "Greetings," "Preparing for the Dress Shop Tour," "Meeting Time and Place," and "Call End."
[0348] Here, if you select ‘Meeting time and place’ (1611) from the second summary information displayed on the first screen (1610), detailed information about the ‘Meeting time and place’ can be displayed on the second screen (1620).
[0349] Here, the second summary information displayed on the second screen (1620) may include content such as ‘See you in front of A mart at 9:00 AM on Saturday.’
[0350] For example, a user might say, "Let's meet in front of A Mart at 9 AM on Saturday!" while on a call with a caller. The caller might respond, "Okay. Then let's meet there on Saturday!" The user might then respond, "Okay." In this case, since the actual appointment time and location have been confirmed as "9 AM on Saturday in front of A Mart," the second summary information can accurately reflect the call content (appointment time and location).
[0351] However, if the other party responds by saying, "No, let's meet at 10 AM on Saturday!", then if the user responds by saying, "Okay," the recorded voice data (second voice data) of the call may only include information indicating that the user said, "Let's meet at 9 AM on Saturday in front of A mart!" and "Okay." Even in this case, the electronic device (100) may obtain second summary information including the information, "Let's meet at 9 AM on Saturday in front of A mart from the voice data."
[0352] Here, the actual appointment time and location are confirmed as '10 AM on Saturday in front of A mart', but the second summary information may contain incorrect information.
[0353] In this way, when the electronic device (100) summarizes based on voice data that only records the user's speech, it may include inaccurate information.
[0354] At this time, the electronic device (100) can identify that the second summary information includes information corresponding to a key information type, such as 'meeting time and place'. Here, since the key information and key information types have been described above in FIG. 3, any redundant description will be omitted.
[0355] Accordingly, the electronic device (100) can display caution information. Here, since the caution information has been described above in FIG. 3, a duplicate description will be omitted.
[0356] For example, the electronic device (100) may display caution information such as 'Caution! The speech of the other party to the call is not reflected and may be inaccurate.'
[0357] However, this is only an example, and the electronic device (100) can display whether the other party's speech has been recorded (i.e., whether the second voice data includes voice data corresponding to the other party's speech), such as 'Caution! The other party's speech has not been reflected.'
[0358] Additionally, the electronic device (100) may display warning information more simply, such as "Summary may be inaccurate!" In addition to the examples described above, the electronic device (100) may display various forms of display (images, videos, screen special effects, etc.) to warn that the summary information may be inaccurate.
[0359] Through this, if the conversation content relates to a key topic (e.g., appointment time and location), the electronic device (100) displays the above-mentioned cautionary information, allowing the user to determine the reliability of the summary information. The user can then call the other party again, confirm the exact information via text message, or check the conversation text (conversation log) to confirm the exact information.
[0360] FIG. 17 is a diagram illustrating a conversation record of a video conference according to one or more embodiments of the present disclosure.
[0361] According to FIG. 17, the electronic device (100) can display a video call screen (1710). Here, the video call screen (1710) can display images of multiple participants participating in the video call.
[0362] For example, the electronic device (100) may provide a video call (video conference) function for three or more participants, in addition to a voice call for two participants (the user and the other party). Even in this case, there may be cases where multiple parties (the participants other than the user) do not consent to the recording of the call.
[0363] When a video call ends, the electronic device (100) may display a summary screen (1720). Here, the summary screen (1720) may display second summary information. Here, the second summary information may include an icon (1721) indicating whether the other party's speech is recorded.
[0364] Meanwhile, the summary screen (1720) may include a dialogue text (1722). Here, the user's speech record may be placed on the right, and the speech records of multiple counterparts (Manager A, Manager B, Manager C) may be placed on the left. Here, the speech records may be displayed as multiple speech bubbles (1722-1 to 1722-3), but this is not limited thereto.
[0365] Here, the position of each utterance record can indicate the time of utterance. Furthermore, the times contained within the speech bubbles (e.g., 22s, 57s, 38s) can indicate the duration of the utterance. Duration can be expressed in hours, minutes, or seconds, or simply as h, m, or s.
[0366] That is, the duration of the other party's speech can be expressed by the size (length) of the speech bubble or the number of '.' as shown in Figure 15, and the duration of the speech can also be expressed by h, m, s as shown in Figure 17. However, it is not limited thereto.
[0367] For example, after a user utters, "How should I do this report?", "Manager A" may speak for 22 seconds, then the user may utter, "I think I should look for AI-related benchmarking services," then "Manager B" may speak for 57 seconds, and then "Manager C" may speak for 38 seconds.
[0368] In this case, the electronic device (100) can display a dialogue text (1722) that includes, in order, a speech bubble that says, "How should I do this report?", a speech bubble (1722-1) that says, "Manager A" and "22s," a speech bubble that says, "I think I should look for an AI-related benchmarking service," a speech bubble (1722-2) that says, "Manager B" and "57s," and a speech bubble (1722-3) that says, "Manager C" and "38s."
[0369] Through this, the electronic device (100) can display the timing and length of each utterance, etc., in a speech bubble for each of three or more parties without displaying the exact utterance content. Accordingly, the user can intuitively grasp the flow of the conversation (meeting) content through a conversation record containing multiple speech bubbles.
[0370] FIG. 18 is a flowchart illustrating a method for controlling an electronic device according to one or more embodiments of the present disclosure.
[0371] The electronic device (100) can obtain first summary information through a neural network model based on first voice data obtained from a portion of a call (S1810).
[0372] According to one or more embodiments, the electronic device (100) may obtain first summary information through a neural network model based on first voice data corresponding to voice continuously input for a set period of time after a call is initiated among voice data.
[0373] According to one or more embodiments, the electronic device (100) may obtain first summary information based on first voice data corresponding to a plurality of different partial periods during the period during which the call was performed after the call is terminated.
[0374] According to one or more embodiments, the neural network model may correspond to a model trained to output summary information based on voice data acquired by the electronic device (100).
[0375] Next, the electronic device (100) can obtain second summary information after the call is terminated (S1820).
[0376] According to one or more embodiments, the electronic device (100) can identify, after a call is terminated, whether the state of the electronic device is at least one of a plurality of active states capable of summarizing the content of the call corresponding to second voice data generated by recording the call until the end of the call.
[0377] Thereafter, if the electronic device (100) is identified as being in at least one of a plurality of activation states, the first summary information may be updated through at least one neural network model based on the second voice data to obtain second summary information.
[0378] According to one or more embodiments, the electronic device (100) may record a call until the end of the call, and may obtain point-in-time data by recording the point in time at which at least one spoken voice is received from at least one of the plurality of participants in the call, excluding the user of the electronic device.
[0379] After the call is terminated, if the state of the electronic device is identified as at least one of a plurality of active states, second summary information can be obtained through at least one neural network model based on the acquired point-in-time data and second voice data corresponding to the user's spoken voice.
[0380] Here, the neural network model may correspond to a model trained to output summary information based on voice data acquired by the electronic device and the time at which at least one spoken voice among multiple participants in the call is received.
[0381] Accordingly, the electronic device (100) can efficiently summarize a voice call without concerns about privacy invasion or personal information leakage by using an AI model (on-device AI model) installed within the electronic device (100).
[0382] Additionally, the electronic device (100) can initiate a full call summary when the user is not using the electronic device (100) and it is charging, or when a call summary is absolutely necessary (e.g., when calling a designated contact). Accordingly, other uses of the user's electronic device (100) (uses other than call functions) can be prevented from being interrupted.
[0383] Additionally, in situations where a quick summary is not required, such as when the user is sleeping with the smartphone charged, the electronic device (100) can generate more accurate summary information using sufficient power and time.
[0384] Meanwhile, in Fig. 18, the order is mapped for all steps for convenience of explanation, but it is of course not necessarily limited to the order of steps that are not related to the order or can be performed in parallel.
[0385] Meanwhile, the methods according to at least some of the various embodiments of the present disclosure described above can be implemented in the form of applications that can be installed on existing electronic devices.
[0386] Additionally, the methods according to at least some of the various embodiments of the present disclosure described above can be implemented with only a software upgrade or a hardware upgrade for an existing electronic device.
[0387] Additionally, the methods according to at least some of the various embodiments of the present disclosure described above may also be performed through an embedded server provided in an electronic device, or an external server of at least one of the electronic devices.
[0388] Meanwhile, according to one embodiment of the present disclosure, the various embodiments described above can be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device is a device that can call instructions stored from the storage medium and operate according to the called instructions, and may include an electronic device (e.g., electronic device (100)) according to the disclosed embodiments. When an instruction is executed by a processor, the processor can perform a function corresponding to the instruction directly or by using other components under the control of the processor. The instruction may include code generated or executed by a compiler or an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the 'non-transitory storage medium' means a tangible device and does not include a signal (e.g., electromagnetic wave), and this term is used to refer to a case where data is stored semi-permanently in the storage medium and a case where data is stored temporarily. No distinction is made. For example, a 'non-transitory storage medium' may include a buffer in which data is temporarily stored. According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a commodity. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones).In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily created in a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0389] Various embodiments of the present disclosure may be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device may include an electronic device (e.g., an electronic device (100)) according to the disclosed embodiments, which is a device capable of calling instructions stored in the storage medium and operating according to the called instructions.
[0390] When the above-described instruction is executed by the processor, the processor may perform the function corresponding to the instruction directly or by utilizing other components under the control of the processor. The instruction may include code generated or executed by a compiler or interpreter.
[0391] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.
Claims
1. In electronic devices, A memory storing at least one instruction and at least one neural network model; communication device; Mike; and At least one processor connected to the communication device, the microphone and the memory for controlling the electronic device; wherein said at least one neural network model comprises a model trained to output summary information based on voice data acquired by said electronic device, At least one processor, Based on first voice data obtained from a portion of a call through at least one of the communication device and the microphone, first summary information is obtained through the at least one neural network model, An electronic device, wherein after the call is terminated, if the state of the electronic device is identified as at least one of a plurality of activation states capable of summarizing the content of the call corresponding to second voice data generated by recording the call until the end of the call, the electronic device updates the first summary information through the at least one neural network model based on the second voice data to obtain second summary information.
2. In paragraph 1, At least one processor, An electronic device that obtains the first summary information through the neural network model based on the first voice data corresponding to voice continuously input for a set period of time after the call is initiated among the voice data.
3. In paragraph 1, At least one processor, An electronic device that obtains the first summary information based on first voice data corresponding to a plurality of different partial periods during the period during which the call was performed, after the call is terminated.
4. In paragraph 1, The above multiple activation states are: An electronic device comprising at least one of a first activation state in which a call is made with a contact specified as the callee according to a user operation input and the call is terminated, and a second activation state in which the electronic device is charging and the user operation input is not detected for a set period of time after the call is terminated.
5. In paragraph 1, The at least one neural network model includes a classification model trained to identify the content of the call as one of a plurality of types based on the second voice data, and a plurality of type-specific summary models trained to output second summary information corresponding to each of the plurality of types, At least one processor, Based on the second voice data obtained above, the classification model identifies a type corresponding to the content of the call among the plurality of types, and An electronic device that obtains the second summary information by using a summary model corresponding to the identified type among the plurality of type-specific summary models.
6. In paragraph 5, The above types of multiples are: Includes at least one of business calls, reservation calls, inquiry calls, and routine calls; At least one processor, If the above type is identified as a business call, the second summary information including at least one of the topic (Agenda), to-do, and schedule corresponding to the content of the call is obtained, If the above type is identified as a reservation call, the second summary information including at least one of a reservation schedule, number of people, and reservation location is obtained, An electronic device that, when the above type is identified as an inquiry call, obtains second summary information including the inquiry and a response corresponding to the inquiry.
7. In paragraph 1, The at least one neural network model includes a model trained to output summary information based on voice data acquired by the electronic device and a time point at which at least one spoken voice among the plurality of participants of the call is received through the microphone, At least one processor, Recording the call through the microphone until the end of the call, and obtaining point-in-time data by recording the point in time when at least one spoken voice is received from at least one of the parties excluding the user of the electronic device among the multiple participants of the call, An electronic device, wherein, after the call is terminated, if the state of the electronic device is identified as at least one of the plurality of activation states, the electronic device obtains the second summary information through the at least one neural network model based on the acquired point-in-time data and the second voice data corresponding to the user's spoken voice obtained through the microphone.
8. In paragraph 7, At least one processor, An electronic device that records the time at which a voice of at least one of the at least one counterparty received through the microphone is received when the intensity of the voice is identified as being higher than a predetermined intensity.
9. In paragraph 7, including display; At least one processor, Controlling the display to display caution information together with the second summary information when information included in at least one of the first summary information and the second summary information is identified as corresponding to at least one of a plurality of predetermined main information types; The above-described plurality of primary information types include at least one of a time type and a location type, An electronic device wherein the above-mentioned attention information includes information on whether the second voice data includes voice data corresponding to at least one spoken voice of the other party.
10. In paragraph 1, including display; The above neural network model includes a model trained to convert input voice data into text and output it, At least one processor, An electronic device that controls the display to display the text obtained by inputting the second voice data into the neural network model, together with the second summary information, separately for each of the plurality of participants making the call.
11. In a method for controlling an electronic device, A step of obtaining first summary information through at least one neural network model based on first voice data obtained from a portion of a call; and After the call is terminated, if the state of the electronic device is identified as at least one of a plurality of activation states capable of summarizing the content of the call corresponding to second voice data generated by recording the call until the end of the call, a step of obtaining second summary information by updating the first summary information through the at least one neural network model based on the second voice data; A control method, wherein the at least one neural network model comprises a model learned to output summary information based on voice data acquired by the electronic device.
12. In paragraph 11, The step of obtaining the above first summary information is: A control method comprising: a step of obtaining the first summary information through the neural network model based on first voice data corresponding to voices continuously input for a set period of time after the call is initiated among the voice data; 13. In paragraph 11, The step of obtaining the above first summary information is: A control method comprising: a step of obtaining the first summary information based on first voice data corresponding to a plurality of different partial periods during the period during which the call was performed, after the call is terminated.
14. In paragraph 11, The above multiple activation states are: A control method comprising at least one of a first activation state in which a call is made with a contact specified as the callee according to a user operation input and the call is terminated, and a second activation state in which the electronic device is charging and the user operation input is not detected for a set period of time after the call is terminated.
15. A non-transitory computer-readable recording medium storing computer instructions that, when executed by a processor of an electronic device, cause the electronic device to perform an operation, the operation comprising: A step of obtaining first summary information through at least one neural network model based on first voice data obtained from a portion of a call; and After the call is terminated, if the state of the electronic device is identified as at least one of a plurality of activation states capable of summarizing the content of the call corresponding to second voice data generated by recording the call until the end of the call, a step of obtaining second summary information by updating the first summary information through the at least one neural network model based on the second voice data; A non-transitory computer-readable recording medium, wherein the at least one neural network model comprises a model learned to output summary information based on voice data acquired by the electronic device.
Citation Information
Patent Citations
Call key information acquisition method and device
CN113590828A
Service process for mobile phone application of call end alarm service and contents
KR101631292B1
Game apparatus
KR102536329B1
Systems For Summarizing Contact Center Calls And Methods Of Using Same
US20220038577A1
KR20190059151A