Call voice processing method, electronic device, and program
The conversion of call audio to text with speaker information addresses the challenge of missed audio, allowing users to review past calls through user inputs and screen interactions, enhancing communication quality.
Patent Information
- Application Number
- JP2025118489
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-15
- Filing Date
- 2025-07-14
- Publication Date
- 2026-01-27
AI Technical Summary
Users often miss call audio due to unexpected problems or multitasking, making it difficult to check the contents of past calls during real-time communication.
A method and electronic device that converts real-time call audio to text, displays it with speaker information, and allows users to review past call content via user inputs or screen switching.
Enables users to confirm call content both audibly and visually, supporting smooth call continuation and easy review of past conversations even in multitasking environments.
Smart Images

Figure 2026012660000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a call audio processing method and an electronic device. [Background technology]
[0002] With the advancement of communication technology, there has been active development of technologies related to voice and video communication. For example, centered on the Internet, which is developing as an integrated network environment, VoIP (Voice over Internet Protocol) enables everything from one-to-one voice and video calls to multi-party voice and video calls involving many users. Furthermore, by being applied to various fields such as SNS (Social Network Service) and games, VoIP has evolved into a multi-party interactive immersive communication technology that converts voices, music, and other sounds input by many participants into three-dimensional sound, giving participants a sense of immersion.
[0003] However, in a telephone call, the voices of the participants are transmitted and received in real time, so it is possible that a user may miss the call audio. For example, a user may be unable to hear the call audio due to an unexpected problem occurring during the call, or may be unable to hear the call audio when using another application during a call in a multitasking environment. However, if a user misses the call audio, it is difficult to check the contents of past calls during the call because the call is conducted in real time. Therefore, there is a need for the development of technology that allows users to check the contents of past calls during a call or to check the contents of calls even when the screen is switched in a multitasking environment, etc. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Korean Patent Registration No. 10-2801805 Summary of the Invention [Problem to be solved by the invention]
[0005] The present disclosure provides a speech processing method and electronic device that solves the above-mentioned problems. [Means for solving the problem]
[0006] The present disclosure can be embodied in numerous forms, including as a method, an apparatus (system) and / or a computer program product.
[0007] According to one embodiment of the present disclosure, a call audio processing method executed by at least one processor may include steps of converting real-time call audio to text, displaying the converted text on a first screen together with information about the speaker of the real-time call audio, and, in response to receiving a first user input, displaying a first text converted from call audio at a first time point, which is a predetermined time before the current time, on the first screen together with information about the speaker of the call audio at the first time point.
[0008] A program for causing a computer to execute the call voice processing method according to an embodiment of the present disclosure may be provided. Also, a program recorded on a computer-readable recording medium may be provided for causing a computer to execute the call voice processing method according to an embodiment of the present disclosure.
[0009] An electronic device according to one embodiment of the present disclosure includes a display, a memory, and at least one processor connected to the display and the memory and configured to execute at least one computer-readable program stored in the memory, wherein the at least one program may include instructions for converting real-time call audio into text, displaying the converted text on a first screen of the display together with information about a speaker of the real-time call audio, and, in response to receiving a first user input, displaying first text converted from call audio at a first time point that is a predetermined time before the present time on the first screen together with information about the speaker of the call audio at the first time point. [Effects of the Invention]
[0010] One embodiment of the present disclosure can improve service quality related to calls by displaying text converted from real-time call audio along with information about the speaker, allowing the content of real-time calls to be confirmed not only as audio information but also as visual information.
[0011] Furthermore, an embodiment of the present disclosure supports the user in smoothly continuing a call by displaying text converted from past call voice and information about the speaker in response to a user input.
[0012] Furthermore, one embodiment of the present disclosure supports users in easily checking the content of a call even in a multitasking environment by displaying text converted from real-time call audio and information about the speaker on the screen after switching screens.
[0013] The effects of the present disclosure are not limited to the above, and other effects not described will be obvious to a person with ordinary skill in the technical field to which the invention pertains from the claims. [Brief explanation of the drawings]
[0014] Embodiments of the present disclosure are described below with reference to the accompanying drawings, in which like reference numerals indicate similar elements, but are not limited to the drawings.
[0015] [Figure 1] 1 is a diagram illustrating an exemplary electronic device for processing speech audio in accordance with one embodiment of the present disclosure. [Figure 2] 1 is a schematic diagram illustrating a configuration in which an information processing system is communicably connected to a plurality of user terminals in relation to data processing in an embodiment of the present disclosure. [Figure 3] 1 is a block diagram illustrating an internal configuration of a user terminal and an information processing system according to an embodiment of the present disclosure. [Figure 4]FIG. 1 is a diagram illustrating the configuration of an electronic device for processing voice in a call in an embodiment of the present disclosure. [Figure 5] 10A and 10B are diagrams illustrating a method for displaying text converted from a call voice and information about the speaker on an execution screen of a call application in an embodiment of the present disclosure. [Figure 6] 10A and 10B are diagrams illustrating a method for processing real-time call voice in a state where text converted from past call voice and information about the speaker are displayed on the execution screen of a call application in one embodiment of the present disclosure. [Figure 7] 10A and 10B are diagrams illustrating a method for displaying text converted from a call voice and information about the speaker on another execution screen of a call application in an embodiment of the present disclosure. [Figure 8] This figure explains a method for processing real-time call voice in one embodiment of the present disclosure, with text converted from past call voice and information about the speaker displayed on another execution screen of the call application. [Figure 9] This figure explains a method for displaying text converted from a call voice and information about the speaker in PIP format when switched to an execution screen of an application other than a call application in one embodiment of the present disclosure. [Figure 10] 10A and 10B are diagrams illustrating a method for changing the size of the PIP region in one embodiment of the present disclosure. [Figure 11] 10A and 10B are diagrams illustrating a method for displaying text converted from a call voice and information about the speaker on an execution screen of a chat application linked to a call application in an embodiment of the present disclosure. [Figure 12] 10A and 10B are diagrams illustrating a method for processing sounds corresponding to designated patterns in a call voice in an embodiment of the present disclosure. [Figure 13] FIG. 1 illustrates a method for processing call logs in one embodiment of the present disclosure. [Figure 14]10A and 10B are diagrams illustrating a method for displaying text converted from past call voices and information about the speaker, among the call voice processing methods according to an embodiment of the present disclosure. [Figure 15] 10A and 10B are diagrams illustrating a method of displaying text converted from a call voice and information about the speaker on a screen after switching screens, in a call voice processing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0016] According to one embodiment, the step of displaying the first text on the first screen may include a step of displaying information about the first time point on the first screen together with the first text and information about the speaker of the call audio at the first time point.
[0017] According to one embodiment, the first user input may include an input to select a first object displayed on the first screen or an input to scroll in a first direction on the first screen.
[0018] According to one embodiment, the first user input may include an input to select a first object displayed on the first screen or an input to scroll in a first direction on the first screen.
[0019] According to one embodiment, the call audio processing method may further include, in response to receiving a second user input, displaying on the first screen a second text converted from the call audio at a second time point, which is later than the first time point, together with information about the speaker of the call audio at the second time point.
[0020] According to one embodiment, the call voice processing method may further include a step of displaying, after displaying the first text on the first screen, text converted from the current real-time call voice and information about the speaker of the real-time call voice in response to a predetermined time having elapsed or a second user input being received.
[0021] According to one embodiment, the second user input may include an input to select a second object displayed on the first screen or an input to scroll to an end point of a second direction on the first screen that is opposite to the first direction.
[0022] According to one embodiment, the call audio processing method may further include a step of translating the converted text into a language set for a user (e.g., a speaker or a recipient) if the language set for the speaker of the real-time call audio is different from the language set for the recipient of the real-time call audio.
[0023] According to one embodiment, the step of displaying the first text on the first screen includes a step of displaying the first text and information about the speaker of the call voice at the first time in a first area of the first screen, and the call voice processing method may further include a step of displaying text converted from the real-time call voice and information about the speaker of the real-time call voice in a second area of the first screen different from the first area in response to receiving the real-time call voice while the first text and information about the speaker of the call voice at the first time is displayed in the first area of the first screen.
[0024] According to an embodiment, the method for processing call voice may further include a step of storing log information about the call based on the text converted from the real-time call voice and information about the speaker of the real-time call voice.
[0025] According to one embodiment, storing log information about the call may include storing text converted from the real-time call audio and information about the speaker of the real-time call audio immediately after the call ends or in response to receiving a second user input.
[0026] According to one embodiment, storing log information about the call may include storing information about the time point of the real-time call audio, along with information about the text into which the real-time call audio is converted and the speaker of the real-time call audio.
[0027] According to an embodiment, the call voice processing method may further include a step of displaying log information related to the call on a second screen different from the first screen.
[0028] According to one embodiment, the step of displaying log information about the call on the second screen may include a step of displaying log information about the call on the second screen in response to receiving a second user input selecting a second object displayed on a third screen that is opened immediately after the call ends.
[0029] According to one embodiment, the call audio processing method may further include a step of displaying specific text in the log information about the call displayed on the second screen differently from other text in response to receiving a second user input selecting specific text displayed on the first screen.
[0030] According to one embodiment, the call voice processing method may further include a step of checking whether the user has subscribed to a specified service, and a step of determining whether or not to save log information related to the call based on whether the user has subscribed to the specified service.
[0031] According to one embodiment, the call audio processing method may further include transmitting log information regarding the call to an external electronic device in response to receiving the second user input.
[0032] According to one embodiment, the call voice processing method may further include a step of checking whether a user of the external electronic device has subscribed to a predetermined service, and a step of determining whether or not to send log information related to the call based on whether the user of the external electronic device has subscribed to the predetermined service.
[0033] According to one embodiment, the step of transmitting call log information to an external electronic device may include transmitting a portion of the call log information, or call log information converted into another format, to the external electronic device if the user of the external electronic device does not subscribe to the specified service.
[0034] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the accompanying drawings. However, in the following description, specific descriptions of well-known functions, configurations, etc. may be omitted if they may unnecessarily obscure the gist of the present disclosure.
[0035] In the accompanying drawings, the same or corresponding components are denoted by the same reference numerals. In addition, in the following description of the embodiments, duplicated descriptions of the same or corresponding components may be omitted. However, the omission of a description of a component does not mean that the component is not included in a certain embodiment.
[0036] The advantages and features of the disclosed embodiments, as well as methods for achieving them, will become apparent from the following examples, which are provided in conjunction with the accompanying drawings. However, the present disclosure is not limited to the examples disclosed below, and may be realized in various different forms. These examples are provided merely to complete the disclosure and to enable those skilled in the art to fully understand the scope of the invention.
[0037] The terms used in this specification will be briefly explained, and the disclosed embodiments will be specifically described. The terms used in this specification are currently commonly used and general terms that have been selected as much as possible while taking into consideration the function of the present disclosure. However, these terms may differ depending on the intentions of engineers in the relevant field, legal precedents, the emergence of new technologies, etc. In addition, in certain cases, the applicant may arbitrarily select terms, and in such cases, their meanings will be described in detail in the description of the invention. Therefore, the terms used in this disclosure should be defined not simply by their names, but based on the meanings of the terms and the overall content of this disclosure.
[0038] In this specification, the singular includes the plural unless the context clearly dictates otherwise, and the plural includes the singular unless the context clearly dictates otherwise. When the entire specification is described as including a certain element, it does not mean that other elements are excluded, but that other elements may also be included, unless specifically stated to the contrary.
[0039] Furthermore, the terms "module" or "module" as used herein refer to a software or hardware component, and the "module" or "module" performs a certain function. However, the meaning of "module" or "module" is not limited to software or hardware. A "module" or "module" may reside on an addressable storage medium and be configured to be executed by one or more processors. Thus, by way of example, a "module" or "module" may include at least one of a software component, an object-oriented software component, a component such as a class component or task component, a process, a function, an attribute, a procedure, a subroutine, a segment of program code, a driver, firmware, microcode, a circuit, data, a database, a data structure, a table, an array, or a variable.
[0040] According to one embodiment of the present disclosure, a "module" or a "unit" may be implemented with a processor and memory. "Processor" should be interpreted broadly to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, etc. In some environments, "processor" may refer to an ASIC, a PLD, an FPGA, etc. "Processor" may also refer to a combination of processing devices, such as a combination of a DSP and a microprocessor, a combination of multiple microprocessors, or a combination of one or more microprocessors coupled with a DSP core. Furthermore, "memory" should be interpreted broadly to include any electronic component capable of storing electronic information. "Memory" may refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), nonvolatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or marking data storage devices, registers, etc.
[0041] Furthermore, terms such as first, second, A, B, (a), (b), etc. used in the following embodiments are used only to distinguish one component from another, and do not limit the essence or order of the components.
[0042] Furthermore, in the following embodiments, when a component is described as being "connected," "coupled," or "connected" to another component, it should be understood that the component may be directly connected or coupled to the other component, but that another component may be "connected," "coupled," or "connected" between each component.
[0043] Furthermore, the terms "comprises" and / or "comprising" used in the following embodiments do not preclude the presence or addition of one or more other components, steps, operations and / or elements to a listed component, step, operation and / or element.
[0044] Various embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings.
[0045] FIG. 1 is a diagram illustrating an electronic device 100 that processes voice in a call according to an embodiment of the present disclosure. Referring to FIG. 1, the electronic device 100 may perform a call function. For example, the electronic device 100 may process voice of a user 112 received through a microphone and transmit the processed voice data to the electronic device 102 of the other party 114, and may receive processed voice data of the other party 114 from the electronic device 102 of the other party 114 through a communication module and output the processed voice data to a speaker. Here, the call may include a voice call in which real-time voice is transmitted and received, and a video call in which video is transmitted and received along with voice, and may include calls based on various protocols such as VoIP.
[0046] According to one embodiment, the electronic device 100 may convert real-time call audio into text during a call and display the converted text 122 on the screen 120 together with information 124 about the speaker. In this way, the electronic device 100 may provide real-time call content not only as audio information but also as visual information. Furthermore, the electronic device 100 may support a user in clearly understanding the flow of conversation according to the call content by displaying the text 122 converted from the call audio together with the information 124 about the speaker. Here, the text 122 converted from the call audio and the information 124 about the speaker may be displayed on the screen 120 in an interactive format between the call participants. For example, the electronic device 100 may display the call content in an interactive format supported by a chat application or the like.
[0047] According to an embodiment, in response to receiving a user input, the electronic device 100 may display on the screen 120 text converted from a call voice from a predetermined time ago and information about the speaker. For example, the electronic device 100 may support the user in reviewing past call content. Here, the user input may include at least one of an input to select an object (e.g., a button object) displayed on the screen 120 or an input to scroll (or swipe) in a predetermined direction on the screen 120. In this way, the electronic device 100 may support the user in easily reviewing past call content that is pushed out and disappears due to the size limitation of a predetermined area of the screen 120 (e.g., the display is closed after being displayed by other call content). For example, if an object displayed on the screen 120 is a button object that can return to a time before a predetermined time, the user can review the call content from a time before the predetermined time by selecting the object. As another example, if the user input is a scroll input, the electronic device 100 may change the time point of the call content to be displayed depending on the scroll direction. For example, when receiving an input to scroll in a first direction (e.g., upward) on screen 120, electronic device 100 may display on screen 120 text converted from the voice of a call at a time before the currently displayed time and information about the speaker, and when receiving an input to scroll in a second direction (e.g., downward) on screen 120, electronic device 100 may display on screen 120 text converted from the voice of a call at a time after the currently displayed time and information about the speaker.
[0048] According to an embodiment, when switching screens, the electronic device 100 may display text converted from the call voice and information about the speaker on the switched screen. For example, the electronic device 100 may support a user in easily checking the contents of a call even in a multitasking environment. Here, the switched screen may be an execution screen of an application other than the call application or another execution screen of the call application. For example, if the switched screen is an execution screen of an application other than the call application, the electronic device 100 may display the text converted from the call voice and information about the speaker on the switched screen in a picture-in-picture (PIP) format. For another example, if the switched screen is an execution screen of a chat application linked to the call application, the electronic device 100 may display the text converted from the call voice and information about the speaker on the switched screen in an interactive format supported by the chat application. As yet another example, if the screen after switching is another execution screen of the call application (when switching from a first execution screen of the call application to a second execution screen of the call application), the electronic device 100 may display text converted from the call voice and information about the speaker in an interactive format between the call participants on the screen after switching.
[0049] 2 is a schematic diagram illustrating a configuration in which an information processing system 230 is communicatively connected to multiple user terminals 210_1, 210_2, and 210_3 for data processing according to an embodiment of the present disclosure. The information processing system 230 may include a system that can provide a data processing service (e.g., a telephone voice processing infrastructure service). In one embodiment, the information processing system 230 may include one or more server devices and / or databases that can store, provide, and execute computer-executable programs (e.g., downloadable applications) and data related to the data processing service, or one or more distributed computing devices and / or distributed databases of a cloud computing service infrastructure. For example, the information processing system 230 may include a separate system (e.g., a server) for the data processing service.
[0050] The data processing services provided by the information processing system 230 can be provided to users via data processing applications, web browser applications, etc. installed on each of the multiple user terminals 210_1, 210_2, and 210_3.
[0051] A plurality of user terminals 210_1, 210_2, and 210_3 may communicate with the information processing system 230 via a network 220. The network 220 may be configured to enable communication between the plurality of user terminals 210_1, 210_2, and 210_3 and the information processing system 230. Depending on the installation environment, the network 220 may be configured as, for example, a wired network such as Ethernet, a wired home network (Power Line Communication), a telephone line communication device, and RS-serial communication, a mobile communication network, a wireless network such as WLAN (Wireless LAN), Wi-Fi, Bluetooth, and ZigBee, or a combination thereof. The communication method is not limited to, and may include not only a communication method utilizing a communication network that the network 220 may include (e.g., a mobile communication network, a wired Internet, a wireless Internet, a broadcast network, a satellite network, etc.), but also short-range wireless communication between the user terminals (210_1, 210_2, and 210_3).
[0052] For example, a plurality of user terminals 210_1, 210_2, 210_3 may transmit data processing requests, instructions related to user requests for data processing, to the information processing system 230 via the network 220, and the information processing system 230 may receive them.
[0053] 2 illustrates a mobile phone terminal 210_1, a tablet terminal 210_2, and a PC terminal 210_3 as examples of user terminals, but is not limited thereto, and the user terminals 210_1, 210_2, and 210_3 may be any computing devices capable of wired and / or wireless communication and capable of installing and executing a data processing application or the like. For example, the user terminals may include smartphones, mobile phones, navigation systems, computers, notebooks, digital broadcasting terminals, PDAs, PMPs, tablet PCs, game consoles, wearable devices, Internet of Things (IoT) devices, virtual reality (VR) devices, augmented reality (AR) devices, etc. Also, while FIG. 2 illustrates three user terminals 210_1, 210_2, and 210_3 communicating with the information processing system 230 via the network 220, the present invention is not limited thereto, and a different number of user terminals may be configured to communicate with the information processing system 230 via the network 220.
[0054] FIG. 3 is a block diagram illustrating the internal configuration of a user terminal 210 and an information processing system 230 according to an embodiment of the present disclosure. The user terminal 210 refers to any computing device capable of executing data processing applications and capable of wired / wireless communication, and may include, for example, the mobile phone terminal 210_1, tablet terminal 210_2, and PC terminal 210_3 of FIG. 2 . As illustrated, the user terminal 210 may include a memory 312, a processor 314, a communication module 316, and an input / output interface 318. Similarly, the information processing system 230 may include a memory 332, a processor 334, a communication module 336, and an input / output interface 338. As illustrated in FIG. 3 , the user terminal 210 and the information processing system 230 may be configured to communicate information and / or data over the network 220 using the respective communication modules 316 and 336. Furthermore, the input / output device 320 may be configured to input information and / or data to the user terminal 210 via the input / output interface 318, or to output information and / or data generated by the user terminal 210.
[0055] The memory 312, 332 may include any non-transitory computer-readable recording medium. According to one embodiment, the memory 312, 332 may include a non-volatile mass storage device such as a read-only memory (ROM), a disk drive, a solid-state drive (SSD), or a flash memory. As another example, a non-volatile mass storage device such as a ROM, an SSD, a flash memory, or a disk drive may be included in the user terminal 210 or the information processing system 230 as a separate permanent storage device distinct from the memory. The memory 312, 332 may also store an operating system and at least one program code (e.g., code for an application related to a data processing service, etc.).
[0056] These software components may be loaded from a computer-readable recording medium separate from the memories 312, 332. Such separate computer-readable recording medium may include a recording medium directly connectable to the user terminal 210 and the information processing system 230, such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, or a memory card. As another example, the software components may be loaded into the memories 312, 332 via the communication modules 316, 336 rather than from a computer-readable recording medium. For example, at least one program may be loaded into the memories 312, 332 based on a computer program (e.g., an application related to a data processing service) to be installed by a file provided via the network 220 by a developer or a file distribution system that distributes application installation files.
[0057] The processors 314, 334 may be configured to process computer program instructions by performing basic arithmetic, logic, and input / output operations. The instructions may be provided to the processors 314, 334 by the memories 312, 332 or the communications modules 316, 336. For example, the processors 314, 334 may be configured to execute instructions that they receive according to program code stored in a storage device, such as the memories 312, 332.
[0058] The communication modules 316, 336 may provide configurations or functions for the user terminal 210 and the information processing system 230 to communicate with each other via the network 220, and may provide configurations or functions for the user terminal 210 and / or the information processing system 230 to communicate with other user terminals or other systems (e.g., a separate cloud system, etc.). For example, a request or data (e.g., a data processing request or data, etc.) generated by the processor 314 of the user terminal 210 in accordance with program code stored in a storage device such as the memory 312 may be transmitted to the information processing system 230 via the network 220 under the control of the communication module 316. Conversely, a control signal or command provided under the control of the processor 334 of the information processing system 230 may be received by the user terminal 210 via the communication module 316 of the user terminal 210 via the communication module 336 and the network 220.
[0059] The input / output interface 318 may be a means for interfacing with the input / output device 320. As an example, the input device may include a device such as a camera including an audio sensor and / or an image sensor, a keyboard, a microphone, or a mouse, while the output device may include a device such as a display, a speaker, or a haptic feedback device. As another example, the input / output interface 318 may be a means for interfacing with a device that has an input and output configuration or function integrated therein, such as a touchscreen. While the input / output device 320 is not shown in FIG. 3 as being included in the user terminal 210, this is not a limitation and the device may be configured as an integrated device with the user terminal 210. Furthermore, the input / output interface 338 of the information processing system 230 may be a means for interfacing with an input or output device (not shown) that may be connected to the information processing system 230 or included in the information processing system 230. While the input / output interfaces 318 and 338 are shown in FIG. 3 as separate elements from the processors 314 and 334, this is not a limitation and the input / output interfaces 318 and 338 may be configured as being included in the processors 314 and 334.
[0060] The user terminal 210 and the information processing system 230 may include more components than those shown in FIG. 3 . However, it is not necessary to explicitly illustrate most of the conventional components. In one embodiment, the user terminal 210 may be implemented to include at least a portion of the input / output device 320 described above. The user terminal 210 may also include other components such as a transceiver, a GPS (Global Positioning System) module, a camera, various sensors, a database, etc. For example, if the user terminal 210 is a smartphone, it may include components typically found in smartphones, such as an acceleration sensor, a gyro sensor, a microphone module, a camera module, various physical buttons, buttons using a touch panel, an input / output port, a vibrator, etc.
[0061] According to one embodiment, the processor 314 of the user terminal 210 may be configured to run a data processing application or a web browser application that provides data processing services. At this time, program code associated with the application may be loaded into the memory 312 of the user terminal 210. While the application is running, the processor 314 of the user terminal 210 may receive information and / or data provided by the input / output device 320 via the input / output interface 318 or from the information processing system 230 via the communication module 316, process the received information and / or data, and store it in the memory 312. The information and / or data may also be provided to the information processing system 230 via the communication module 316.
[0062] While the data processing application is running, the processor 314 may receive voice data, text, images, video, etc., entered or selected via an input device such as a touch screen, keyboard, camera including an audio sensor and / or image sensor, microphone, etc., connected to the input / output interface 318, and store the received voice data, text, images, and / or video, etc., in the memory 312 or provide the received voice data, text, images, and / or video, etc., to the information processing system 230 via the communication module 316 and the network 220. In one embodiment, the processor 314 may receive user input entered via the input device and provide data / requests corresponding to the received user input to the information processing system 230 via the network 220 and the communication module 316.
[0063] The processor 314 of the user terminal 210 may transmit and output information and / or data to an input / output device 320 via an input / output interface 318. For example, the processor 314 of the user terminal 210 may output processed information and / or data through an output device 320, such as a display-capable device (e.g., a touchscreen, a display, etc.) or an audio-capable device (e.g., a speaker).
[0064] The processor 334 of the information processing system 230 may be configured to manage, process, and / or store information and / or data received from the plurality of user terminals 210 and / or the plurality of external systems. The information and / or data processed by the processor 334 may be provided to the user terminal 210 via the communication module 336 and the network 220.
[0065] 4 is a diagram illustrating a configuration of an electronic device 400 for processing telephone call audio according to one embodiment of the present disclosure. Referring to FIG. 4, the electronic device 400 for processing telephone call audio (e.g., the electronic device 100 in FIG. 1 may include a processor 410), a display 420, a memory 430, and a communication module 440. However, the configuration of the electronic device 400 is not limited thereto. According to various embodiments, the electronic device 400 may omit at least one of the above components and may further include at least one other component.
[0066] The processor 410 may execute software (or programs) to control at least one other component (e.g., a hardware or software component) of the electronic device 400 connected to the processor 410, and may perform various data processing or calculations. According to one embodiment, as at least part of the data processing or calculations, the processor 410 may load instructions or data received from another component (e.g., the communications module 440) into volatile memory, process the instructions or data stored in the volatile memory, and store the resulting data in non-volatile memory.
[0067] Display 420 may provide visual information. According to one embodiment, display 420 may display text converted from speech and information about the speaker. Display 420 may include, for example, a touch sensor configured to detect a touch or a pressure sensor configured to measure the amount of force generated by a touch, and may include sensor circuitry or control circuitry for controlling the sensor.
[0068] The memory 430 may store various data used by at least one component (e.g., the processor 410) of the electronic device 400. The data may include, for example, input data or output data for software (or programs) and associated instructions. The memory 430 may include volatile or non-volatile memory. According to one embodiment, the memory 430 may store a call log.
[0069] The memory 430 may include at least one instruction related to processing the call audio. The at least one instruction may include, for example, instructions related to acquiring the call audio, converting the call audio to text, identifying the speaker of the call audio, displaying the call content, and saving a call log. The memory 430 may include a call audio acquisition module 431, a text conversion module 433, a speaker identification module 435, a call content display module 437, and a call log saving module 439. However, the types of modules included in the memory 430 are categorized according to the function of the instructions, and the types and number of modules are not limited thereto. Furthermore, the modules included in the memory 430 (or the instructions included in the modules) are executed by the processor 410 and may be implemented by the processor 410 itself.
[0070] The call audio acquisition module 431 may acquire call audio. For example, the call audio acquisition module 431 may acquire user voice data through a microphone. The call audio acquisition module 431 may also acquire voice data of the other party from an external electronic device through the communication module 440.
[0071] The text conversion module 433 may convert the speech acquired via the speech audio acquisition module 431 into text. For example, the text conversion module 433 may perform a STT (Speech-To-Text) function. The text conversion module 433 may remove noise from the acquired speech data using a filter and divide the continuous speech data into multiple frames. The text conversion module 433 may then extract features from each divided frame and convert the speech data into text via a speech recognition model based on the extracted features. Here, the speech recognition model may include, for example, an acoustic model that analyzes and classifies the speech data into phonemes, syllables, etc., and a language model that converts the phonemes and syllables extracted from the speech data into words or sentences.
[0072] The speaker identification module 435 may identify the speaker of the acquired audio. According to one embodiment, the speaker identification module 435 may identify the speaker of the audio based on the source from which the audio was acquired. For example, if the audio is acquired through a microphone of the electronic device 400, the speaker identification module 435 may identify the speaker of the acquired audio as the user of the electronic device 400. As another example, if the audio is acquired from an external electronic device via the communication module 440 of the electronic device 400, the speaker identification module 435 may identify the speaker of the acquired audio as the user of the external electronic device, i.e., the other party of the call. According to one embodiment, the speaker identification module 435 may identify the speaker through voiceprint recognition of the audio. For example, the speaker identification module 435 may analyze the audio to extract unique features (e.g., frequency spectrum) of the audio and use the extracted unique features to identify the speaker of the audio. According to an embodiment, the speaker identification module 435 may identify a speaker of the call voice based on at least one of a speech recognition result of the call voice or speaker-related information included in a text converted from the call voice, where the speaker-related information may include information that can identify the speaker (e.g., the speaker's name, the speaker's nickname, or a name for the speaker).
[0073] The call content display module 437 may display the call content during a call on the display 420. For example, the call content display module 437 may display text converted from the call voice on the screen of the display 420 during a call.
[0074] According to one embodiment, the call content display module 437 may display text converted from the real-time call voice on a screen together with information about the speaker, and may display at least one of the current time or the call duration (e.g., the duration of the call) on a screen together with the text converted from the call voice and the information about the speaker.
[0075] According to one embodiment, in response to receiving a first user input, the call content display module 437 may display on a screen a first text obtained by converting a call voice at a first time point, which is a predetermined time (e.g., 10 seconds) before the current time point, together with information about the speaker of the call voice at the first time point. In this case, the call content display module 437 may display on a screen information about the first time point together with the first text and the information about the speaker of the call voice at the first time point. Here, the information about the first time point may include at least one of time information about the first time point or information about the time elapsed from the start of the call to the first time point. In addition, the first user input may include at least one of an input to select an object (e.g., a button object) displayed on the screen or an input to scroll / swipe the screen in a first direction (e.g., scroll up).
[0076] According to one embodiment, in response to receiving a second user input, the call content display module 437 may display on the screen a second text obtained by converting a call voice from a second time point after the first time point, along with information about the speaker of the call voice from the second time point. For example, when the call content display module 437 receives a second user input while displaying the call content from a previous time point (first time point), the call content display module 437 may display the call content from the second time point, which is a predetermined time after the current time point. Here, the second user input may include at least one of an input to select an object (e.g., a button object) displayed on the screen or an input to scroll (e.g., scroll downward) / swipe on the screen in a second direction opposite to the first direction. This allows the call content display module 437 to change the time point of the call content to be displayed based on the scrolling direction.
[0077] According to one embodiment, after displaying the first text on the screen, the call content display module 437 may display text converted from the current real-time call voice and information about the speaker on the screen when a predetermined time has elapsed or a third user input has been received. For example, the call content display module 437 may return to a screen state displaying the real-time call content when a predetermined time has elapsed or a third user input has been received while displaying the past call content. Here, the third user input may include at least one of an input for selecting an object (e.g., a button object) displayed on the screen or an input for scrolling to an end point of a second direction (e.g., downward) on the screen.
[0078] According to one embodiment, when real-time call voice is received while displaying the first text and information about the speaker of the call voice at a first time point in a first region of the screen, the call content display module 437 may display text converted from the real-time call voice and information about the speaker of the real-time call voice in a second region of the screen different from the first region. For example, when real-time call voice is received while displaying the content of a past call, the real-time call voice may be displayed in the second region while maintaining the display state of the content of the past call. Here, the second region may be a region fixed below the first region.
[0079] According to one embodiment, when switching screens, the call content display module 437 may display text converted from real-time call voice on the switched screen together with information about the speaker. In this case, the call content display module 437 may display at least one of the current time or the call duration (e.g., the duration of the call) on the switched screen together with the text converted from the call voice and information about the speaker. Herein, the switched screen may include at least one of an execution screen of an application other than the call application, an execution screen of a chat application linked to the call application, or another execution screen of the call application. Hereinafter, for convenience of explanation, the screen before switching will be referred to as the first screen, and the switched screen will be referred to as the second screen.
[0080] According to an embodiment, when the first screen is an execution screen of a call application and the second screen is an execution screen of an application other than the call application, the call content display module 437 may display text converted from the call voice and information about the speaker in a PIP format on the second screen. In this case, the PIP area set in the second screen may have a size sufficient to include a predetermined number of texts converted from the call voice.
[0081] According to an embodiment, when the first screen is an execution screen of a calling application and the second screen is an execution screen of a chat application linked to the calling application, the call content display module 437 may display text converted from the call voice on the second screen in an interactive format supported by the chat application based on information about the speaker of the call voice. In this case, the call content display module 437 may display an object (e.g., an image object) indicating the call status on the second screen to distinguish from when the chat application displays chat content, i.e., to indicate that a call is in progress rather than a chat. In one embodiment, the call content display module 437 may support the execution of a function of the chat application while interactively displaying the text converted from the call voice on the second screen. For example, the call content display module 437 may interactively display the text converted from the call voice on the second screen, while also displaying text entered via the chat application (e.g., chat input text) on the second screen in chronological order. In other words, the electronic device 400 may support the user to simultaneously talk and chat with the other party. Here, we have described the case where the second screen is a chat application that works in conjunction with a calling application, but this is not limited to this, and the second screen can be linked to the calling application and can include an execution screen of an SNS application, a message application, etc. that supports an interactive format.
[0082] According to an embodiment, when the first screen is a first execution screen of the calling application and the second screen is a second execution screen of the calling application, the call content display module 437 may display text converted from the call voice on the second screen in an interactive format between the call participants based on information about the speaker of the call voice. In this case, the call content display module 437 may display an object (e.g., an image object) indicating the call status on the second screen to indicate that a call is in progress. Here, the first execution screen of the calling application may be an initial execution screen or a main screen of the calling application, and the second execution screen of the calling application may be a screen switched from the first execution screen in response to receiving a user input, and the area where the text converted from the call voice is displayed may occupy most of the screen.
[0083] According to one embodiment, the call content display module 437 may activate a vibration actuator or provide a visual effect on the second screen when the text converted from the call voice includes predetermined text. For example, the call content display module 437 may vibrate the electronic device 400 or provide a visual effect on the second screen when predetermined text is detected in the call content. According to one embodiment, the predetermined text may include information related to the user. The information related to the user may include information that can identify the user (e.g., the user's name, the user's nickname, or a nickname for the user).
[0084] According to one embodiment, when the call voice is identified as a sound corresponding to a predetermined pattern, the call content display module 437 may display an image mapped to the predetermined pattern on a screen together with information about the speaker of the call voice. Here, the sound corresponding to the predetermined pattern may include at least one of a sound generated by the speaker's actions (e.g., laughter, coughing, etc.) or a sound generated by an external object (e.g., a car sound, music, etc.). In addition, the image mapped to the predetermined pattern (e.g., a sticker image) may include at least one of an image with a shape similar to the speaker's actions or an image with a shape similar to an external object.
[0085] According to one embodiment, the call content display module 437 can display information about the time of the real-time call audio on at least one of the first screen and the second screen, along with text converted from the real-time call audio and information about the speaker of the real-time call audio. Here, the information about the time of the real-time call audio may include at least one of time information when the real-time call audio was received (e.g., current time information) or information about the time elapsed from the start of the call to the time when the real-time call audio was received (e.g., the current time). According to one embodiment, the information about the time of the real-time call audio displayed on the first screen and the information about the time of the real-time call audio displayed on the second screen may take different forms. For example, if the information about the time of the real-time call audio displayed on the first screen includes current time information, the information about the time of the real-time call audio displayed on the second screen may include time information elapsed from the start of the call to the current time. For another example, if the information about the time of the real-time call audio displayed on the first screen includes time information elapsed from the start of the call to the current time, the information about the time of the real-time call audio displayed on the second screen may include current time information.
[0086] The call log storage module 439 may store the call logs. According to one embodiment, the call log storage module 439 may perform general processing of the call logs. For example, the call log storage module 439 may perform processing related to storing the call logs, viewing the call logs, sharing the call logs, and modifying the call logs.
[0087] According to one embodiment, the call log storage module 439 may store log information related to the call based on the text converted from the real-time call voice and information about the speaker of the real-time call voice. For example, the call log storage module 439 may store the text converted from the real-time call voice and information about the speaker of the real-time call voice in the memory 430 immediately after the call ends or in response to receiving a fourth user input. In this case, the call log storage module 439 may store information about the time point of the real-time call voice together with the text converted from the real-time call voice and information about the speaker of the real-time call voice. Here, the information about the time point of the real-time call voice may include at least one of information about the time when the real-time call voice was received (information about the current time at the time of the call) or information about the time elapsed from the start of the call to the time when the real-time call voice was received (information about the current time at the time of the call).
[0088] According to one embodiment, the call log storage module 439 can display log information related to a call on a screen. For example, the call log storage module 439 may display log information related to a call on a third screen that is separate from the first and second screens. According to one embodiment, the call log storage module 439 may display log information related to a call on the third screen in response to receiving a fourth user input in which an object (e.g., a button object) displayed on the fourth screen that is opened immediately after the call ends is selected.
[0089] According to one embodiment, when a fifth user input is received in which specific text displayed on the first screen is selected, the call log storage module 439 may display (e.g., highlight) the specific text differently from other text in the call-related log information displayed on the third screen. For example, when specific text is selected from text converted from real-time call audio, the call log storage module 439 may store the selected text differently from other text when saving the call-related log information, and based on the result, may display the selected text differently from other text when displaying the call-related log information.
[0090] According to one embodiment, the call log storage module 439 may check whether the user subscribes to a predetermined service (e.g., a premium service) and determine whether to store call log information based on whether the user subscribes to the predetermined service. For example, the call log storage module 439 may store call log information only if the user subscribes to the predetermined service.
[0091] According to one embodiment, the call log storage module 439 may transmit log information about the call to an external electronic device in response to receiving a sixth user input. For example, in response to receiving a sixth user input in which a button object supporting sharing of a call log is selected, the call log storage module 439 may share log information about the call with the other party's electronic device or another user's electronic device.
[0092] According to one embodiment, the call log storage module 439 may check whether the user of the external electronic device subscribes to a predetermined service (e.g., a premium service) and determine whether to send call log information based on whether the user of the external electronic device subscribes to the predetermined service. For example, the call log storage module 439 may restrict sharing of call log information if the user of the external electronic device does not subscribe to the predetermined service. For another example, if the user of the external electronic device does not subscribe to the predetermined service, the call log storage module 439 may transmit to the external electronic device a portion of the call log information (e.g., restricted information) or information converted from the call log information into a different format (e.g., an image captured from the log information).
[0093] According to one embodiment, the call log storage module 439 may modify at least a portion of the call log information displayed on the third screen in response to receiving a seventh user input, such as a user modifying text included in the call log information.
[0094] According to one embodiment, when a call is a multi-party call, the call log storage module 439 may search for the speech content of a specific user from the log information related to the multi-party call by filtering. As an example, in response to receiving an eighth user input, the call log storage module 439 may search for the speech content of a specific user from the log information related to the multi-party call displayed on the third screen and display the speech content differently from the speech content of other users. As another example, in response to receiving an eighth user input, the call log storage module 439 may search for the speech content of a specific user from the log information related to the multi-party call and display only the speech content of the searched specific user on the third screen.
[0095] The communication module 440 (or communication circuitry) may support the establishment of a direct (e.g., wired) or wireless communication channel between the electronic device 400 and an external electronic device (e.g., the electronic device 102 of FIG. 1 ) and the execution of communication via the established communication channel. According to one embodiment, the electronic device 400 may receive voice data of a call partner from the external electronic device via the communication module 440.
[0096] 5 is a diagram illustrating a method for displaying text 534 converted from a call voice and information 532 about the speaker on an execution screen 500 of a call application according to an embodiment of the present disclosure. Referring to FIG. 5, a processor (e.g., processor 410 in FIG. 4) of an electronic device (e.g., electronic device 100 in FIG. 1 or electronic device 400 in FIG. 4) that processes the call voice can display text 534 converted from the call voice and information 532 about the speaker of the call voice on the execution screen 500 of the call application. For example, the processor can display the contents of the call on the screen during a call.
[0097] The execution screen 500 of the calling application may include information about the call party 510, a call duration 520, and an object supporting call termination 550. The information about the call party 510 may include at least one of an image 512 representing the call party or the name of the call party 514. The call duration 520 may include information about the time elapsed since the start of the call. The object supporting call termination 550, when selected by user input, may send a signal to the processor to terminate the call.
[0098] According to one embodiment, the execution screen 500 of the calling application may include a first area 530 in which text 534 converted from the call voice and information about the speaker of the call voice 532 are displayed. The first area 530 may be set to a predetermined position and a predetermined size on the execution screen 500 of the calling application.
[0099] According to one embodiment, the first area 530 may display text converted from real-time speech and text converted from speech at a time point a predetermined time before the current time point (hereinafter referred to as the "first time point"), along with information about the speaker of the speech. To display the text converted from the speech at the first time point, the processor may receive a user input generated in the first area 530. Here, the user input may be an input to select an object 542 (e.g., a button object) displayed in the first area 530, or an input to select a predetermined object 542 in the first area 530. The inputs may include at least one of inputs to scroll 544 (or swipe) in a predetermined direction in the first area 530. As one example, when an object 542 displayed in the first area 530 is selected, the processor may display text converted from the call voice at the first time point and information about the speaker of the call voice at the first time point in the first area 530. As another example, when an input to scroll 544 (or swipe) in a predetermined direction in the first area 530 is received, the processor may display text converted from the call voice at the first time point and information about the speaker of the call voice at the first time point in the first area 530.
[0100] According to one embodiment, the call voice may be converted into text based on at least one of information related to the speaker of the call voice or information related to the recipient of the call voice. Here, the information related to the speaker of the call voice is related to the language set for the speaker of the call voice, and may include at least one of the nationality of the speaker of the call voice, the language used by the speaker of the call voice, or the language set for the electronic device used by the speaker (e.g., electronic device 100 of FIG. 1 ). Also, the information related to the recipient of the call voice is related to the language set for the recipient of the call voice, and may include at least one of the nationality of the recipient of the call voice, the language used by the recipient of the call voice, or the language set for the electronic device used by the recipient of the call voice (e.g., electronic device 102 of FIG. 1 ). As one example, the call voice may be converted into text based on the language set for the speaker of the call voice (or the text converted from the call voice may be translated into the language set for the speaker of the call voice). As another example, the call voice may be converted into text based on the language set for the recipient of the call voice (or the text converted from the call voice may be translated into the language set for the recipient of the call voice). As yet another example, if the language set for the speaker of the call voice and the language set for the recipient of the call voice are different, the call voice may be converted to text (or the text converted from the call voice may be translated into the language set for the user of the electronic device) based on the language set for the user of the electronic device (e.g., the language set for the speaker of the call voice or the language set for the recipient of the call voice).
[0101] According to one embodiment, text converted from the call voice may be translated into a predetermined language (e.g., a language set for the speaker of the call voice, a language set for the recipient of the call voice, or a language set by the system or a user), and a displayable translation function-related object may be displayed in first area 530 or an area adjacent to first area 530. For example, if the language set for the speaker of the call voice is different from the language set for the recipient of the call voice, the processor may display the translation function-related object in first area 530 or an area adjacent to first area 530. In this case, when the translation function-related object is selected based on user input, the processor may provide text converted from the call voice based on the predetermined language (e.g., the language set for the recipient of the call voice).
[0102] 6 is a diagram illustrating a method for processing real-time call audio in a state where text 624 converted from past call audio and information 622 about the speaker are displayed on an execution screen 600 of a call application according to an embodiment of the present disclosure. Referring to FIG. 6, a processor (e.g., processor 410 of FIG. 4 ) of an electronic device (e.g., electronic device 100 of FIG. 1 or electronic device 400 of FIG. 4 ) that processes call audio may display text 624 converted from call audio at a time point a predetermined time before the current time point (hereinafter referred to as a “first time point”) and information 622 about the speaker of the call audio at the first time point, as in a first state 602, on the execution screen 600 of the call application. For example, the processor may display the contents of past calls on the screen during a call.
[0103] According to one embodiment, in response to receiving a user input (hereinafter referred to as a "first user input"), the processor may display text 624 converted from the call voice at the first time point and information 622 about the speaker of the call voice at the first time point in a first region 620 of the screen (e.g., first region 530 in FIG. 5). For example, when an input is received to select a first object 610 (e.g., object 542 in FIG. 5) displayed in the first region 620 or an input to scroll in a first direction (e.g., upward) in the first region 620 (e.g., scroll 544 in FIG. 5), the execution screen 500 of the calling application shown in FIG. 5 may be changed to a first state 602 of the execution screen 600 of the calling application shown in FIG. 6. That is, the text converted from the real-time call voice displayed in the first area 620 (e.g., text 534 in FIG. 5) and information about the speaker of the real-time call voice (e.g., information about the speaker 532 in FIG. 5) are scrolled in chronological order from the current time to the first time point, and finally, the text 624 converted from the call voice at the first time point and the information 622 about the speaker of the call voice at the first time point are displayed in the first area 620.
[0104] According to one embodiment, in response to receiving a user input (hereinafter referred to as a second user input), the processor may display text converted from a call voice at a second time point after the first time point and information about the speaker of the call voice at the second time point in the first area 620 of the screen. For example, when the processor receives a second user input while displaying the call content at a past time point (first time point), the processor may display the call content at the second time point a predetermined time after the currently displayed time point (first time point). Here, the second user input may include at least one of an input to select a second object (e.g., a button object) displayed in the first area 620 or an input to scroll in a second direction (e.g., downward) opposite to the first direction on the first area 620. This allows the processor to change the time point of the call content to be displayed based on the scroll direction. The second object may be an object that supports returning to a time point (second time point) after a predetermined time. According to one embodiment, the processor may change the first object 610 to a second object upon receiving the first user input. According to another embodiment, the processor may display the first object 610 together with the second object in the first area 620 .
[0105] According to one embodiment, the processor may display text 624 converted from the call voice at a first time point in the first area 620, and then, after a predetermined time has elapsed or a user input (hereinafter referred to as a third user input) is received, display text converted from the current real-time call voice and information about the speaker in the first area 620. For example, when a predetermined time has elapsed or a third user input is received while displaying the content of a past call, the processor may return to a screen state displaying the content of the real-time call (e.g., the state of the execution screen 500 of the call application shown in FIG. 5). Here, the third user input may include at least one of an input to select a third object (e.g., a button object) displayed in the first area 620 or an input to scroll to the end of the second direction (e.g., downward) on the first area 620.
[0106] According to one embodiment, when real-time call audio is received in a state where text 624 converted from the call audio at a first time point and information 622 about the speaker of the call audio at the first time point are displayed in a first region 620 of the screen (e.g., first state 602), the processor may display text 634 converted from the real-time call audio and information 632 about the speaker of the real-time call audio in a second region 630 of the screen different from the first region 620, as in a second state 604. For example, when real-time call audio is received in a state where content of a past call is displayed, the processor may display the content of the real-time call in a predetermined region (e.g., second region 630) while maintaining the display state of the content of the past call. According to one embodiment, second region 630 may be a region fixed at the bottom of first region 620.
[0107] 7 is a diagram illustrating a method for displaying text 724, 726 converted from a call voice and information 722 about the speaker on another execution screen 700 of a call application according to an embodiment of the present disclosure, and FIG. 8 is a diagram illustrating a method for processing real-time call voice in a state where text 726 converted from a past call voice and information about the speaker are displayed on another execution screen 700 of the call application according to an embodiment of the present disclosure. With reference to FIG. 7 and FIG. 8, a processor (e.g., processor 410 of FIG. 4) of an electronic device (e.g., electronic device 100 of FIG. 1 or electronic device 400 of FIG. 4) that processes call voice can display text 724, 726, 744 converted from the real-time call voice on a switched screen together with information 722, 742 about the speaker when switching screens. At this time, the processor may display at least one of information 710 about the other party of the call, the current time, or the call duration 720 (e.g., the duration of the call), along with text 724, 726, 744 converted from the call voice and information 722, 742 about the speaker, on the screen after switching. Hereinafter, for convenience of explanation, the screen before screen switching will be referred to as the first screen, and the screen after screen switching will be referred to as the second screen 700. For example, the first screen may be a first execution screen of the call application (e.g., execution screen 500 of the call application in FIG. 5 or execution screen 600 of the call application in FIG. 6) and represents an initial execution screen or main screen of the call application, and the second screen 700 may be a second execution screen of the call application and may be a screen switched from the first execution screen in response to receiving a user input. According to one embodiment, the area where text 724, 726, and 744 converted from the call voice is displayed on the second screen 700 may occupy most of the screen.
[0108] According to one embodiment, the processor may display text 724, 726, 744 converted from the call voice on the second screen 700 in an interactive format between the call participants based on information 722, 742 about the speaker of the call voice. At this time, the processor may display an object (e.g., an image object) indicating the call status on the second screen 700 to indicate that the call is in progress. According to one embodiment, if the speaker of the call voice is the user of the electronic device, the processor may display only text 726 converted from the voice spoken by the user on the second screen 700, excluding information about the user.
[0109] According to one embodiment, in response to receiving a user input (hereinafter referred to as a “first user input”), the processor may display text converted from a call voice at a time point (hereinafter referred to as a “first time point”) prior to the current time point and information about the speaker of the call voice at the first time point on the second screen 700. For example, when an input to scroll 730 (e.g., scroll 544 in FIG. 5 ) in a first direction (e.g., upward) on a first object (e.g., an input or second screen 700 in which object 542 in FIG. 5 or object 610 in FIG. 6 is selected) displayed on the second screen 700 is received, text 724, 726 converted from the real-time call voice and information 722 about the speaker of the real-time call voice displayed on the second screen 700 are scrolled in chronological order from the current time point to the first time point, and ultimately, the text converted from the call voice at the first time point and information about the speaker of the call voice at the first time point may be displayed on the second screen 700.
[0110] According to one embodiment, in response to receiving a user input (hereinafter referred to as a "second user input"), the processor may display on the second screen 700 text converted from a call voice at a second time point that is later than the first time point and information about the speaker of the call voice at the second time point. For example, when the processor receives a second user input while displaying the call content at a past time point (first time point), the processor may display the call content at the second time point that is a predetermined time after the time point currently being displayed (first time point). Here, the second user input may include at least one of an input to select a second object (e.g., a button object) displayed on the second screen 700 or an input to scroll the second screen 700 in a second direction (e.g., downward) opposite to the first direction. This allows the processor to change the time point of the call content to be displayed based on the scroll direction.
[0111] According to one embodiment, the processor displays text converted from the call voice at a first time point on the second screen 700, and then, after a predetermined time has elapsed or a user input (hereinafter referred to as a "third user input") has been received, can display text 724, 726 converted from the current real-time call voice and information 722 about the speaker on the second screen 700. For example, when a predetermined time has elapsed or a third user input has been received while displaying the content of a past call, the processor can return to a screen state displaying the content of the real-time call. Here, the third user input can include at least one of an input to select a third object (e.g., a button object) displayed on the second screen 700 or an input to scroll to the end of the second direction (e.g., downward) on the second screen 700.
[0112] According to one embodiment, when real-time call voice is received while the processor is displaying text converted from the call voice at a first time point and information about the speaker of the call voice at the first time point on the second screen 700, the processor may display text converted from the real-time call voice 744 and information about the speaker of the real-time call voice 742 in a predetermined area 740 of the second screen 700. For example, when the processor receives real-time call voice while displaying content of a past call, the processor may display the content of the real-time call in the predetermined area 740 while maintaining the display state of the content of the past call. According to one embodiment, the predetermined area 740 may be a fixed area at the bottom of the second screen 700.
[0113] FIG. 9 is a diagram illustrating a method for displaying text 914 converted from a call voice in a PIP format and information 912 about the speaker when an execution screen 900 of an application other than the call application is switched to, according to an embodiment of the present disclosure. FIG. 10 is a diagram illustrating a method for changing the size of a PIP area 910 according to an embodiment of the present disclosure. With reference to FIGS. 9 and 10 , a processor (e.g., processor 410 of FIG. 4 ) of an electronic device (e.g., electronic device 100 of FIG. 1 or electronic device 400 of FIG. 4 ) that processes the call voice can display text 914 converted from the real-time call voice together with information 912 about the speaker on a switched screen when switching screens. At this time, the processor can display at least one of the current time or the duration of the call (e.g., the duration of the call) on the switched screen together with the text 914 converted from the call voice and the information 912 about the speaker. Here, the switched screen may include an execution screen 900 of an application other than the call application (e.g., a game application). For ease of explanation, the screen before the screen is switched will be referred to as the first screen, and the screen after the screen is switched will be referred to as the second screen 900. For example, the first screen may be an execution screen of a calling application (for example, the execution screen 500 of the calling application in FIG. 5 or the execution screen 600 of the calling application in FIG. 6), and the second screen 900 may be an execution screen of an application other than the calling application.
[0114] According to one embodiment, the processor may overlay and display text 914 converted from the call voice and information 912 about the speaker in a PIP format on the second screen 900. In this case, the PIP area 910 set on the second screen 900 may have a size that can include a predetermined number or more of the text 914 converted from the call voice. In addition, the processor may display an object 920 (e.g., an image object) indicating the call status on the second screen 900 to indicate that a call is in progress. For example, the object 920 indicating the call status may be displayed in the PIP area 910.
[0115] According to one embodiment, the processor may display resizable objects 932, 934 in the PIP region 910. As one example, when a first object 932 displayed in the PIP region 910 is selected, the processor may increase the size of the PIP region 910, as shown in FIG. 10. As another example, when a second object 934 displayed in the PIP region 910 is selected, the processor may decrease the size of the PIP region 910, as shown in FIG. 9. According to one embodiment, the processor may change the first object 932 to the second object 934 in response to receiving user input selecting the first object 932, and change the second object 934 to the first object 932 in response to receiving user input selecting the second object 934. For example, the processor may toggle between displaying resizable objects 932, 934 in the PIP region 910.
[0116] According to one embodiment, in response to receiving a user input (hereinafter referred to as a “first user input”), the processor may display text converted from a call voice at a time point a predetermined time before the current time (hereinafter referred to as a “first time point”) and information about the speaker of the call voice at the first time point in PIP area 910. For example, when an input is received to select a first object (e.g., object 542 in FIG. 5 or object 610 in FIG. 6 ) displayed in PIP area 910 or to scroll in a first direction (e.g., upward) in PIP area 910 (e.g., scroll 544 in FIG. 5 ), text 914 converted from the real-time call voice and information 912 about the speaker of the real-time call voice, which are displayed in PIP area 910, may be scrolled in chronological order from the current time point to the first time point, and ultimately, the text converted from the call voice at the first time point and information about the speaker of the call voice at the first time point may be displayed in PIP area 910.
[0117] According to one embodiment, in response to receiving a user input (hereinafter referred to as a "second user input"), the processor may display, in the PIP area 910, text converted from the call audio at a second time point (later than the first time point) and information about the speaker of the call audio at the second time point. For example, when the processor receives a second user input while displaying the call audio at a past time point (first time point), the processor may display the call audio at the second time point, which is a predetermined time after the currently displayed time point (first time point). Here, the second user input may include at least one of an input to select a second object (e.g., a button object) displayed in the PIP area 910 or an input to scroll in a second direction (e.g., downward) opposite to the first direction on the PIP area 910. This allows the processor to change the time point of the call audio to be displayed based on the scrolling direction.
[0118] According to one embodiment, the processor may display text 914 converted from the current real-time call voice and information 912 about the speaker in the PIP area 910 after a predetermined time has elapsed or a user input (hereinafter referred to as a "third user input") has been received, after which the processor may return to a screen state displaying the real-time call voice while displaying the past call voice. Here, the third user input may include at least one of an input to select a third object (e.g., a button object) displayed in the PIP area 910 or an input to scroll to the end of the PIP area 910 in a second direction (e.g., downward).
[0119] According to one embodiment, when real-time call audio is received while displaying text converted from the call audio at a first time point and information about the speaker of the call audio at the first time point in the PIP area 910, the processor may display text converted from the real-time call audio 914 and information about the speaker of the real-time call audio 912 in a predetermined area of the PIP area 910. For example, when real-time call audio is received while displaying content of a past call, the processor may display the content of the real-time call in a predetermined area while maintaining the display state of the content of the past call. According to one embodiment, the predetermined area may be a fixed area at the bottom of the PIP area 910.
[0120] FIG. 11 is a diagram illustrating a method for displaying text 1114, 1116 converted from a call voice and information 1112 about the speaker on an execution screen 1100 of a chat application linked with a call application according to an embodiment of the present disclosure. Referring to FIG. 11 , a processor (e.g., processor 410 of FIG. 4 ) of an electronic device (e.g., electronic device 100 of FIG. 1 or electronic device 400 of FIG. 4 ) that processes the call voice can display, on a switched screen, text 1114, 1116 converted from real-time call voice, along with information 1112 about the speaker. At this time, the processor can display at least one of the current time or the duration of the call (e.g., the duration of the call), on the switched screen, along with the text 1114, 1116 converted from the call voice and the information 1112 about the speaker. Here, the switched screen may include the execution screen 1100 of the chat application linked with the call application. 11 shows a case where the screen after switching is an execution screen 1100 of a chat application that cooperates with a calling application, but is not limited to this, and the screen after switching can include an execution screen of an SNS application, a message application, or the like that can cooperate with the calling application and supports an interactive format. Also, for convenience of explanation, the screen before the screen switching will be referred to as the first screen, and the screen after the screen switching will be referred to as the second screen 1100. For example, the first screen can be an execution screen of a calling application (e.g., execution screen 500 of the calling application in FIG. 5 or execution screen 600 of the calling application in FIG. 6), and the second screen 1100 can be an execution screen of a chat application that cooperates with the calling application.
[0121] According to one embodiment, the processor may display text 1114, 1116 converted from the call voice on the second screen 1100 in an interactive format supported by the chat application based on information 1112 about the speaker of the call voice. In this case, the processor may display an object 1120 indicating the call status (e.g., a text box object) on the second screen 1100 to distinguish from when the chat application displays chat content, i.e., to indicate that a call is in progress rather than a chat. According to one embodiment, when the speaker of the call voice is the user of the electronic device, the processor may display only text 1116 converted from the voice spoken by the user on the second screen 1100, excluding information about the user.
[0122] According to one embodiment, the processor can support execution of a function of a chat application while interactively displaying text 1114, 1116 converted from the call voice on second screen 1100. For example, the processor can interactively display text 1114, 1116 converted from the call voice on second screen 1100 while simultaneously displaying text 1118 input via the chat application (e.g., chat input text) on second screen 1100 in chronological order. In other words, the electronic device can support a user to simultaneously talk and chat with a call partner.
[0123] According to one embodiment, in response to receiving a user input (hereinafter referred to as a “first user input”), the processor may display text converted from a call voice at a time point a predetermined time before the current time point (hereinafter referred to as a “first time point”) and information about the speaker of the call voice at the first time point on the second screen 1100. For example, when an input is received to select a first object (e.g., object 542 in FIG. 5 or object 610 in FIG. 6 ) displayed on the second screen 1100 or an input to scroll (e.g., scroll 544 in FIG. 5 ) in a first direction (e.g., upward) on the second screen 1100, the text 1114, 1116 converted from the real-time call voice and the information 1112 about the speaker of the real-time call voice displayed on the second screen 1100 may be scrolled in chronological order from the current time point to the first time point, and eventually the text converted from the call voice at the first time point and the information about the speaker of the call voice at the first time point may be displayed on the second screen 1100.
[0124] According to one embodiment, in response to receiving a user input (hereinafter referred to as a "second user input"), the processor may display on the second screen 1100 text converted from the call audio at a second time point after the first time point and information about the speaker of the call audio at the second time point. For example, when the processor receives a second user input while displaying the call audio from a past time point (first time point), the processor may display the call audio from the second time point a predetermined time after the currently displayed time point (first time point). Here, the second user input may include at least one of an input to select a second object (e.g., a button object) displayed on the second screen 1100 or an input to scroll on the second screen 1100 in a second direction (e.g., downward) opposite to the first direction. This allows the processor to change the time point of the call audio to be displayed based on the scroll direction.
[0125] According to one embodiment, after displaying text converted from the call voice at a first time on the second screen 1100, the processor may display text 1114, 1116 converted from the current real-time call voice and information 1112 about the speaker on the second screen 1100 after a predetermined time has elapsed or when a user input (hereinafter referred to as a "third user input") is received. For example, the processor may return to a screen state displaying real-time call content after a predetermined time has elapsed or a third user input has been received while displaying past call content. Here, the third user input may include at least one of an input to select a third object (e.g., a button object) displayed on the second screen 1100 or an input to scroll to the end of a second direction (e.g., downward) on the second screen 1100.
[0126] According to one embodiment, when real-time call audio is received while the processor is displaying text converted from the call audio at a first time point and information about the speaker of the call audio at the first time point on the second screen 1100, the processor may display the text converted from the real-time call audio and information about the speaker of the real-time call audio in a predetermined area of the second screen 1100. For example, when the processor receives real-time call audio while displaying content of a past call, the processor may display the content of the real-time call in a predetermined area while maintaining the display state of the content of the past call. According to one embodiment, the predetermined area may be a fixed area at the bottom of the second screen 1100.
[0127] 12 is a diagram illustrating a method for processing sounds corresponding to designated patterns in a call voice according to an embodiment of the present disclosure. Referring to FIG. 12, a processor (e.g., processor 410 in FIG. 4) of an electronic device (e.g., electronic device 100 in FIG. 1 or electronic device 400 in FIG. 4) that processes the call voice may display text 1214 converted from the call voice and information 1212 about the speaker of the call voice on an execution screen 1200 of a call application. For example, the processor may display the contents of the call on the screen during a call.
[0128] According to one embodiment, when the call voice is identified as a sound corresponding to a designated pattern, the processor may display an image 1224 mapped to the designated pattern on a screen together with information 1222 about the speaker of the call voice. Here, the sound corresponding to the designated pattern may include at least one of a sound generated by the speaker's actions (e.g., laughter, coughing, etc.) or a sound generated by an external object (e.g., a car sound, music, etc.). Furthermore, the image 1224 mapped to the designated pattern may be, for example, a sticker image, and may include at least one of an image with a shape similar to the speaker's actions or an image with a shape similar to an external object.
[0129] According to one embodiment, when the converted call voice text 1214 includes specific text, the processor may at least one of activate a vibration actuator, apply a visual effect to an area where the converted call voice text 1214 is displayed, or apply a visual effect to an area where the specific text in the converted call voice text 1214 is displayed. For example, when the processor detects specific text in the call content, the processor may vibrate the electronic device or apply a visual effect to an area where the converted call voice text 1214 or the specific text is displayed. According to one embodiment, the specific text may include information related to the user. The information related to the user may include information that can identify the user (e.g., the user's name, the user's nickname, or a nickname for the user).
[0130] FIG. 13 is a diagram illustrating a method for processing a call log according to an embodiment of the present disclosure. Referring to FIG. 13, a processor (e.g., processor 410 in FIG. 4) of an electronic device (e.g., electronic device 100 in FIG. 1 or electronic device 400 in FIG. 4) that processes call audio may store log information related to the call based on text 1324, 1326 converted from the real-time call audio and information 1322 related to the speaker of the real-time call audio. For example, the processor may store the text 1324, 1326 converted from the real-time call audio and information 1322 related to the speaker of the real-time call audio in a memory (e.g., memory 430 in FIG. 4) immediately after the call ends or in response to receiving a user input. In this case, the processor may store information related to the time point of the real-time call audio together with the text 1324, 1326 converted from the real-time call audio and information 1322 related to the speaker of the real-time call audio. Here, the information regarding the time point of the real-time call voice may include at least one of time information when the real-time call voice was received (current time information at the time of the call) or time information elapsed from the start of the call to the time when the real-time call voice was received (current time information at the time of the call).
[0131] According to one embodiment, the processor may display log information related to the call on screen 1320. For example, in response to receiving a user input selecting an object 1310 (e.g., a button object) included in screen 1300 displayed immediately after the call ends as in first state 1302, the processor may display the log information related to the call on screen 1320 as in second state 1304. Figure 13 shows a state in which the processor outputs the log information related to the call on screen 1320 in a popup format.
[0132] According to one embodiment, when a user input is received selecting specific text displayed on a screen displaying text converted from call voice (e.g., the call application execution screen 500 of FIG. 5, the call application execution screen 600 of FIG. 6, the call application execution screen 700 of FIGS. 7 and 8, the execution screen 900 of an application other than the call application of FIGS. 9 and 10, the execution screen 1100 of a chat application linked to the call application of FIG. 11, or the call application execution screen 1200 of FIG. 12), the processor may display the specific text differently (e.g., highlighted) from other text in the call log information displayed on screen 1320. For example, when specific text is selected from the text converted from real-time call voice, the processor can store the selected text separately from other text when saving the call log information, and can also display the selected text separately from other text when displaying the call log information based on that.
[0133] According to one embodiment, the processor may determine whether the user subscribes to a predetermined service (e.g., a premium service) and determine whether to store call log information based on whether the user subscribes to the predetermined service. For example, the processor may store call log information only if the user subscribes to the predetermined service.
[0134] According to one embodiment, the processor may transmit log information about the call to an external electronic device in response to receiving user input, for example, the processor may share log information about the call with the electronic device of the other party or another user in response to receiving user input selecting a button object that supports sharing of call logs.
[0135] According to one embodiment, the processor may check whether the user of the external electronic device subscribes to a predetermined service (e.g., a premium service) and determine whether to transmit call log information based on whether the user of the external electronic device subscribes to the predetermined service. As one example, the processor may restrict sharing of call log information if the user of the external electronic device does not subscribe to the predetermined service. As another example, the processor may transmit to the external electronic device a portion of the call log information (e.g., restricted information) or information converted from the call log information into a different format (e.g., an image capturing at least a portion of the call log information) if the user of the external electronic device does not subscribe to the predetermined service.
[0136] According to one embodiment, the processor may modify at least a portion of the log information for the call in response to receiving user input, for example, a user may modify text included in the log information for the call.
[0137] According to one embodiment, if a call is a multi-party call, the processor can search for the speech content of a specific user from the log information related to the multi-party call by filtering. As one example, in response to receiving a user input, the processor can search for the speech content of a specific user from the log information related to the multi-party call and display it differently from the speech content of other users. As another example, in response to receiving a user input, the processor can search for the speech content of a specific user from the log information related to the multi-party call and display only the speech content of the specific user that has been searched for on the screen 1320.
[0138] 14 is a diagram illustrating a procedure for displaying text converted from past call voices and information about the speaker in a call voice processing method according to an embodiment of the present disclosure. Referring to FIG. 14, a processor (e.g., processor 410 in FIG. 4) of an electronic device (e.g., electronic device 100 in FIG. 1 or electronic device 400 in FIG. 4) that processes call voices may convert real-time call voices into text at step 1410 (S1410). For example, the processor may convert at least one of the user's call voice acquired in real time via a microphone or the call partner's call voice acquired in real time from an external electronic device via a communication module (e.g., communication module 440 in FIG. 4) into text using the STT function.
[0139] At step 1420 (S1420), the processor may display the converted text together with information about the speaker. For example, the processor may identify the speaker of the acquired voice. As an example, if the voice is acquired through a microphone of the electronic device, the processor may identify the speaker of the acquired voice as the user of the electronic device. As another example, if the voice is acquired from an external electronic device through a communication module of the electronic device, the processor may identify the speaker of the acquired voice as the user of the external electronic device, i.e., the other party of the call. As yet another example, the processor may identify the speaker by voiceprint recognition of the voice. Furthermore, the processor may identify the speaker of the call voice based on at least one of the results of speech recognition of the call voice or information related to the speaker included in the text converted from the call voice. Here, the information related to the speaker may include information that can identify the speaker (e.g., the speaker's name, the speaker's nickname, or a nickname for the speaker). Then, the processor may display the text converted from the real-time call voice on a screen together with information about the speaker. At this time, the processor can display at least one of the current time or the duration of the call (for example, the duration of the call) on the screen together with the text converted from the call voice and information about the speaker.
[0140] At step 1430 (S1430), in response to receiving a user input, the processor may display text converted from a call voice at a time prior to the current time, together with information about the speaker. For example, in response to receiving a user input, the processor may display on a screen a first text converted from a call voice at a first time point, which is a predetermined time (e.g., 10 seconds) before the current time, together with information about the speaker of the call voice at the first time point. In this case, the processor may display on a screen information about the first time point, together with the first text and information about the speaker of the call voice at the first time point. Here, the information about the first time point may include at least one of time information about the first time point or information about the time elapsed from the start of the call to the first time point. In addition, the user input may include at least one of an input to select an object (e.g., a button object) displayed on the screen or an input to scroll (or swipe) the screen in a predetermined direction.
[0141] 15 is a diagram illustrating a procedure for displaying text converted from the call voice and information about the speaker on a screen after switching the screen in a call voice processing method according to an embodiment of the present disclosure. Referring to FIG. 15, a processor (e.g., processor 410 in FIG. 4) of an electronic device (e.g., electronic device 100 in FIG. 1 or electronic device 400 in FIG. 4) that processes the call voice may convert the real-time call voice into text at step 1510 (S1510). For example, the processor may convert at least one of the user's call voice acquired in real time via a microphone or the call partner's call voice acquired in real time from an external electronic device via a communication module (e.g., communication module 440 in FIG. 4) into text using the STT function.
[0142] In step 1520 (S1520), the processor may display the converted text on the first screen along with information about the speaker. For example, the processor may identify the speaker of the acquired voice. As an example, if the voice is acquired through a microphone of the electronic device, the processor may identify the speaker of the acquired voice as the user of the electronic device. As another example, if the voice is acquired from an external electronic device via a communication module of the electronic device, the processor may identify the speaker of the acquired voice as the user of the external electronic device, i.e., the call partner. As yet another example, the processor may identify the speaker by voiceprint recognition of the voice. Furthermore, the processor may identify the speaker of the call voice based on at least one of the results of speech recognition of the call voice or speaker-related information included in the text converted from the call voice. Here, the speaker-related information may include information that can identify the speaker (e.g., the speaker's name, the speaker's nickname, or a nickname for the speaker). Then, the processor may display the text converted from the real-time call voice on the first screen along with information about the speaker. At this time, the processor may display at least one of the current time or the duration of the call (for example, the duration of the call) on the first screen together with the text converted from the call voice and information about the speaker.
[0143] At step 1530 (S1530), the processor may switch the first screen to the second screen in response to receiving a user input. Here, if the first screen is an execution screen of the calling application, the second screen may be an execution screen of another application other than the calling application. Alternatively, if the first screen is an execution screen of the calling application, the second screen may be an execution screen of an application that cooperates with the calling application and supports an interactive format (e.g., a chat application, a social networking application, a messaging application, etc.). Alternatively, if the first screen is a first execution screen of the calling application, the second screen may be a second execution screen of the calling application. Here, the first execution screen of the calling application may be an initial execution screen or main screen of the calling application, and the second execution screen of the calling application may be a screen switched from the first execution screen of the calling application in response to receiving a user input, and an area where text converted from the calling voice is displayed may occupy most of the screen.
[0144] In step 1540 (S1540), the processor may display the converted text and information about the speaker on the second screen. For example, the processor may display the text converted from the real-time call voice and information about the speaker on the switched screen (second screen). At this time, the processor may display at least one of the current time or the call duration (e.g., the duration of the call) on the second screen together with the text converted from the call voice and the information about the speaker.
[0145] According to an embodiment, when the first screen is a screen for executing a call application and the second screen is a screen for executing an application other than the call application, the processor may display text converted from the call voice and information about the speaker in a PIP format on the second screen. In this case, the PIP area set in the second screen may have a size sufficient to include at least a predetermined number of pieces of text converted from the call voice.
[0146] According to one embodiment, when the first screen is an execution screen of a calling application and the second screen is an execution screen of a chat application linked to the calling application, the processor may display text converted from the call voice on the second screen in an interactive format supported by the chat application based on information about the speaker of the call voice. At this time, the processor may display an object (e.g., an image object) indicating the call status on the second screen to distinguish it from when the chat application displays chat content, i.e., to indicate that a call is in progress rather than a chat. In one embodiment, the processor may support execution of functions of the chat application while interactively displaying the text converted from the call voice on the second screen. For example, the processor may interactively display the text converted from the call voice on the second screen while also chronologically displaying text entered via the chat application (e.g., chat input text) on the second screen.
[0147] According to one embodiment, when the first screen is a first execution screen of a call application and the second screen is a second execution screen of the call application, the processor may convert the call voice into text and display it on the second screen in an interactive format between the call participants based on information about the speaker of the call voice. At this time, the processor may display an object (e.g., an image object) indicating the call status on the second screen to indicate that a call is in progress.
[0148] The above flowcharts and descriptions are merely examples, and different implementations are possible in some embodiments, for example, the order of the steps may be changed, some steps may be repeated, omitted, or added.
[0149] The above method may be provided as a computer program stored on a computer-readable recording medium for execution by a computer. The medium may permanently or temporarily store the computer-executable program. Furthermore, the medium may be a variety of recording / storage means, including a single or multiple hardware components, and may not be limited to media directly connected to a specific computer system but may also be distributed across a network. Examples of media on which program instructions may be stored include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and ROM, RAM, flash memory, etc. Other examples include recording / storage media managed by app stores that distribute applications or other software providers and distribution sites, servers, etc.
[0150] The methods, operations, or techniques of the present disclosure can be implemented by various means. For example, these techniques may be implemented using hardware, firmware, software, or a combination thereof. Those skilled in the art will appreciate that the various illustrative logic blocks, modules, circuits, and algorithm steps described herein may be implemented using electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described conceptually from a functional perspective. Whether these functions are implemented in hardware or software depends on the particular application and overall system design requirements. Those skilled in the art may implement the described functions in various ways depending on each application, but such implementations should not be interpreted as departing from the scope of the present disclosure.
[0151] In a hardware implementation, the processing unit that performs the techniques may be implemented as one or more ASICs, DSPs, digital signal processing devices, programmable logic devices, FPGAs, processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions of the present disclosure, computers, or combinations thereof.
[0152] Accordingly, the various illustrative logic blocks, modules, and circuits described in connection with this disclosure may be implemented or performed as a general purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination designed to perform the functions described herein. A general purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other combination of microprocessors.
[0153] In a firmware and / or software implementation, the techniques may be implemented as instructions stored on a computer-readable medium, such as RAM, ROM, NVRAM, PROM, EPROM, EEPROM, flash memory, CD, magnetic or marked data storage device, etc. The instructions are executable by one or more processors and may cause the processors to perform certain aspects of the functions described in this disclosure.
[0154] If implemented as software, the techniques described above may be stored on or transmitted over as one or more instructions or code on a computer-readable medium. Computer-readable media includes any medium that facilitates transfer of a computer program from one place to another and encompasses both computer storage media and communication media. Storage media may be any available medium that can be accessed by a computer. By way of non-limiting example, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM and other optical disk storage, magnetic disk storage and other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Furthermore, any connection is properly termed a computer-readable medium.
[0155] For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted wire, DSL, or wireless technologies such as infrared, radio, microwave, etc., the coaxial cable, fiber optic cable, twisted wire, DSL, or wireless technologies such as infrared, radio, microwave, etc. are included within the definition of medium. As used herein, disk and disc include CDs, laser discs, optical discs, DVDs, floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically and discs reproduce data optically using a laser. Combinations of the above are also included within the scope of computer-readable media.
[0156] A software module may reside in RAM, flash memory, ROM, EPROM, EEPROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.
[0157] Although the embodiments described above have been described as utilizing aspects of the subject matter of this disclosure in one or more stand-alone computer systems, the present disclosure is not limited thereto and may be implemented in conjunction with any computing environment, such as a network or distributed computing environment. Furthermore, aspects of the subject matter of this disclosure may be implemented in multiple processing chips or devices, and storage may be affected similarly across multiple devices. These devices may include PCs, network servers, and mobile devices.
[0158] Although the present disclosure has been described herein with reference to several embodiments, those skilled in the art will recognize that various changes and modifications may be made without departing from the scope of the present disclosure, and such changes and modifications are considered to be within the scope of the claims appended hereto.
[0159] This application claims priority based on Patent Application No. 10-2024-0093007 filed with the Korean Intellectual Property Office on July 15, 2024, the entire contents of which are incorporated herein by reference. [Explanation of symbols]
[0160] 100 Electronic equipment 400 Electronic equipment 410 processor 420 Display 430 memory 440 Communication Module
Claims
1. 1. A method of speech processing executed by at least one processor, comprising: converting real-time call audio into text; displaying the converted text on a first screen together with information about the speaker of the real-time call voice; In response to receiving a first user input, a first text converted from a call voice at a first time point, which is a predetermined time before the current time, is displayed on the first screen together with information about the speaker of the call voice at the first time point.
2. The step of displaying the first text on the first screen includes: The method for processing a call voice according to claim 1 , further comprising the step of displaying information about the first time point on the first screen together with the first text and information about a speaker of the call voice at the first time point.
3. The call voice processing method according to claim 1 , wherein the first user input includes an input for selecting a first object displayed on the first screen or an input for scrolling in a first direction on the first screen.
4. 4. The method for processing a call voice according to claim 3, further comprising the step of, in response to receiving a second user input, displaying on the first screen a second text obtained by converting the call voice at a second time point that is later than the first time point, together with information about a speaker of the call voice at the second time point.
5. 4. The method of claim 3, further comprising: displaying, on the first screen, text converted from the real-time call voice at the current time and information about a speaker of the real-time call voice, in response to a predetermined time having elapsed or a second user input being received after the first text is displayed on the first screen.
6. The call voice processing method of claim 5, wherein the second user input includes an input to select a second object displayed on the first screen, or an input to scroll to an end point of a second direction on the first screen that is opposite to the first direction.
7. 2. The method for processing a call voice according to claim 1, further comprising a step of translating the converted text into a language set for the speaker or the recipient of the real-time call voice when the language set for the speaker of the real-time call voice is different from the language set for the recipient of the real-time call voice.
8. The step of displaying the first text on the first screen includes: displaying the first text and information about a speaker of the call voice at the first time point in a first area of the first screen; The call voice processing method includes:
2. The method for processing a call voice according to claim 1, further comprising a step of displaying, in a state in which the first text and information about the speaker of the call voice at the first time are displayed in a first area of the first screen, text converted from the real-time call voice and information about the speaker of the real-time call voice in a second area of the first screen different from the first area in response to receiving the real-time call voice.
9. The method of claim 1 , further comprising: storing log information about the call based on the text converted from the real-time call voice and information about a speaker of the real-time call voice.
10. The step of storing log information relating to the call includes: The method for processing a call voice according to claim 9, further comprising a step of storing, immediately after the call is ended or in response to receiving a second user input, text into which the real-time call voice is converted and information about a speaker of the real-time call voice.
11. The step of storing log information relating to the call includes: The method of claim 9 , further comprising: storing information about a time point of the real-time call voice together with information about a text into which the real-time call voice is converted and information about a speaker of the real-time call voice.
12. The call voice processing method according to claim 9 , further comprising the step of displaying log information relating to the call on a second screen different from the first screen.
13. The step of displaying log information related to the call on the second screen includes:
13. The call voice processing method of claim 12, further comprising a step of displaying log information regarding the call on the second screen in response to receiving a second user input selecting a second object displayed on a third screen that is opened immediately after the call is ended.
14. 13. The call voice processing method of claim 12, further comprising the step of, in response to receiving a second user input selecting specific text displayed on the first screen, displaying the specific text in the log information related to the call displayed on the second screen differently from other text.
15. checking whether the user has subscribed to a predetermined service; The method of claim 9, further comprising: determining whether or not log information relating to the call needs to be saved based on whether or not the user has subscribed to the predetermined service.
16. 16. The method of claim 15, further comprising the step of transmitting log information regarding the call to an external electronic device in response to receiving a second user input.
17. checking whether a user of the external electronic device has subscribed to the predetermined service; The method of claim 16, further comprising: determining whether or not to transmit log information regarding the call based on whether or not a user of the external electronic device subscribes to the predetermined service.
18. transmitting log information about the call to the external electronic device, 18. The method for processing voice calls according to claim 17, further comprising the step of transmitting a part of the log information relating to the call or information obtained by converting the log information relating to the call into another format to the external electronic device if the user of the external electronic device does not subscribe to the specified service.
19. A program for causing a computer to execute the call voice processing method according to any one of claims 1 to 18.
20. 1. An electronic device comprising: The display and Memory and at least one processor coupled to the display and the memory and configured to execute at least one computer-readable program stored in the memory; The at least one program Convert real-time call voice to text, Displaying the converted text on a first screen of the display together with information about the speaker of the real-time call voice; An electronic device comprising instructions for, in response to receiving a first user input, displaying on the first screen a first text converted from a call voice at a first time point that is a predetermined time before the current time, together with information about the speaker of the call voice at the first time point.
Citation Information
Patent Citations
Electronic device for processing user voice during performing call with a plurality of users
KR102801805B1