Electronic device, method, and non-transitory computer-readable recording medium for linking content with user input
The integration of a content service application with generative AI and linking modules addresses the challenge of synchronizing audio and text data, enabling efficient content creation and summarization by converting and linking user inputs, thus enhancing user interaction and productivity.
Patent Information
- Application Number
- PCT/KR2025/006904
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-27
- Filing Date
- 2025-05-21
- Publication Date
- 2026-01-15
AI Technical Summary
Existing electronic devices struggle to effectively integrate and synchronize audio data and text data, particularly in linking user inputs such as handwriting and voice inputs, for efficient content creation and summarization.
The integration of a content service application with a generative AI module and linking module that processes user inputs through optical character recognition (OCR) and automatic speech recognition (ASR) to convert handwriting and voice into text data, and generates summaries using neural networks, while linking these data types in a split view for enhanced content creation.
Enables seamless integration and synchronization of audio and text data, allowing for efficient content creation and summarization, enhancing user interaction and productivity by visually highlighting relevant content based on user inputs.
Smart Images

Figure KR2025006904_15012026_PF_FP_ABST
Abstract
Description
Electronic devices, methods, and non-transitory computer-readable recording media for linking content and user input
[0001] The descriptions below relate to electronic devices, methods, and non-transitory computer-readable recording media that link content and user input.
[0002] Electronic devices can provide content to users. While providing content to users, the electronic device can receive user input. The electronic device can store data representing the received user input as a memo of the content. For example, user input can include not only text or images entered by the user, but also voice input.
[0003] An electronic device is disclosed. The electronic device may include a display, at least one processor including a processing circuit; and a memory storing instructions and including one or more storage media. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, through the display, a first text summarizing text data representing audio data and a second text representing the contents of a document, according to a split view. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify an input for a word within the second text. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the at least one word visually highlighted in the first text based on identifying an utterance corresponding to the word identified according to the input among utterances in text data representing the audio data and identifying at least one word in the first text summarizing the utterance, while displaying the word in the second text visually highlighted according to the input.
[0004] An electronic device is disclosed. The electronic device may include a display, at least one processor including a processing circuit; and a memory storing instructions and including one or more storage media. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, through the display, a first text representing audio data and a second text summarizing a document according to a split view. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify an input for a first word within the first text. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the at least one third word visually highlighted in the second text based on identifying at least one second word corresponding to the first word identified in the input among words in the document and identifying at least one third word in the second text corresponding to the at least one second word in the document, while displaying the first word in the first text visually highlighted in response to the input.
[0005] A method is disclosed. The method can be performed by an electronic device including a display. The method can include an operation of displaying, through the display, a first text summarizing text data representing audio data and a second text representing the contents of a document according to a split view. The method can include an operation of identifying an input for a word in the second text. The method can include an operation of identifying an utterance corresponding to the word identified according to the input among utterances in the text data representing the audio data, and an operation of displaying the at least one word visually highlighted in the first text while displaying the word visually highlighted in the second text according to the input, based on the identification of at least one word in the first text summarizing the utterance.
[0006] A method is disclosed. The method can be performed in an electronic device including a display. The method can include an operation of displaying, through the display, a first text representing audio data and a second text summarizing a document according to a split view. The method can include an operation of identifying an input for a first word in the first text. The method can include an operation of identifying at least one second word corresponding to the first word identified according to the input among words in the document, and an operation of displaying at least one third word in the second text corresponding to the at least one second word in the document, while displaying the first word in the first text visually highlighted according to the input.
[0007] A non-transitory computer-readable storage medium is disclosed. The non-transitory computer-readable storage medium can store a program including instructions. The instructions, when individually or collectively executed by at least one processor of an electronic device including a display, can cause the electronic device to display, through the display, a first text summarizing text data representing audio data and a second text representing the content of a document according to a split view. The instructions, when individually or collectively executed by the at least one processor, can cause the electronic device to identify an input for a word within the second text. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the at least one word visually highlighted in the first text based on identifying an utterance corresponding to the word identified according to the input among utterances in text data representing the audio data and identifying at least one word in the first text summarizing the utterance, while displaying the word in the second text visually highlighted according to the input.
[0008] A non-transitory computer-readable recording medium is disclosed. The non-transitory computer-readable recording medium can store a program including instructions. The instructions, when individually or collectively executed by at least one processor of an electronic device including a display, can cause the electronic device to display, through the display, a first text representing audio data and a second text summarizing a document according to a split view. The instructions, when individually or collectively executed by the at least one processor, can cause the electronic device to identify an input for a first word within the first text. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the at least one third word visually highlighted in the second text based on identifying at least one second word corresponding to the first word identified in the input among words in the document and identifying at least one third word in the second text corresponding to the at least one second word in the document, while displaying the first word in the first text visually highlighted in response to the input.
[0009] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.
[0010] Figure 2 is a block diagram of an electronic device according to one embodiment.
[0011] FIG. 3A is a diagram illustrating an example of a screen of a content service application displayed by an electronic device according to one embodiment.
[0012] FIG. 3b is a diagram illustrating an example of a screen of a content service application including a user interface (UI) for obtaining audio data through a content service application displayed by an electronic device according to one embodiment.
[0013] FIG. 3c is a diagram illustrating an example of a screen in which an electronic device obtains audio data and note data through a content service application according to one embodiment.
[0014] FIG. 3D is a diagram illustrating another example of a screen in which an electronic device obtains audio data and note data through a content service application according to one embodiment.
[0015] FIG. 4 is a diagram showing the relationship between audio data and a document according to one embodiment.
[0016] FIG. 5A is a diagram showing a relationship between text data converted from audio data and a summary that summarizes the text data, according to one embodiment.
[0017] FIG. 5b is a diagram showing the relationship between text data converted from strokes and a summary that summarizes the text data, according to one embodiment.
[0018] FIG. 6A is a diagram illustrating a table showing the relationship between text data converted from audio data and text data converted from strokes, according to one embodiment.
[0019] FIG. 6b is a diagram illustrating a table showing the relationship between text data converted from audio data and a summary according to one embodiment.
[0020] FIG. 6c is a diagram illustrating a table showing the relationship between text data converted from strokes and a summary, according to one embodiment.
[0021] FIG. 7A is a diagram illustrating an example of a screen of note content displaying text data converted from strokes and audio data, according to one embodiment.
[0022] FIG. 7b is a diagram illustrating an example of a screen of note content displaying note content including strokes and summaries according to one embodiment.
[0023] FIG. 7c is a diagram illustrating an example of a screen of note content displaying strokes and a summary of text data converted from strokes, according to one embodiment.
[0024] FIG. 7D is a diagram illustrating an example of a screen of note content displaying a summary of text data converted from strokes and audio data, according to one embodiment.
[0025] FIG. 7e is a diagram illustrating an example of a screen of note content displaying text data and summaries converted from audio data, according to one embodiment.
[0026] FIG. 7F is a diagram illustrating an example of a screen of note content displaying a summary of text data converted from audio data and text data converted from strokes, according to one embodiment.
[0027] FIG. 7g is a diagram illustrating an example of a screen of note content that displays text data converted from audio data and a summary of text data converted from audio data, according to one embodiment.
[0028] FIG. 7h is a diagram illustrating an example of a screen of note content displaying summaries according to one embodiment.
[0029] FIG. 8A is a diagram illustrating an example of a screen of note content in which an electronic device highlights and displays another object corresponding to a selected object, according to one embodiment.
[0030] FIG. 8B is a diagram illustrating another example of a screen of note content in which an electronic device highlights and displays another object corresponding to a selected object, according to one embodiment.
[0031] FIG. 8C is a diagram illustrating another example of a screen of note content in which an electronic device highlights and displays objects according to a playback point in time, according to one embodiment.
[0032] FIG. 8D is a diagram illustrating another example of a screen of note content in which an electronic device highlights and displays objects according to a playback point in time, according to one embodiment.
[0033] FIG. 8E is a diagram illustrating another example of a screen of note content in which an electronic device highlights and displays objects based on user input, according to one embodiment.
[0034] FIG. 9A is a diagram illustrating another example of a screen of note content displaying text data converted from strokes and audio data, according to one embodiment.
[0035] FIG. 9b is a diagram illustrating another example of a screen of note content displaying a summary of text data converted from strokes and audio data, according to one embodiment.
[0036] FIG. 9c is a diagram illustrating an example of a screen of note content displaying a translation of a summary of text data converted from strokes and audio data, according to one embodiment.
[0037] FIG. 9D is a diagram illustrating an example of a screen of note content in which an electronic device highlights and displays another object corresponding to a selected object, according to one embodiment.
[0038] FIG. 10A is a diagram showing the relationship between pages and utterance sets within a document, according to one embodiment.
[0039] FIG. 10b is a diagram illustrating an example of a screen in which an electronic device displays a page and a set of utterances corresponding to the page, according to one embodiment.
[0040] FIGS. 11A to 11C are diagrams illustrating examples of visual objects indicating synchronization within a document displayed by an electronic device, according to one embodiment.
[0041] FIG. 12 is a flowchart illustrating the operation of an electronic device according to one embodiment.
[0042] FIG. 13A is a diagram illustrating an example of a screen of note content displaying text data converted from strokes and audio data, according to one embodiment.
[0043] FIG. 13b is a diagram illustrating an example of a progress bar according to one embodiment.
[0044] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments.
[0045] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0046] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0047] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0048] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0049] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0050] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0051] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0052] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0053] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0054] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0055] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0056] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0057] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0058] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0059] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0060] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0061] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0062] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for realizing eMBB, a loss coverage (e.g., 664 dB or less) for realizing mMTC, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 6 ms or less for round trip) for realizing URLLC.
[0063] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0064] According to various embodiments, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0065] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0066] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0067] Figure 2 is a block diagram of an electronic device according to one embodiment.
[0068] The electronic device (101) of FIG. 2 may correspond to the electronic device (101) of FIG. 1. The electronic device (101) of FIG. 2 may be described with reference to the electronic device (101) of FIG. 1.
[0069] In one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), a display (260), and a microphone (270). For example, the processor (120) may refer to a plurality of processors that collectively perform a plurality of operations by dividing them among the processors. The display (260) of FIG. 2 may correspond to the display module (160) of FIG. 1. The microphone (270) of FIG. 2 may be included as a part of the input module (150) of FIG. 1. In one embodiment, the program (140) of the memory (130) may include a content service application (APP) (210), a generative artificial intelligence (AI) module (220), and a linking module (230).
[0070] In one embodiment, the content service APP (210), the generative AI module (220), and the linkage module (230) may correspond to the application (146) of FIG. 1. For example, the content service APP (210), the generative AI module (220), and the linkage module (230) may include instructions that may be executed by the processor (120).
[0071] In one embodiment, the content service APP (210) can display (or output) (or play) content. In one embodiment, the content can include documents, images, videos, and / or audio. In one embodiment, the documents can include notes written directly by the user.
[0072] In one embodiment, the content service APP (210) may obtain a user's handwriting input and / or typing input through the input module (150) (or display (260)). Hereinafter, handwriting input and typing input may be referred to as text input.
[0073] In one embodiment, the content service APP (210) can receive voice input through a microphone (270).
[0074] In one embodiment, the content service APP (210) may store text, images, voices, and / or videos identified (or received) by user input as notes (or note content) (hereinafter, notes). In one embodiment, the note may include at least one of the content of a document (e.g., document (420) of FIG. 4), a user's text input entered into the document (e.g., handwriting input (e.g., multiple strokes (421) of FIG. 4), text data converted from the text input (e.g., text data (423) of FIG. 4), or a summary of the text data converted from the text input (e.g., summary (425) of FIG. 4), a voice input (e.g., audio data (410) of FIG. 4, or multiple utterances (411) of FIG. 4), text data converted from the voice input (e.g., text data (413) of FIG. 4), or a summary of the text data converted from the voice input (e.g., summary (415) of FIG. 4). According to an embodiment, the content service APP (210) may include at least one of the content of a document, the user's text input entered into the document (e.g., handwriting input), text data converted from the text input, voice input, or text data converted from the voice input. It may include summaries (e.g., summaries of documents and user text input, documents, user text input, and voice input). In one embodiment, the content service APP (210) may temporally and / or semantically link the content of a document, the user's text input entered into the document, text data converted from the text input, voice input, text data converted from the voice input, and / or summaries thereof. For example, the linking module (230) may generate relationship information (or mapping information) that temporally and / or semantically link the content of a document, the user's text input entered into the document, text data converted from the text input, voice input, text data converted from the voice input, and / or summaries thereof.Below, relationship information can be explained with reference to Fig. 4.
[0075] In one embodiment, the content service APP (210) may provide optical character recognition (OCR) and automatic speech recognition (ASR) (or speech to text (STT)) for user input.
[0076] In one embodiment, the content service APP (210) can convert handwriting received through handwriting input via OCR into text data (or character data) that the electronic device (101) can process. In one embodiment, the content service APP (210) can store the text data (or character data) converted from the handwriting as data included in a note. However, the present invention is not limited thereto. In one embodiment, the content service APP (210) can request the generative AI module (220) to convert handwriting received through handwriting input into text data (or character data) that the electronic device (101) can process.
[0077] In one embodiment, the electronic device (101) can detect user input (e.g., touch input, pen input, stylus pen input, finger input, gesture input, and / or segment input) through the display (260). In one embodiment, the electronic device (101) can perform an input command according to the detected user input.
[0078] For example, when a drawing input (or sketch, drawing) is detected based on a user input (e.g., touch input, pen input, stylus pen input, finger input, gesture input, and / or segment input), the electronic device (101) may generate points, lines, and images corresponding to the drawing input, and display the generated points, lines, and images through the display (260). The electronic device (101) may generate a drawing object according to the drawing input. In one embodiment, the drawing object may represent points, lines, and images. In one embodiment, the drawing input (or sketch, drawing) may be analyzed by the generative AI module (220). In one embodiment, an object represented by the drawing input (or sketch, drawing) may be identified by the analysis of the generative AI module (220). In one embodiment, for analysis by the generative AI module (220), a drawing input (or sketch, drawing), texts surrounding the drawing input (or sketch, drawing), and / or content (e.g., images, videos) associated with a user account may be input as prompts to the generative AI module (220).
[0079] For example, when a segment input is detected, the electronic device (101) may monitor the start time when the segment input starts, the end time when the segment input ends, and / or the movement path (e.g., movement, movement gesture) from the start time to the end time of the segment input. The electronic device (101) may determine whether the segment input corresponds to a set character (e.g., letter, text), a symbol (e.g., asterisk, underline, emphasis), and / or a command (e.g., format, copy, paste). In one embodiment, when the segment input corresponds to a set input, the electronic device (101) may generate a character and / or symbol corresponding to the segment input. The electronic device (101) may extract the character and / or symbol corresponding to the segment input and provide it to the user. For example, the electronic device (101) may display the character and / or symbol corresponding to the segment input through the display (260). In one embodiment, the electronic device (101) may perform a command (e.g., formatting, copying, pasting) corresponding to the segment input when the segment input corresponds to a set command. Hereinafter, a single character, symbol, or command identified through handwriting input by a user input (e.g., multiple strokes (421) of FIG. 4) may be referred to as an object (or input object).
[0080] In one embodiment, the content service APP (210) can convert voice received by voice input through ASR (or STT) into text data (or character data) that the electronic device (101) can process. In one embodiment, the content service APP (210) can store the text data (or character data) converted from voice as data included in a note. However, the present invention is not limited thereto. In one embodiment, the content service APP (210) can request the generative AI module (220) to convert voice received by voice input through ASR (or STT) into text data (or character data) that the electronic device (101) can process.
[0081] In one embodiment, the content service APP (210) may request a summary of text data (or character data) converted from handwriting from the generative AI module (220). In one embodiment, the content service APP (210) may request a summary of text data (or character data) converted from speech from the generative AI module (220).
[0082] In one embodiment, the generative AI module (220) may include an AI model including a plurality of parameters related to a neural network having a structure based on an encoder and a decoder, such as a transformer. In one embodiment, the generative AI module (220) may include a bi-directional model based on learning about an encoder (e.g., bidirectional encoder representations from transformers (BERT)) or an auto-encoding model (e.g., a diffusion model). In one embodiment, the generative AI module (220) may include an auto-regressor model based on learning about a decoder (e.g., a generative pre-trained transformer (GPT)). In one embodiment, the generative AI module (220) may include a sequence-to-sequence model (e.g., stable diffusion, DALL-E 2) based on learning about an encoder and a decoder. In one embodiment, the generative AI module (220) may include a large language model (LLM) for processing natural language based on a massive number of parameters, but is not limited thereto. The generative AI module (220) may include parameters for driving a neural network such as a convolutional neural network (CNN), a recurrent neural network (RNN), a feedforward neural network (FNN), and / or a long short-term memory (LSTM).
[0083] In one embodiment, the generative AI module (220) can generate content based on a prompt. For example, the generative AI module (220) can generate processed content from original content based on the prompt. For example, the processed content can be a summary of the original content. For example, if the original content is text data converted from handwriting, the generative AI module (220) can generate a summary of the text data. For example, if the original content is text data converted from speech, the generative AI module (220) can generate a summary of the text data. In one embodiment, the prompt can include at least one sentence (or at least one word) to guide the generation of a summary based on text data converted from handwriting. In one embodiment, the prompt can include at least one sentence (or at least one word) to guide the generation of a summary based on text data converted from speech. In one embodiment, the generative AI module (220) can generate processed content based on a prompt that includes original content (e.g., text data converted from speech, text data converted from handwriting).
[0084] In one embodiment, the generative AI module (220) may generate a one-sentence summary for text data of a specified length (e.g., 250 words or 5,000 words). In one embodiment, the generative AI module (220) may identify a topic (or main content) in original content. In one embodiment, the generative AI module (220) may generate one or more sentences for a paragraph that includes content related to the topic (or main content) identified in the original content. In one embodiment, the generative AI module (220) may not generate a sentence for a paragraph that does not include content related to the topic (or main content) identified in the original content.
[0085] In one embodiment, the linking module (230) may link data included in a note. For example, the linking module (230) may link a document portion and / or an audio data portion included in the note. For example, the document portion may include handwriting, text data converted from handwriting, and a summary thereof. For example, the audio data portion may include audio data, text data converted from audio data, and a summary thereof.
[0086] For example, the linking module (230) can generate information (or relationship information) (or mapping information) that links text data converted from handwriting and text data converted from speech. For example, the linking module (230) can generate information (or relationship information) (or mapping information) that links text data converted from speech and a summary thereof. For example, the linking module (230) can generate information (or relationship information) that links a summary from handwriting and a summary from speech. In one embodiment, the information (or relationship information) that links a summary from handwriting and a summary from speech may include information that links a specific section among a plurality of sections of speech (or video), text extracted from speech (or video), and / or a summary of text extracted from speech (or video). In one embodiment, information (or relationship information) linking a summary from handwriting and a summary from speech may link content of a document, user text input entered into a document, text data converted from text input, speech input, text data converted from speech input, and / or a summary thereof with a specific section among a plurality of sections of speech (or video), text extracted from speech (or video), and / or a summary of text extracted from speech (or video).
[0087] Hereinafter, with reference to the drawings, an operation of an electronic device (101) generating a note, an operation of processing the generated note, and an operation of providing the processed note to a user are described.
[0088] Hereinafter, with reference to FIGS. 3A to 3D, the operation of the electronic device (101) generating a note is described.
[0089] FIG. 3A is a diagram illustrating an example of a screen of a content service application displayed by an electronic device according to an embodiment. FIG. 3B is a diagram illustrating an example of a screen of a content service application including a UI (user interface) for obtaining audio data through a content service application displayed by an electronic device according to an embodiment. FIG. 3C is a diagram illustrating an example of a screen for obtaining audio data and note data through a content service application by an electronic device according to an embodiment. FIG. 3D is a diagram illustrating another example of a screen for obtaining audio data and note data through a content service application by an electronic device according to an embodiment.
[0090] FIGS. 3A to 3D may be described with reference to the electronic device (101) (or components of the electronic device (101)) of FIG. 2. The operations described with reference to FIGS. 3A to 3D may be performed by the electronic device (101) as the processor (120) executes instructions included in the content service APP (210), the generative AI module (220), and the linkage module (230) of FIG. 2.
[0091] In one embodiment, the electronic device (101) may receive a user input for executing a content service APP (210). In one embodiment, the electronic device (101) may execute the content service APP (210) based on the user input for executing the content service APP (210). In one embodiment, referring to FIG. 3A, the electronic device (101) may display a screen (301) of the content service APP (210) through the display (260) based on executing the content service APP (210).
[0092] In one embodiment, the screen (301) may include an area (311) in which data corresponding to a user input received through the input module (150) is recorded. For example, handwriting (or a handwriting image) and / or text (340) according to a typed input may be input into the area (311). For example, a symbol (350) (e.g., an asterisk, an underline, an emphasis) according to a handwriting input may be input into the area (311). In one embodiment, the text (340) and / or symbol (350) according to the handwriting input may be input into the area (311) so as to overlap at least a portion of the content (331). For example, a symbol (350) corresponding to the underline of "4th Industrial Revolution" in the content (331) may be input. However, the present invention is not limited thereto. For example, an image and / or a video may be input into the area (311).
[0093] In one embodiment, the screen (301) may include an area (313) for providing additional functions related to user input. In one embodiment, the area (313) may include icons for setting characteristics (e.g., color, input tool (e.g., pen, marker, keyboard)) of data corresponding to user input received through the input module (150). In one embodiment, the area (313) may include icons for setting changes (e.g., delete, go back, align, correct text) of data corresponding to user input received through the input module (150).
[0094] In one embodiment, the screen (301) may include an affordance (315) for adding various objects, such as images, audio, video, or pictures. In one embodiment, the affordance may be referenced as an icon, a visual object, or an executable object.
[0095] In one embodiment, the electronic device (101) may obtain a user input for selecting an affordance (315) of the screen (301). Referring to FIG. 3B, in response to the user input for selecting the affordance (315) of the screen (301), the electronic device (101) may display a user interface (UI) (305) for adding various objects, such as images, audio files, or pictures, on the screen (301).
[0096] In one embodiment, the electronic device (101) may obtain a user input for selecting a visual object (320) for voice recording on a UI (305) on a screen (301).
[0097] In one embodiment, the electronic device (101) may receive a voice input (or record a voice) through the microphone (270) in response to a user input selecting a visual object (320) for voice recording in the UI (305). Referring to FIG. 3C, the electronic device (101) may display a visual object (335) indicating that recording is in progress while receiving a voice input through the microphone (270). In one embodiment, the visual object (335) may include information about the recording time. For example, the visual object (335) may indicate the elapsed time of the recording. For example, the visual object (335) may indicate the elapsed time of the recording in seconds and minutes. For example, the visual object (335) may indicate the elapsed time of the recording in seconds, minutes, and hours. In one embodiment, the electronic device (101) may display a visual object indicating that recording is in progress in a notification area (e.g., an upper area of the display (260)) while receiving a voice input through the microphone (270). In one embodiment, the visual object indicating that recording is in progress displayed in the notification area (e.g., an upper area of the display (260)) may blink.
[0098] In one embodiment, the electronic device (101) may obtain a user input for selecting a visual object (325) to retrieve content (331) from a UI (305) on a screen (301) (or to attach it to a document (e.g., document (420) of FIG. 4). In one embodiment, the electronic device (101) may display the content (331) within an area (311) in response to a user input for selecting a visual object (325) to retrieve content (331) from the UI (305). In one embodiment, the content (331) may be a document in a specified format (e.g., portable document format (PDF)).
[0099] In one embodiment, the electronic device (101) may obtain handwriting input and / or typing input. In one embodiment, the electronic device (101) may obtain user input (e.g., handwriting input and / or typing input) for inputting characters in an overlapping state over the content (331) while the content (331) is displayed.
[0100] Referring to FIG. 3D, the electronic device (101) may obtain multiple handwriting inputs through the input module (150) (e.g., a touchscreen or a digitizer) while recording voice through the microphone (270). In one embodiment, the handwriting input may include at least one stroke, but is not limited thereto. For example, the electronic device (101) may obtain multiple handwriting inputs through the input module (150) (e.g., a touchscreen or a digitizer) after recording voice through the microphone (270). For example, the electronic device (101) may obtain multiple handwriting inputs through the input module (150) (e.g., a touchscreen or a digitizer) before recording voice through the microphone (270). In one embodiment, the electronic device (101) may, in response to a handwriting input for entering letters and / or symbols into the content (331), display letters and / or symbols according to the handwriting input in an overlapping state on the content (331). For example, referring to FIG. 3D , the electronic device (101) may, in response to a handwriting input for entering symbols in a margin within the content (331), display symbols (350) according to the handwriting input in an overlapping state on the content (331).
[0101] In one embodiment, the electronic device (101) may process voice input and / or document data. For example, processing voice input may include converting voice input into audio data. For example, processing voice input may include converting audio data into text data. For example, processing voice input may include summarizing text data based on audio data. For example, processing document data may include converting text input (e.g., handwritten input or typed input) into text data. For example, processing document data may include summarizing text data based on text input.
[0102] In one embodiment, the electronic device (101) may generate relationship information between voice input and / or document data. For example, the relationship information may include relationship information between text data based on audio data and text data based on text input. For example, the relationship information may include relationship information between a summary related to audio data and a summary related to text input. For example, the relationship information may include relationship information between text data based on audio data and a summary related to audio data. For example, the relationship information may include relationship information between text data based on text input and a summary related to the text input. In one embodiment, the relationship information may also be referred to as mapping information.
[0103] Hereinafter, with reference to FIGS. 4 to 6c, an operation of an electronic device (101) processing voice input and / or document data, or generating relationship information between voice input and / or document data, is described.
[0104] Hereinafter, with reference to FIGS. 4 to 6c, the operation of processing the generated note by the electronic device (101) will be described.
[0105] FIG. 4 is a diagram illustrating a relationship between audio data and a document, according to an embodiment. FIG. 5a is a diagram illustrating a relationship between text data converted from audio data and a summary summarizing the text data, according to an embodiment. FIG. 5b is a diagram illustrating a relationship between text data converted from strokes and a summary summarizing the text data, according to an embodiment. FIG. 6a is a diagram illustrating a table illustrating a relationship between text data converted from audio data and text data converted from strokes, according to an embodiment. FIG. 6b is a diagram illustrating a table illustrating a relationship between text data converted from audio data and a summary, according to an embodiment. FIG. 6c is a diagram illustrating a table illustrating a relationship between text data converted from strokes and a summary, according to an embodiment.
[0106] FIGS. 4 to 6c may be described with reference to the electronic device (101) (or components of the electronic device (101)) of FIG. 2.
[0107] Referring to FIG. 4, the electronic device (101) can obtain text data (413) from audio data (410) through ASR (or STT). For example, the electronic device (101) can identify a plurality of pronunciations from a plurality of utterances (411) (or voice input) in the audio data (410) through ASR (or STT). For example, the electronic device (101) can obtain text data (413) representing a plurality of pronunciations corresponding to the plurality of utterances (411) through ASR (or STT). In one embodiment, the text data (413) may be data in which the format of the plurality of utterances (411) is changed. For example, texts included in the text data (413) may correspond to the plurality of utterances (411).
[0108] In one embodiment, the electronic device (101) may obtain text data (423) from a document (420) through OCR. For example, the electronic device (101) may identify a plurality of characters from a plurality of strokes (421) (or text input) in the document (420) through OCR. For example, the electronic device (101) may obtain text data (423) from a plurality of characters corresponding to the plurality of strokes (421) through OCR. In one embodiment, the plurality of characters corresponding to the plurality of strokes (421) may include at least one of uppercase letters, lowercase letters, numbers, fractions, ligatures, punctuation marks, mathematical symbols, phonetic symbols, stress symbols, currency symbols, decorative symbols, or emphasis symbols. The text data (423) may be displayed while maintaining the shape (e.g., length, position, thickness, or color) of the plurality of strokes (421) input by the user. For example, the electronic device (101) can display text data (423) by applying handwriting. For example, the electronic device (101) can display text data (423) by applying the user's handwriting.
[0109] In one embodiment, the electronic device (101) may convert a plurality of strokes (421) input by a user into a designated font supported by the electronic device (101) and display them as text data (423). For example, in the text data (423) converted into a font, lines constituting letters of the font may be mapped to be associated with each of the plurality of strokes (421) input by the user.
[0110] In one embodiment, the plurality of strokes (421) may include a user-drawn image (e.g., a drawing object). For example, the drawing object may be displayed around text data (423). For example, the drawing object may be displayed while maintaining the shape of the plurality of strokes (421) input by the user. For example, the drawing object may be converted and displayed through the generative AI module (220). For example, the generative AI module (220) may generate an image or video from the drawing object based on text surrounding the drawing object and / or audio data that is semantically and / or temporally associated with the drawing object. For example, lines included in the image or video converted through the generative AI module (220) may be mapped to be associated with each of the plurality of strokes (421) input by the user. For example, one or more lines constituting the drawing object may be temporally and / or semantically associated with audio data (410). For example, one or more lines constituting a drawing object may be temporally and / or semantically linked to text data (413). For example, as audio data (413) is played, one or more lines constituting an image generated by a drawing object or AI module may be visually highlighted and displayed.
[0111] In one embodiment, at least some of the plurality of strokes (421) may be acquired while the plurality of utterances (411) are acquired (or while the audio data (410) is recorded). In one embodiment, at least some of the plurality of strokes (421) may be input while being superimposed on the content (331) of the document (420). For example, in one embodiment, at least some of the plurality of strokes (421) may be input at a location where the content (331) of the document (420) is displayed.
[0112] In one embodiment, the electronic device (101) can detect user input (e.g., touch input, pen input, stylus pen input, finger input, gesture input, and / or segment input) through the display (260). In one embodiment, the electronic device (101) can perform an input command according to the detected user input.
[0113] For example, when a drawing input (or sketch, drawing) is detected based on a user input (e.g., touch input, pen input, stylus pen input, finger input, gesture input, and / or segment input), the electronic device (101) may generate points, lines, and images corresponding to the drawing input, and display the generated points, lines, and images through the display (260). The electronic device (101) may generate a drawing object according to the drawing input. In one embodiment, the drawing object may represent points, lines, and images. In one embodiment, the drawing input (or sketch, drawing) may be analyzed by the generative AI module (220). In one embodiment, an object represented by the drawing input (or sketch, drawing) may be identified by the analysis of the generative AI module (220). In one embodiment, for analysis by the generative AI module (220), a drawing input (or sketch, drawing), texts surrounding the drawing input (or sketch, drawing), and / or content (e.g., images, videos) associated with a user account may be input as prompts to the generative AI module (220).
[0114] For example, when a segment input is detected, the electronic device (101) may monitor the start time when the segment input starts, the end time when the segment input ends, and / or the movement path (e.g., movement, movement gesture) from the start time to the end time of the segment input. The electronic device (101) may determine whether the segment input corresponds to a set character (e.g., letter, text), a symbol (e.g., asterisk, underline, emphasis), and / or a command (e.g., format, copy, paste). In one embodiment, when the segment input corresponds to a set input, the electronic device (101) may generate a character and / or symbol corresponding to the segment input. The electronic device (101) may extract the character and / or symbol corresponding to the segment input and provide it to the user. For example, the electronic device (101) may display the character and / or symbol corresponding to the segment input through the display (260). In one embodiment, the electronic device (101) may perform a command (e.g., formatting, copying, pasting) corresponding to a segment input when the segment input corresponds to a set command. Hereinafter, a single character, symbol, or command identified through handwriting input (e.g., multiple strokes (421)) by a user input may be referred to as an object.
[0115] In one embodiment, the electronic device (101) can obtain a summary (415) from text data (413) using the generative AI module (220). For example, the electronic device (101) can generate a summary of one sentence by inputting sentences containing characters within a specified length (e.g., 250 words or 5,000 words) among the sentences in the text data (413) into the generative AI module (220). For example, referring to FIG. 5A, the electronic device (101) can generate one sentence (561) of the summary (415) by inputting a plurality of sentences (551) of the text data (413) into the generative AI module (220). For example, the electronic device (101) can generate one sentence (563) of the summary (415) by inputting multiple sentences (553) of text data (413) into the generative AI module (220). However, the present invention is not limited thereto. For example, the electronic device (101) can generate multiple sentences (561, 563) by inputting text data (413) into the generative AI module (220).
[0116] In one embodiment, the electronic device (101) can obtain a summary (425) from text data (423) using the generative AI module (220). For example, the electronic device (101) can generate a summary of one sentence by inputting sentences containing characters within a specified length (e.g., 250 words or 5,000 words) among the sentences in the text data (423) into the generative AI module (220). For example, referring to FIG. 5B, the electronic device (101) can generate one sentence (581) of the summary (425) by inputting a plurality of sentences (571) of the text data (423) into the generative AI module (220). For example, the electronic device (101) can generate one sentence (583) of a summary (425) by inputting multiple sentences (573) of text data (423) into the generative AI module (220).
[0117] Referring to FIG. 5A, the electronic device (101) can identify one or more words (511 to 526) in text data (413) through the linking module (230). For example, the electronic device (101) can identify a word of a specified format in the text data (413). For example, the word of the specified format may be a noun. However, the present invention is not limited thereto. For example, the word of the specified format may be an adjective or a verb. In one embodiment, the word of the specified format may be referred to as a keyword.
[0118] In one embodiment, the electronic device (101) may identify, through the linking module (230), information about one or more words (511 to 526) identified in text data (413). For example, the information about the one or more words (511 to 526) may include information (e.g., a sequence) indicating a location at which the one or more words (511 to 526) are identified. For example, the information about the one or more words (511 to 526) may include a timestamp at which the one or more words (511 to 526) are identified in the audio data (410). For example, referring to table (601) of FIG. 6A, the electronic device (101) may generate, through the linking module (230), data representing information about one or more words (511 to 526) identified in the text data (413). For example, the electronic device (101) may, through the linking module (230), generate a table that maps one or more words (511 to 526) to a sequence indicating the location of one or more words (511 to 526) and a timestamp at which one or more words (511 to 526) were identified.
[0119] Referring to FIG. 5B, the electronic device (101) can identify one or more words in text data (423) through the linking module (230). For example, the electronic device (101) can identify a word of a specified format in the text data (423). For example, the word of the specified format may be a noun. However, the present invention is not limited thereto. For example, the word of the specified format may be an adjective or a verb. According to an embodiment, the electronic device (101) can identify one or more words separated by a specified character (e.g., a blank space (or space)) in the text data (423).
[0120] In one embodiment, the electronic device (101) may, through the linking module (230), identify information about one or more words identified in text data (423). For example, the information about the one or more words may include information (e.g., a sequence) indicating the location at which the one or more words are identified. For example, the information about the one or more words may include a timestamp at which the one or more words were entered into the document (420). For example, referring to table (603) of FIG. 6A, the electronic device (101) may, through the linking module (230), generate data indicating information about one or more words identified in the text data (423). For example, the electronic device (101) may, through the linking module (230), generate a table that maps one or more words to a sequence indicating the location of the one or more words and a timestamp at which the one or more words are identified.
[0121] In one embodiment, referring to FIG. 5A, the electronic device (101) can identify one or more words (531 to 540) in the summary (415). For example, the electronic device (101) can identify a word of a specified format in the summary (415). In one embodiment, referring to FIG. 5B, the electronic device (101) can identify one or more words in the summary (425). For example, the electronic device (101) can identify a word of a specified format in the summary (425). For example, the word of the specified format may be a noun, but is not limited thereto. For example, the word of the specified format may be an adjective or a verb.
[0122] In one embodiment, the electronic device (101) may generate relationship information (430) for audio data (410) and / or documents (420) via the linking module (230).
[0123] In one embodiment, the electronic device (101) may generate relationship information (431) indicating a relationship between text data (413) and text data (423) through the linking module (230). For example, the electronic device (101) may generate relationship information (431) that maps words identified with the same timestamp in text data (413) and text data (423) to each other through the linking module (230). For example, the electronic device (101) may generate relationship information (431) that maps words identified with the same and / or similar timestamps in text data (413) and text data (423) to each other during note writing through the linking module (230). For example, referring to FIG. 6A, the electronic device (101) can map "20th century", which is identified by the same timestamp as "4th industrial revolution" in the text data (413). For example, referring to FIG. 6A, the electronic device (101) can map "electrical and electronic equipment", which is identified by the same timestamp as "traditional industry" in the text data (413).
[0124] In one embodiment, the electronic device (101) may generate relationship information (433) indicating a relationship between text data (413) and a summary (415) through the linking module (230). For example, the electronic device (101) may generate relationship information (433) that maps corresponding words in the text data (413) and the summary (415) to each other through the linking module (230). For example, the electronic device (101) may generate relationship information (433) that maps corresponding words in sentences (551, 553) in the text data (413) and one sentence (561, 563) in the summary (415) that summarizes the sentences (551, 553) in the text data (413) to each other through the linking module (230). For example, the electronic device (101) can map words in the text data (413) and words in the summary (415) that have the same and / or similar meanings to each other through the linking module (230). For example, the electronic device (101) can map words included in a unit sentence input to the generative AI module (220) and words included in a unit sentence generated from the generative AI module (220) to each other through the linking module (230). For example, the electronic device (101) can map words (511 to 526) in a sentence (551) and words (531 to 540) in a sentence (561) to each other. For example, the electronic device (101) can map words (511 to 526) within a sentence (551) and words (531 to 540) within a sentence (561) to each other based on sequence and / or meaning.
[0125] For example, referring to FIG. 6B, the electronic device (101) may, through the linking module (230), generate relationship information (433) that maps words (e.g., words in sequences 1 to 6 of table (601)) identified in sentences (551) in text data (413) to words (e.g., words in sequences 1 to 6 of table (611)) identified in one sentence (561) in a summary (415) summarizing sentences (551). For example, referring to FIG. 6B, the electronic device (101) may generate, through the linking module (230), relationship information (433) that maps words (e.g., words in sequences 49 to 50 of table (601)) identified in sentences (553) in text data (413) to words (e.g., words in sequences 7 to 8 of table (611)) identified in one sentence (563) in a summary (415) summarizing sentences (553). For example, the electronic device (101) may generate relationship information (433) indicating that a word in a first sequence of table (601) (“4th Industrial Revolution”) and a word in a first sequence of table (611) (“4th Industrial Revolution”) correspond to each other. For example, the electronic device (101) can generate relationship information (433) indicating that the word (“artificial intelligence”) of the 50th sequence of the table (601) and the word (“artificial intelligence”) of the 8th sequence of the table (611) correspond to each other.
[0126] In one embodiment, the electronic device (101) may generate relationship information (435) indicating a relationship between text data (423) and a summary (425) through the linking module (230). For example, the electronic device (101) may generate relationship information (435) that maps corresponding words in the text data (423) and the summary (425) to each other through the linking module (230). For example, the electronic device (101) may generate relationship information (435) that maps corresponding words in sentences (571, 573) in the text data (423) and one sentence (581, 583) in the summary (425) that summarizes the sentences (571, 573) in the text data (423) to each other through the linking module (230). For example, the electronic device (101) can map words in the text data (423) and words in the summary (425) that have the same and / or similar meanings to each other through the linking module (230). For example, the electronic device (101) can map words included in a unit sentence input to the generative AI module (220) and words included in a unit sentence generated from the generative AI module (220) to each other through the linking module (230). For example, the electronic device (101) can map words in a sentence (571) and words in a sentence (581) to each other. For example, the electronic device (101) can map words in a sentence (571) and words in a sentence (581) to each other based on sequence and / or meaning.
[0127] For example, referring to FIG. 6c, the electronic device (101) may generate relationship information (435) that maps words (e.g., words in sequences 1, 2, 4, and 5 of table (603)) identified in sentences (571) in text data (423) to words (e.g., words in sequences 1 and 2 of table (613)) identified in one sentence (581) in a summary (425) summarizing sentences (571) through the linking module (230). For example, referring to FIG. 6C, the electronic device (101) may generate relationship information (435) that maps words (e.g., words in sequences 23, 24, 26, 27, 29, and 32 of table (603)) identified in sentences (573) in text data (423) to words (e.g., words in sequences 11 to 15 of table (613)) identified in one sentence (583) in a summary (425) summarizing sentences (573) via the linking module (230). In one embodiment, the electronic device (101) may generate relationship information (435) between tables (603, 613). For example, the electronic device (101) can generate relationship information (435) indicating that the words of the first and second sequences of the table (603) (“1st to 2nd,” “Industrial Revolution”) correspond to each other and the words of the first sequence of the table (613) (“1st to 2nd Industrial Revolution”). For example, the electronic device (101) can generate relationship information (435) indicating that the words of the 32nd sequence of the table (603) (“Digital Infrastructure”) correspond to the words of the 15th sequence of the table (613) (“Digital Infrastructure”).
[0128] According to an embodiment, the electronic device (101) may generate text data (423) and / or a summary (425) in real time based on obtaining text input (e.g., handwriting input and typing input). Similarly, the electronic device (101) may generate text data (413) and / or a summary (415) in real time based on obtaining voice input. According to an embodiment, the electronic device (101) may generate text data (423), a summary (425), text data (413) and / or a summary (415) after note writing is completed (e.g., after document (420) and audio data (410) are stored).
[0129] In FIG. 4, relationship information (430) is illustrated as including relationship information (431) between text data (413, 423), relationship information (433) between text data (413) and a summary (415), and relationship information (435) between text data (423) and a summary (425), but this is only an example.
[0130] For example, the electronic device (101) may further include relationship information between the text data (413) and the summary (425). For example, the electronic device (101) may generate relationship information between the text data (423) and the summary (415) based on relationship information (431) between the text data (413, 423) and the text data (423) and the summary (425) and relationship information (435) between the text data (423) and the summary (425). For example, the electronic device (101) may generate relationship information between the text data (413) and the summary (425) by using relationship information (431) indicating a relationship between the timestamp information of the text data (413) and the timestamp information of the text data (423) and relationship information (435) indicating a relationship between a sequence of the text data (423) and a sequence of the summary (425). For example, considering that “the 20th century” of the text data (413) is mapped to “the 4th industrial revolution” of the text data (423), and that “the 4th industrial revolution” of the text data (423) is mapped to “the 4th industrial revolution” of the summary (425), the electronic device (101) can generate relationship information between the text data (413) and the summary (425) that maps “the 20th century” of the text data (413) and “the 4th industrial revolution” of the summary (425).
[0131] For example, the electronic device (101) may further include relationship information between the text data (423) and the summary (415). For example, the electronic device (101) may generate relationship information between the text data (423) and the summary (415) based on relationship information (431) between the text data (413, 423) and the text data (413) and relationship information (433) between the text data (413) and the summary (415). For example, the electronic device (101) may generate relationship information between the text data (423) and the summary (415) by using relationship information (431) indicating a relationship between the timestamp information of the text data (423) and the timestamp information of the text data (413) and relationship information (433) indicating a relationship between a sequence of the text data (413) and a sequence of the summary (415). For example, considering that the “4th industrial revolution” of the text data (423) is mapped to the “20th century” of the text data (413), and that the “20th century” of the text data (413) is mapped to the “20th century” of the summary (415), the electronic device (101) can generate relationship information between the text data (413) and the summary (425) that maps the “4th industrial revolution” of the text data (423) and the “20th century” of the summary (415).
[0132] For example, the electronic device (101) may further include relationship information between the summary (415) and the summary (425). For example, the electronic device (101) may generate relationship information between the summary (415) and the summary (425) based on relationship information (431) between text data (413, 423), relationship information (433) between text data (413) and the summary (415), and relationship information (435) between text data (423) and the summary (425). For example, considering that “the 20th century” of the text data (413) is mapped to “the 4th industrial revolution” of the text data (423), “the 4th industrial revolution” of the text data (423) is mapped to “the 4th industrial revolution” of the summary (425), and “the 20th century” of the text data (413) is mapped to “the 20th century” of the summary (415), the electronic device (101) can generate relationship information between the summary (415) and the summary (425) that maps “the 4th industrial revolution” of the summary (425) and “the 20th century” of the summary (415).
[0133] Hereinafter, relationship information (430) may be exemplified as utilizing not only relationship information (431, 433, 435), but also various relationship information between text data (413), text data (423), summary (415), and / or summary (425). Hereinafter, various relationship information between text data (413), text data (423), summary (415), and / or summary (425) may be referred to as relationship information (430).
[0134] In FIG. 4, it is illustrated that relationship information (430) is generated based on a sequence of timestamps and / or keywords, but this is merely an example. According to an embodiment, the electronic device (101) may generate relationship information (430) based on the meaning between keywords. For example, if the timestamps of keywords with the same and / or similar meanings are similar (e.g., keywords identified within a specified time range), the electronic device (101) may map and manage the corresponding keywords within the relationship information (430).
[0135] Hereinafter, with reference to FIGS. 7a to 8c, an operation of an electronic device (101) providing a processed note to a user is described.
[0136] FIG. 7A is a diagram illustrating an example of a screen of note content displaying text data converted from strokes and audio data according to an embodiment. FIG. 7B is a diagram illustrating an example of a screen of note content displaying note content including strokes and summaries according to an embodiment. FIG. 7C is a diagram illustrating an example of a screen of note content displaying strokes and a summary of text data converted from strokes according to an embodiment. FIG. 7D is a diagram illustrating an example of a screen of note content displaying a summary of text data converted from strokes and audio data according to an embodiment. FIG. 7E is a diagram illustrating an example of a screen of note content displaying text data converted from audio data and summaries according to an embodiment. FIG. 7F is a diagram illustrating an example of a screen of note content displaying text data converted from audio data and a summary of text data converted from strokes according to an embodiment. FIG. 7g is a diagram illustrating an example of a screen of note content that displays text data converted from audio data and a summary of the text data converted from audio data, according to one embodiment. FIG. 7h is a diagram illustrating an example of a screen of note content that displays summaries, according to one embodiment.
[0137] FIG. 7a may be described with reference to the electronic device (101) (or components of the electronic device (101)) of FIG. 2.
[0138] Referring to FIG. 7A, the electronic device (101) can display a screen (701) of a content service APP (210) through a display (260). For example, in response to an execution request (or a viewing request) for a note (e.g., a file including a document (420) and audio data (410)), the electronic device (101) can display a screen (701) of the content service APP (210) that displays the note.
[0139] In one embodiment, the screen (701) may be a screen divided into a plurality of regions (710, 730). The plurality of regions (710, 730) may be arranged left and right within the screen (701). For example, the region (710) may be arranged to the left of the region (730). The plurality of regions (710, 730) may be arranged so as not to overlap each other within the screen (701). In one embodiment, the size and / or position of the plurality of regions (710, 730) may be adjusted according to a user input. For example, the electronic device (101) may enlarge the region (710) to the right (or enlarge in the direction of the region (730)) or enlarge the region (730) to the left (or enlarge in the direction of the region (710)) based on the user input. For example, the electronic device (101) can arrange multiple areas (710, 730) vertically based on user input.
[0140] In one embodiment, the region (710) may display text data (423) composed of a plurality of strokes (421). In one embodiment, the region (730) may display text data (413) converted from audio data (410). In one embodiment, the region (730) may display one or more visual objects (731, 733, 735). For example, the visual object (731) may indicate that the text data (413) is the original content that has not been summarized. For example, the visual object (733) may be a visual object for requesting a translation of the text data (413). For example, the visual object (735) may be a visual object representing the speaker of the text data (413). In one embodiment, the speaker of the text data (413) may be distinguished based on inputting a plurality of utterances (411) into the generative AI module (220). For example, the speaker of the text data (413) may be distinguished through the generative AI module (220) based on the tone and / or pronunciation characteristics of the plurality of utterances (411).
[0141] In one embodiment, text data (423) composed of a plurality of strokes (421) displayed in an area (710) and text data (413) converted from audio data (410) displayed in an area (730) may correspond to each other. In one embodiment, the text data (423) may be data for which relationship information is formed with the audio data (410) after the document (420) is stored. In one embodiment, the plurality of strokes (421) may mean raw data of strokes input by a user before the document (420) is stored.
[0142] For example, the electronic device (101) may display text data (423) and text data (413) in regions (710, 730) so that they are synchronized based on the relationship information (430). For example, the electronic device (101) may display text data (413) that is temporally and / or semantically mapped to text data (423) displayed in region (710) in region (730) based on the relationship information (430).
[0143] For example, the electronic device (101) may display words of text data (413) that are temporally and / or semantically mapped to words of text data (423) displayed in the area (710) based on the relationship information (430), in the area (730). For example, referring to FIG. 6A, when displaying “20th century” (741) of text data (423) in the area (710), the electronic device (101) may display “4th industrial revolution” (743) of text data (413) corresponding to “20th century” (741) of text data (423) in the area (730). For example, the electronic device (101) may display paragraphs (or sentences) including words of the text data (413) that are temporally and / or semantically mapped to words of the text data (423) displayed in the area (710) based on the relationship information (430), in the area (730). For example, when displaying “20th century” of the text data (423) in the area (710), the electronic device (101) may display a paragraph (or sentence) including “4th industrial revolution” of the text data (413) corresponding to “20th century” of the text data (423) in the area (730).
[0144] In one embodiment, when text data (423) displayed in an area (710) is scrolled (or moved), the electronic device (101) can scroll (or move) text data (413) displayed in an area (730) in accordance with the text data (423) displayed in the area (710) according to the scrolling (or moving). In one embodiment, when text data (413) displayed in an area (730) is scrolled (or moved), the electronic device (101) can scroll (or move) text data (423) displayed in an area (710) in accordance with the text data (413) displayed in the area (730).
[0145] In FIG. 7A, the screen (701) is illustrated as being divided into a plurality of regions (710, 730), but this is merely an example. According to an embodiment, among the plurality of regions (710, 730), the region (730) may be a screen that overlaps the region (710). For example, while the electronic device (101) displays the screen (701) displaying the region (710), the electronic device (101) may display the region (730) by overlapping it over the region (710) based on a user input requesting the display of text data (423). For example, while the electronic device (101) displays the screen (701) displaying the region (710), the electronic device (101) may divide the screen (701) into the region (710) and the region (730) based on a user input requesting the display of text data (423).
[0146] Referring to FIG. 7b, the electronic device (101) can display a screen (702) of a content service APP (210) through a display (260).
[0147] In one embodiment, the screen (702) may be a screen divided into a plurality of regions (710, 730). The plurality of regions (710, 730) may be arranged left and right within the screen (702). For example, the region (710) may be arranged to the left of the region (730). The plurality of regions (710, 730) may be arranged so as not to overlap each other within the screen (702).
[0148] In one embodiment, the region (710) may display text data (423) composed of strokes. In one embodiment, the region (730) may display summaries (415, 425) of text data converted from strokes and audio data. For example, the summaries (415, 425) may be arranged vertically within the region (730). For example, the summary (425) may be arranged above the summary (415) within the region (730). The summaries (415, 425) may be arranged so as not to overlap each other within the screen (702).
[0149] In one embodiment, the area (730) may display one or more visual objects (751, 753, 755). For example, the visual object (751) may indicate that the summaries (415, 425) are content summarized from the text data (413, 423). For example, the visual object (753) may be a visual object indicating that it is a summary of the text data (423). For example, the visual object (755) may be a visual object indicating that it is a summary of the text data (413).
[0150] For example, the electronic device (101) may display text data (423) and summaries (415, 425) in regions (710, 730) so that they are synchronized based on relationship information (430). For example, the electronic device (101) may display summaries (415, 425) that are temporally and / or semantically mapped to text data (423) displayed in region (710) in region (730) based on relationship information (430).
[0151] For example, the electronic device (101) may display words of summaries (415, 425) that are temporally and / or semantically mapped to words of text data (423) displayed in the area (710) based on the relationship information (430), in the area (730). For example, referring to FIG. 6C, when displaying “20th century” (741) of text data (423) in the area (710), the electronic device (101) may display “20th century” (745) of the summary (425) corresponding to “20th century” (741) of text data (423) in the area (730). For example, referring to FIGS. 6A and 6C, when the electronic device (101) displays “20th century” (741) of text data (423) in an area (710), it may display “4th industrial revolution” (747) of the summary (415) corresponding to “4th industrial revolution” of text data (413) corresponding to “20th century” (741) of text data (423) in an area (730).
[0152] In one embodiment, when text data (423) displayed in an area (710) is scrolled (or moved), the electronic device (101) can scroll (or move) the summaries (415, 425) displayed in an area (730) according to the text data (423) displayed in the area (710) according to the scrolling (or moving). In one embodiment, when the summaries (415, 425) displayed in an area (730) are scrolled (or moved), the electronic device (101) can scroll (or move) the text data (423) displayed in an area (710) according to the summaries (415, 425) displayed in the area (730) according to the scrolling (or moving).
[0153] In FIG. 7B, the screen (702) is illustrated as being divided into a plurality of regions (710, 730), but this is merely an example. Depending on the embodiment, the region (730) among the plurality of regions (710, 730) may be a screen that overlaps the region (710). For example, while the electronic device (101) displays the screen (702) displaying the region (710), the electronic device (101) may display the region (730) by overlapping it over the region (710) based on a user input requesting the display of the summaries (415, 425). For example, while the electronic device (101) displays the screen (702) displaying the region (710), the electronic device (101) may divide the screen (702) into the region (710) and the region (730) based on a user input requesting the display of the summaries (415, 425).
[0154] Referring to FIG. 7c, the electronic device (101) can display the screen (703) of the content service APP (210) through the display (260).
[0155] In one embodiment, the screen (703) may be a screen divided into a plurality of regions (710, 730). The plurality of regions (710, 730) may be arranged left and right within the screen (703). For example, the region (710) may be arranged to the left of the region (730). The plurality of regions (710, 730) may be arranged so as not to overlap each other within the screen (703).
[0156] In one embodiment, the area (710) may display text data (423) composed of strokes. In one embodiment, the area (730) may display a summary (425) of text data converted from strokes.
[0157] In one embodiment, the area (730) may display one or more visual objects (751, 753). For example, the visual object (751) may indicate that the summary (425) is content summarized from text data (423). For example, the visual object (753) may be a visual object indicating that it is a summary of text data (423).
[0158] For example, the electronic device (101) may display text data (423) and a summary (425) in regions (710, 730) so that they are synchronized based on the relationship information (430). For example, the electronic device (101) may display a summary (425) that is temporally and / or semantically mapped to the text data (423) displayed in the region (710) in the region (730), based on the relationship information (430).
[0159] For example, the electronic device (101) may display, in the area (730), words of the summary (425) that are temporally and / or semantically mapped to words of the text data (423) displayed in the area (710) based on the relationship information (430). For example, referring to FIG. 6C, when the electronic device (101) displays “20th century” of the text data (423) in the area (710), the electronic device (101) may display “20th century” of the summary (425) corresponding to “20th century” of the text data (423) in the area (730).
[0160] In one embodiment, when text data (423) displayed in an area (710) is scrolled (or moved), the electronic device (101) can scroll (or move) the summary (425) displayed in an area (730) to match the text data (423) displayed in the area (710) according to the scrolling (or moving). In one embodiment, when the summary (425) displayed in an area (730) is scrolled (or moved), the electronic device (101) can scroll (or move) the text data (423) displayed in an area (710) to match the summary (425) displayed in the area (730) according to the scrolling (or moving).
[0161] In FIG. 7C, the screen (703) is illustrated as being divided into a plurality of regions (710, 730), but this is merely an example. Depending on the embodiment, the region (730) among the plurality of regions (710, 730) may be a screen that overlaps the region (710). For example, while the electronic device (101) displays the screen (703) displaying the region (710), the electronic device (101) may display the region (730) by overlapping it over the region (710) based on a user input requesting the display of the summary (425). For example, while the electronic device (101) displays the screen (703) displaying the region (710), the electronic device (101) may divide the screen (703) into the regions (710) and (730) based on a user input requesting the display of the summary (425).
[0162] Referring to FIG. 7d, the electronic device (101) can display the screen (704) of the content service APP (210) through the display (260).
[0163] FIG. 7d illustrates that, compared to FIG. 7c, a summary (415) may be displayed in the region (730) instead of the summary (425). In one embodiment, one or more visual objects (751, 755) may be displayed in the region (730). For example, the visual object (751) may indicate that the summary (415) is content summarized from text data (413). For example, the visual object (755) may be a visual object indicating that it is a summary of text data (413).
[0164] In one embodiment, the screen (704) may be a screen divided into a plurality of regions (710, 730). The plurality of regions (710, 730) may be arranged left and right within the screen (704). For example, the region (710) may be arranged to the left of the region (730). The plurality of regions (710, 730) may be arranged so as not to overlap each other within the screen (704).
[0165] In one embodiment, the area (710) may display text data (423) composed of strokes. In one embodiment, the area (730) may display a summary (415) of text data converted from audio data.
[0166] In one embodiment, the area (730) may display one or more visual objects (751, 753). For example, the visual object (751) may indicate that the summary (415) is content summarized from the text data (413). For example, the visual object (755) may be a visual object indicating that it is a summary of the text data (413).
[0167] For example, the electronic device (101) may display text data (423) and a summary (415) in regions (710, 730) so that they are synchronized based on the relationship information (430). For example, the electronic device (101) may display a summary (415) that is temporally and / or semantically mapped to the text data (423) displayed in the region (710) in the region (730), based on the relationship information (430).
[0168] For example, the electronic device (101) may display, in the area (730), words of the summary (415) that are temporally and / or semantically mapped to words of the text data (423) displayed in the area (710) based on the relationship information (430). For example, referring to FIG. 6C, when the electronic device (101) displays “20th century” of the text data (423) in the area (710), the electronic device (101) may display “4th Industrial Revolution” of the summary (415) corresponding to “20th century” of the text data (423) in the area (730).
[0169] In one embodiment, when text data (423) displayed in an area (710) is scrolled (or moved), the electronic device (101) can scroll (or move) the summary (415) displayed in an area (730) to match the text data (423) displayed in the area (710) according to the scrolling (or moving). In one embodiment, when the summary (415) displayed in an area (730) is scrolled (or moved), the electronic device (101) can scroll (or move) the text data (423) displayed in an area (710) to match the summary (415) displayed in the area (730) according to the scrolling (or moving).
[0170] In FIG. 7D, the screen (704) is illustrated as being divided into a plurality of regions (710, 730), but this is merely an example. Depending on the embodiment, the region (730) among the plurality of regions (710, 730) may be a screen that overlaps the region (710). For example, while the electronic device (101) displays the screen (704) displaying the region (710), the electronic device (101) may display the region (730) by overlapping it over the region (710) based on a user input requesting the display of the summary (415). For example, while the electronic device (101) displays the screen (704) displaying the region (710), the electronic device (101) may divide the screen (704) into the regions (710) and (730) based on a user input requesting the display of the summary (415).
[0171] Referring to FIG. 7e, the electronic device (101) can display a screen (705) of a content service APP (210) through a display (260). Compared to FIG. 7b, FIG. 7e can display text data (413) converted from audio data instead of text data (423) composed of strokes in an area (710). Descriptions of FIG. 7e that overlap with those of FIG. 7b may not be repeated.
[0172] For example, the electronic device (101) may display text data (413) and summaries (415, 425) in regions (710, 730) so that they are synchronized based on the relationship information (430). For example, the electronic device (101) may display summaries (415, 425) that are temporally and / or semantically mapped to text data (413) displayed in region (710) in region (730) based on the relationship information (430). For example, the electronic device (101) may display words of summaries (415, 425) that are temporally and / or semantically mapped to words of text data (413) displayed in region (710) in region (730) based on the relationship information (430).
[0173] In one embodiment, when text data (413) displayed in an area (710) is scrolled (or moved), the electronic device (101) can scroll (or move) the summaries (415, 425) displayed in an area (730) according to the text data (413) displayed in the area (710) according to the scrolling (or moving). In one embodiment, when the summaries (415, 425) displayed in an area (730) are scrolled (or moved), the electronic device (101) can scroll (or move) the text data (413) displayed in an area (710) according to the summaries (415, 425) displayed in the area (730).
[0174] Referring to FIG. 7F, the electronic device (101) can display a screen (706) of a content service APP (210) through a display (260). Compared to FIG. 7C, FIG. 7F can display text data (413) converted from audio data instead of text data (423) composed of strokes in an area (710). Descriptions of FIG. 7F that overlap with those of FIG. 7C may not be repeated.
[0175] For example, the electronic device (101) may display text data (413) and a summary (425) in regions (710, 730) so that they are synchronized based on the relationship information (430). For example, the electronic device (101) may display a summary (425) that is temporally and / or semantically mapped to text data (413) displayed in region (710) in region (730) based on the relationship information (430). For example, the electronic device (101) may display words of the summary (425) that are temporally and / or semantically mapped to words of the text data (413) displayed in region (710) in region (730) based on the relationship information (430).
[0176] In one embodiment, when text data (413) displayed in an area (710) is scrolled (or moved), the electronic device (101) can scroll (or move) the summary (425) displayed in an area (730) to match the text data (413) displayed in an area (710) according to the scrolling (or moving). In one embodiment, when the summary (425) displayed in an area (730) is scrolled (or moved), the electronic device (101) can scroll (or move) the text data (413) displayed in an area (710) to match the summary (425) displayed in an area (730) according to the scrolling (or moving).
[0177] Referring to FIG. 7g, the electronic device (101) can display a screen (707) of a content service APP (210) through a display (260). Compared to FIG. 7d, FIG. 7g can display text data (413) converted from audio data instead of text data (423) composed of strokes in an area (710). Descriptions of FIG. 7g that overlap with those of FIG. 7d may not be repeated.
[0178] For example, the electronic device (101) may display text data (413) and a summary (415) in regions (710, 730) so that they are synchronized based on the relationship information (430). For example, the electronic device (101) may display a summary (415) that is temporally and / or semantically mapped to text data (413) displayed in region (710) in region (730) based on the relationship information (430). For example, the electronic device (101) may display words of the summary (415) that are temporally and / or semantically mapped to words of the text data (413) displayed in region (710) in region (730) based on the relationship information (430).
[0179] In one embodiment, when text data (413) displayed in an area (710) is scrolled (or moved), the electronic device (101) can scroll (or move) the summary (415) displayed in an area (730) according to the text data (413) displayed in the area (710) according to the scrolling (or moving). In one embodiment, when the summary (415) displayed in an area (730) is scrolled (or moved), the electronic device (101) can scroll (or move) the text data (413) displayed in an area (710) according to the summary (415) displayed in an area (710) according to the scrolling (or moving).
[0180] Referring to FIG. 7h, the electronic device (101) can display a screen (708) of a content service APP (210) through a display (260). Compared to FIG. 7g, FIG. 7h can display a summary (425) of text data (423) composed of strokes instead of text data (413) converted from audio data in an area (710). Descriptions of FIG. 7h that overlap with those of FIG. 7g may not be repeated.
[0181] For example, the electronic device (101) may display the summary (425) and the summary (415) in the areas (710, 730) so that they are synchronized based on the relationship information (430). For example, the electronic device (101) may display the summary (415) in the area (730) that is temporally and / or semantically mapped to the text data (413) displayed in the area (710) based on the relationship information. For example, the electronic device (101) may display the words of the summary (415) that are temporally and / or semantically mapped to the words of the summary (425) displayed in the area (710) based on the relationship information.
[0182] In one embodiment, when the summary (425) displayed in the area (710) is scrolled (or moved), the electronic device (101) can scroll (or move) the summary (415) displayed in the area (730) to match the summary (425) displayed in the area (710) according to the scrolling (or moving). In one embodiment, when the summary (415) displayed in the area (730) is scrolled (or moved), the electronic device (101) can scroll (or move) the summary (425) displayed in the area (710) to match the summary (415) displayed in the area (730) according to the scrolling (or moving).
[0183] As described above, the electronic device (101) can summarize and provide various data to the user. Therefore, users using the content service APP (210) can more easily identify the connection between their handwritten data and recorded voice, and this connection can function as an index in the vast amount of note content.
[0184] FIG. 8A is a diagram illustrating an example of a screen of note content in which an electronic device highlights and displays another object corresponding to a selected object, according to one embodiment.
[0185] Fig. 8a, compared to Fig. 7b, may illustrate a situation in which an electronic device (101) receives a user input (810). Descriptions of Fig. 8a that overlap with those of Fig. 7b may not be repeated.
[0186] Referring to the screen (801) of FIG. 8A, the electronic device (101) may receive a user input (810). In one embodiment, the user input (810) may be an input for selecting at least one word (or sentence). In one embodiment, the user input (810) may be a touch input (e.g., a drag input) and / or a hovering input.
[0187] In one embodiment, the electronic device (101) may identify at least one other word (or sentence) corresponding to the selected at least one word (or sentence) based on a user input (810) selecting at least one word (or sentence). For example, the electronic device (101) may identify at least one other word (or sentence) in another content based on at least one word (or sentence) selected in one content. In one embodiment, the electronic device (101) may identify at least one other word (or sentence) in another content corresponding to the selected at least one word (or sentence) in one content based on relationship information (430).
[0188] In one embodiment, the electronic device (101) may identify at least one other word (or sentence) in another content based on a sequence assigned to at least one selected word (or sentence) in one content. For example, referring to FIG. 8A, the electronic device (101) may identify at least one other word (or sentence) in another content (e.g., text data (423) and / or summary (425)) based on keywords (e.g., “4th industrial revolution,” “traditional industry,” “business,” “paradigm,” “economic society in general,” and “digital transformation speed”) to which sequences are assigned among words included in sentences selected in the summary (415).
[0189] For example, the electronic device (101) can identify keywords (e.g., “20th century,” “electrical and electronic equipment,” “development,” “computer,” “Internet,” and “emergence”) among words in the text data (423) that have corresponding timestamp information based on keywords (e.g., “4th industrial revolution,” “traditional industry,” “business,” “paradigm,” “economic society in general,” and “digital transformation speed”) to which sequences are assigned among words included in the sentences selected in the summary (415). For example, the electronic device (101) can identify keywords (e.g., “4th industrial revolution,” “traditional industry,” “business,” “paradigm,” “economic society in general,” and “digital transformation speed”) to which sequences are assigned among the words included in the sentences selected in the summary (415) within the text data (413) related to the summary (415) within the text data (413). The keywords (e.g., “4th industrial revolution,” “traditional industry,” “business,” “paradigm,” “economic society in general,” and “digital transformation speed”) to which sequences are assigned are identified. For example, the electronic device (101) can identify keywords (e.g., “20th century,” “electrical and electronic equipment,” “development,” “computer,” “Internet,” and “emergence”) in text data (423) that have timestamp information corresponding to keywords (e.g., “4th industrial revolution,” “traditional industry,” “business,” “paradigm,” “economic society in general,” and “digital transformation speed”) assigned the same and / or similar sequences in text data (413).
[0190] In one embodiment, the electronic device (101) may highlight and display at least one selected word (or sentence) (820) based on a user input (810) selecting at least one word (or sentence). In one embodiment, the electronic device (101) may highlight and display at least one selected word (or sentence) (820) based on a user input (810) selecting at least one word (or sentence) (820), while highlighting and displaying at least one other word (or sentence) (830) in text data (423) having timestamp information corresponding to the selected at least one word (or sentence) (820). According to an embodiment, the electronic device (101) may highlight and display at least one other word (or sentence) (830) in the text data (423) having timestamp information corresponding to at least one word (or sentence) (820) selected by a user input (810), based on the fact that at least one other word (or sentence) (830) in the text data (423) having timestamp information corresponding to at least one word (or sentence) (820) selected by a user input (810) is not identified.
[0191] In one embodiment, the electronic device (101) may highlight and display a page that includes at least one other word (or sentence) (830) in the text data (423) having timestamp information that corresponds to at least one word (or sentence) (820) selected by a user input (810), based on the fact that at least one other word (or sentence) (830) in the text data (423) having timestamp information that corresponds to at least one word (or sentence) (820) selected by a user input (810) is not identified. In one embodiment, the electronic device (101) may not highlight at least one other word (or sentence) (830) in the text data (423) having timestamp information that corresponds to at least one word (or sentence) (820) selected by a user input (810) is not identified. If at least one other word (or sentence) within the text data (423) is not highlighted, the electronic device (101) may display an indicator indicating that the associated handwriting input (or multiple strokes (421)) is not identified.
[0192] For example, referring to FIG. 8A, the electronic device (101) may highlight and display at least one sentence (820) selected based on a user input (810) and at least one sentence (830) corresponding to the sentence (820). According to an embodiment, the electronic device (101) may highlight and display other sentences corresponding to the at least one sentence (820) selected based on the user input (810) within a plurality of areas (e.g., a summary (425)). For example, referring to FIG. 8A, the electronic device (101) may highlight and display at least one sentence (820) selected based on a user input (810), at least one sentence (830) corresponding to the sentence (820), and at least one sentence (835) corresponding to the sentence (820). According to an embodiment, the electronic device (101) may highlight and display at least one word selected based on a user input (810) and at least one word corresponding to the selected at least one word.
[0193] For example, highlighting may include changing the characteristics of the selected sentences (820, 830) (e.g., font, size, color, visual effects (e.g., typography (e.g., shadow, neon effect)). For example, highlighting may include changing the characteristics of the selected sentences (820, 830) to be different from the characteristics of sentences surrounding the selected sentences (820, 830).
[0194] In FIG. 8A, a user input (810) for selecting a sentence (820) within a summary (415) is illustrated, but this is merely an example. In an embodiment, the user input may be a user input for selecting a sentence (830) within text data (423). The electronic device (101) may highlight and display at least one sentence (830) selected based on the user input and at least one sentence (820) corresponding to the sentence (830). In an embodiment, the electronic device (101) may highlight and display at least one sentence (830) selected based on the user input and at least one sentence (830) within the summary (425) corresponding to the sentence (830).
[0195] In FIG. 8A, the electronic device (101) is illustrated as receiving a user input (810) while the screen (801) is divided into a plurality of regions (710, 730), but this is merely an example. According to an embodiment, the electronic device (101) may display the region (730) based on receiving a user input (e.g., a touch input (e.g., a drag input) and / or a hovering input) within the region (710) while displaying the screen (801) displaying the region (710). For example, the electronic device (101) may receive a user input for selecting a sentence (830) within the region (710) while displaying the screen (801) displaying the region (710). For example, while displaying a screen (801) displaying an area (710), the electronic device (101) may display an area (730) including a summary (415) including a sentence (820) corresponding to the sentence (830) based on receiving a user input for selecting a sentence (830) within the area (710). For example, while displaying a screen (801) displaying an area (710), the electronic device (101) may display an area (730) including a sentence (820) as a pop-up screen around the sentence (830) within the area (710).
[0196] FIG. 8B is a diagram illustrating another example of a screen of note content in which an electronic device highlights and displays another object corresponding to a selected object, according to one embodiment.
[0197] Fig. 8b, compared to Fig. 7g, may illustrate a situation in which an electronic device (101) receives a user input (811). Descriptions of Fig. 8b that overlap with those of Fig. 7g may not be repeated.
[0198] Referring to the screen (802) of FIG. 8B, the electronic device (101) may receive a user input (811). In one embodiment, the user input (811) may be an input for selecting at least one word (or sentence). In one embodiment, the user input (811) may be a touch input (e.g., a drag input) and / or a hovering input.
[0199] In one embodiment, the electronic device (101) can identify at least one other word (or sentence) corresponding to the selected at least one word (or sentence) based on a user input (811) selecting at least one word (or sentence).
[0200] For example, referring to FIG. 8B, the electronic device (101) can identify at least one other word (or sentence) in other content (e.g., text data (413)) based on keywords (e.g., “4th industrial revolution,” “traditional industry,” “business,” “paradigm,” “economic society in general,” and “digital transformation speed”) to which a sequence is assigned among words included in a sentence selected in the summary (415).
[0201] For example, the electronic device (101) can identify keywords (e.g., “4th industrial revolution,” “traditional industry,” “business,” “paradigm,” “economic society in general,” and “digital transformation speed”) to which sequences are assigned among the words included in the sentences selected in the summary (415) within the text data (413) related to the summary (415) within the text data (413). The keywords (e.g., “4th industrial revolution,” “traditional industry,” “business,” “paradigm,” “economic society in general,” and “digital transformation speed”) to which sequences are assigned are identified.
[0202] In one embodiment, the electronic device (101) may highlight and display at least one selected word (or sentence) based on a user input (811) selecting at least one word (or sentence). In one embodiment, the electronic device (101) may highlight and display at least one selected word (or sentence) while highlighting and displaying at least one other word (or sentence) corresponding to the selected at least one word (or sentence).
[0203] For example, referring to FIG. 8B, the electronic device (101) may highlight and display at least one sentence (821) selected based on a user input (811) and at least one sentence (831) corresponding to the sentence (821). For example, highlighting may include changing characteristics (e.g., font, size, color, visual effects (e.g., typography (e.g., shadow, neon effect))) of the selected sentences (821, 831). For example, highlighting may include changing characteristics of the selected sentences (821, 831) to be different from characteristics of sentences surrounding the selected sentences (821, 831).
[0204] In FIG. 8B, a user input (811) for selecting a sentence (821) within a summary (415) is illustrated, but this is merely an example. In some embodiments, the user input may be a user input for selecting a sentence (831) within text data (413). For example, the electronic device (101) may highlight and display at least one sentence (831) selected based on the user input and at least one sentence (821) corresponding to the sentence (831).
[0205] In FIG. 8B, the electronic device (101) is illustrated as receiving a user input (811) while the screen (801) is divided into a plurality of regions (710, 730), but this is merely an example. According to an embodiment, the electronic device (101) may display the region (730) based on receiving a user input (e.g., a touch input (e.g., a drag input) and / or a hovering input) within the region (710) while displaying the screen (801) displaying the region (710). For example, the electronic device (101) may receive a user input for selecting a sentence (831) within the region (710) while displaying the screen (801) displaying the region (710). For example, while displaying a screen (801) displaying an area (710), the electronic device (101) may display an area (730) including a summary (415) including a sentence (821) corresponding to the sentence (831) based on receiving a user input for selecting a sentence (831) within the area (710). For example, while displaying a screen (801) displaying an area (710), the electronic device (101) may display an area (730) including a sentence (821) as a pop-up screen around the sentence (831) within the area (710).
[0206] FIG. 8C is a diagram illustrating another example of a screen of note content in which an electronic device highlights and displays objects according to a playback time point, according to an embodiment of the present invention. FIG. 8D is a diagram illustrating another example of a screen of note content in which an electronic device highlights and displays objects according to a playback time point, according to an embodiment of the present invention.
[0207] FIGS. 8C and 8D can illustrate a situation in which an electronic device (101) receives user input, compared to FIG. 7B. Descriptions of FIGS. 8C and 8D that overlap with those of FIG. 7B may not be repeated.
[0208] Referring to screen (803) of FIG. 8c, the electronic device (101) can receive user input. In one embodiment, the user input may be an input for selecting a visual object (840) for playback.
[0209] In one embodiment, the electronic device (101) may sequentially highlight and display the plurality of strokes (421) in the document (420) based on a timestamp at which the plurality of strokes (421) are received, based on a user input selecting a visual object (840) for playback. In one embodiment, the electronic device (101) may sequentially play the plurality of utterances (411) in the audio data (410) based on a timestamp at which the plurality of utterances (411) are received, based on a user input selecting a visual object (840) for playback. In one embodiment, the electronic device (101) may sequentially highlight and display the words in the text data (413) converted from the plurality of utterances (411) in the audio data (410) and / or the summary (415) based on a user input selecting a visual object (840) for playback. In one embodiment, the electronic device (101) may sequentially highlight and display words in the text data (413) converted from a plurality of utterances (411) in the audio data (410) and / or the summary (415) based on a user input selecting a visual object (840) for playback, in the order in which the strokes were input. In one embodiment, the electronic device (101) may highlight and display a next input stroke while the previous stroke is highlighted and displayed during playback of the audio data (410). In one embodiment, the electronic device (101) may maintain the highlighted stroke while playback of the audio data (410) is maintained.
[0210] For example, referring to FIG. 8C, the electronic device (101) may, while sequentially playing back multiple utterances (411) in the audio data (410), highlight and display "in" (851) in accordance with the timestamp at which strokes for "ㅇ", "ㅣ", and "ㄴ") for "인" (851) among the "cognitive labor" in the text data (423) are input. For example, referring to FIG. 8C, the electronic device (101) may, at the time at which the utterance acquired at the time at which the strokes for "인" (851) among the "cognitive labor" in the text data (423) are input, highlight and display a word (e.g., "인" (861)) in the summary (415) corresponding to "인" (851) is played back. For example, the electronic device (101) may highlight and display a word (e.g., “recognition” (861)) corresponding to a voice input obtained during the time a stroke for “in” (851) in the text data (423) is input.
[0211] For example, referring to FIG. 8d, the electronic device (101) can highlight and display "ㅈ" (852) at the time when the utterance acquired at the time when the stroke for "ㅈ" (852) among "cognitive labor" in the text data (423) is input is played. For example, referring to FIG. 8d, the electronic device (101) can highlight and display a word (e.g., "labor" (862)) in the summary (415) corresponding to "ㅈ" (852) at the time when the utterance acquired at the time when the stroke for "ㅈ" (852) among "cognitive labor" in the text data (423) is played. For example, the electronic device (101) can highlight and display a word (e.g., "labor" (862)) corresponding to the utterance acquired at the time when the stroke for "ㅈ" (852) among the text data (423) is input.
[0212] In one embodiment, the electronic device (101) may keep the highlighted strokes highlighted during playback according to the selection of the visual object (840). For example, the electronic device (101) may highlight and display "in" (851) at the time when the utterance obtained at the time when the stroke for "in" (851) among "cognitive labor" in the text data (423) is input is played, and then highlight and display "j" at the time when the utterance obtained at the time when the stroke for "j" (852) among the "cognitive labor" in the text data (423) is input is played, while also highlighting and displaying "j".
[0213] According to an embodiment, the electronic device (101) may sequentially highlight and display a plurality of strokes (421) within a document (420) based on a timestamp at which the strokes are received, based on a user input (812) selecting a visual object (840) for playback, and may sequentially highlight and display words within a summary (425) of text data (423). For example, the electronic device (101) may sequentially highlight and display a word (e.g., “cognition”) within the summary (425) corresponding to “in” (851) at a time at which an utterance acquired at a time at which strokes for “in” (851) among “cognitive labor” within the text data (423) are input is played. For example, the electronic device (101) can highlight and display a word (e.g., “labor”) in the summary (425) corresponding to “ㅈ” (852) at a time when the utterance acquired at the time when the stroke for “ㅈ” (852) among “cognitive labor” in the text data (423) is input is played.
[0214] In FIGS. 8C and 8D , only strokes for letters are highlighted, but this is merely an example. According to an embodiment, the electronic device (101) may highlight and display various marks (e.g., highlights, underlines, circles, and asterisks) that may be input according to a user input in accordance with the playback time. In addition, the electronic device (101) may highlight marks based on words completed by strokes, in addition to highlighting marks on a stroke-by-stroke basis. For example, when two or more strokes are required to complete a letter, the electronic device (101) may highlight only a single stroke, or highlight a letter related to the single stroke. For example, the electronic device (101) may highlight "ㅣ" or highlight and display "지" related to "ㅣ" in accordance with the time when a stroke for a part of "지" (e.g., "ㅣ") among "cognitive labor" in the text data (423) is input. Here, a stroke may be an input identified while the user maintains a touch input on the display (260) (e.g., after pressing the touch input but before releasing it).
[0215] According to an embodiment, the electronic device (101) may change the properties of a character based on the playback timing of an input for changing the properties of the character. For example, when writing a note, if a user inputs "cognitive labor" and then changes the properties of the "cognitive labor" (e.g., color, thickness, font, or pen type), the electronic device (101) may change the properties of the "cognitive labor" that are highlighted and displayed according to the playback timing.
[0216] According to an embodiment, the electronic device (101) may scroll (or move) the text data (423) displayed in the area (710) based on an input for selecting a visual object (840) for playback, and when the text data (423) displayed in the area (710) is scrolled (or moved), the summaries (415, 425) displayed in the area (730) may scroll (or move) in accordance with the text data (423) displayed in the area (710) according to the scrolling (or moving). In one embodiment, the electronic device (101) may scroll (or move) the summaries (415, 425) displayed in the area (730) according to the scrolling (or moving) when the page displayed in the area (710) is scrolled (or moved). In one embodiment, when the text data (423) displayed in the area (710) is scrolled (or moved) as the note content is played, the electronic device (101) may change (or move) the object that is highlighted in the played content (e.g., the text data (423)).
[0217] FIG. 8E is a diagram illustrating another example of a screen of note content in which an electronic device highlights and displays objects based on user input, according to one embodiment.
[0218] FIG. 8e illustrates a situation in which, in a viewing situation in which a document (420) is viewed again after a document (420) in which multiple strokes (421) have been input is saved, at least one other corresponding word (or sentence) is highlighted and displayed based on an input for selecting words included in content (331) within the document (420).
[0219] Referring to screen (804) of FIG. 8E, the electronic device (101) may receive a user input for selecting text (871) (“4th Industrial Revolution”). In one embodiment, the user input for selecting text (871) (“4th Industrial Revolution”) may be a touch input (e.g., a drag input) and / or a hovering input directed toward the text (e.g., “4th Industrial Revolution”).
[0220] In one embodiment, the electronic device (101) may identify at least one other word (or sentence) corresponding to at least one selected word (or sentence) based on a user input selecting text (871) (“4th Industrial Revolution”). For example, referring to FIG. 8E , the electronic device (101) may identify at least one other word (or sentence) in other content (e.g., text data (413) converted from a plurality of utterances (411)) based on keywords (e.g., “4th Industrial Revolution,” “traditional industry,” “business,” “paradigm,” “economic society in general,” and “digital transformation”) to which sequences are assigned among words included in the content (331). For example, the electronic device (101) can identify a keyword (881) (e.g., “4th Industrial Revolution”) corresponding to the selected text (871) (“4th Industrial Revolution”) within the content (331) among the words within the text data (413).
[0221] In one embodiment, the electronic device (101) may highlight and display at least one selected keyword (881) (e.g., "4th Industrial Revolution") based on a user input selecting text (871) ("4th Industrial Revolution"). For example, highlighting may include changing the characteristics (e.g., font, size, color, visual effects (e.g., typography (e.g., shadow, neon effect))) of the selected keyword (881) (e.g., "4th Industrial Revolution"). For example, highlighting may include changing the characteristics of the selected keyword (881) (e.g., "4th Industrial Revolution") to be different from the characteristics of words surrounding the selected keyword (881) (e.g., "4th Industrial Revolution"). In addition, according to the user input, as objects are highlighted, the playback time of the content being played may also move to a related playback time (e.g., the playback time of the selected keyword (881).
[0222] Similarly, referring to screen (804) of FIG. 8E, the electronic device (101) may receive a user input selecting text (872) (“digital conversion”). In one embodiment, the electronic device (101) may identify a keyword (882) (e.g., “digital conversion”) corresponding to the selected text (872) (“digital conversion”) within the content (331) among words within the text data (413), based on the user input selecting the text (872) (“digital conversion”).
[0223] In one embodiment, the electronic device (101) may highlight and display at least one selected keyword (882) (e.g., “digital conversion”) based on a user input selecting text (872) (“digital conversion”). Additionally, as objects are highlighted based on the user input, the playback point of the content being played may also move to an associated playback point (e.g., the playback point of the selected keyword (882)).
[0224] FIG. 9A is a diagram illustrating another example of a screen of note content that displays text data converted from strokes and audio data, according to an embodiment. FIG. 9B is a diagram illustrating another example of a screen of note content that displays a summary of text data converted from strokes and audio data, according to an embodiment. FIG. 9C is a diagram illustrating an example of a screen of note content that displays a translation of a summary of text data converted from strokes and audio data, according to an embodiment. FIG. 9D is a diagram illustrating an example of a screen of note content that displays another object corresponding to a selected object by highlighting it, according to an embodiment.
[0225] FIG. 9a may be described with reference to the electronic device (101) (or components of the electronic device (101)) of FIG. 2.
[0226] Referring to FIG. 9A, the electronic device (101) can display a screen (901) of a content service APP (210) through a display (260). For example, in response to an execution request (or a viewing request) for a note (e.g., a file including a document (420) and audio data (410)), the electronic device (101) can display a screen (901) of the content service APP (210) that displays the note.
[0227] In one embodiment, the screen (901) may be divided into a plurality of regions (910, 950). The plurality of regions (910, 950) may be arranged vertically within the screen (901). For example, the region (910) may be arranged above the region (950). The plurality of regions (910, 950) may be arranged so as not to overlap each other within the screen (901).
[0228] For example, the electronic device (101) may display a user interface (UI) (913) in response to an input to an icon (911) for displaying the UI (913) while displaying an area (910) within the screen (901). For example, the electronic device (101) may, while displaying an area (910) within the screen (901), release the display of the displayed UI (913) in response to an additional input to the icon (911). In one embodiment, the UI (913) may display one or more icons (each of which represents audio data).
[0229] For example, the electronic device (101) may display the area (950) together with the area (910) in response to an input to the UI (913) for controlling the display of the area (950) while displaying the area (910) within the screen (901).
[0230] In one embodiment, text data (423) may be displayed in the region (910). For example, in one embodiment, text data (423) corresponding to a selected icon (e.g., voice 001) may be displayed in the region (910). In one embodiment, text data (413) may be displayed in the region (950). In one embodiment, text data (413) may be displayed in the region (950) based on the selection of the tab for "voice content" among the tabs for "voice content" and "summary." The electronic device (101) may highlight (e.g., underline, bold, change color) the tab for "voice content" based on the selection of the tab for "voice content." In one embodiment, one or more visual objects (731, 733, 735) may be displayed in the region (950). For example, a visual object (731) may indicate that the text data (413) is original content that has not been summarized. For example, a visual object (733) may be a visual object for requesting a translation of the text data (413). For example, a visual object (735) may be a visual object representing a speaker of the text data (413). In one embodiment, the speaker of the text data (413) may be distinguished based on inputting a plurality of utterances (411) into the generative AI module (220). For example, the speaker of the text data (413) may be distinguished through the generative AI module (220) based on the tone and / or pronunciation characteristics of the plurality of utterances (411). According to an embodiment, the speaker of the text data (413) can be distinguished through a generative AI module (220) within the electronic device (101) and another program based on the tone and / or pronunciation characteristics of the plurality of utterances (411).
[0231] In one embodiment, text data (423) displayed in area (910) and text data (413) displayed in area (950) may correspond to each other. For example, the electronic device (101) may display text data (423) and text data (413) in areas (910, 950) so that they are synchronized based on relationship information (430). For example, the electronic device (101) may display text data (413) temporally and / or semantically mapped to text data (423) displayed in area (910) in area (950) based on relationship information (430).
[0232] For example, the electronic device (101) may display words of text data (413) that are temporally and / or semantically mapped to words of text data (423) displayed in the area (910) in the area (950) based on relationship information (430).
[0233] In one embodiment, when text data (423) displayed in an area (910) is scrolled (or moved), the electronic device (101) can scroll (or move) text data (413) displayed in an area (950) in accordance with the text data (423) displayed in the area (910) according to the scrolling (or moving). In one embodiment, when text data (413) displayed in an area (950) is scrolled (or moved), the electronic device (101) can scroll (or move) text data (423) displayed in an area (910) in accordance with the text data (413) displayed in the area (950) according to the scrolling (or moving).
[0234] Referring to FIG. 9B, the electronic device (101) may display a screen (902) of a content service APP (210) through a display (260). In one embodiment, the screen (902) may be a screen in which a summary (415) is displayed in an area (950) instead of text data (413) compared to the screen (901). In one embodiment, the area (950) may display a summary (415) instead of text data (413) based on whether the tab for "summary" is selected among the tabs for "voice content" and the tab for "summary." The electronic device (101) may highlight (e.g., underline, bold, or change color) the tab for "summary" based on whether the tab for "summary" is selected. In one embodiment, at least one visual object (733) may be displayed in the area (950).
[0235] For example, the electronic device (101) may display words of a summary (415) that are temporally and / or semantically mapped to words of text data (423) displayed in an area (910) based on relationship information (430), in an area (950).
[0236] In one embodiment, when text data (423) displayed in an area (910) is scrolled (or moved), the electronic device (101) can scroll (or move) the summary (415) displayed in an area (950) to match the text data (423) displayed in the area (910) according to the scrolling (or moving). In one embodiment, when the summary (415) displayed in an area (950) is scrolled (or moved), the electronic device (101) can scroll (or move) the text data (423) displayed in an area (910) to match the summary (415) displayed in an area (950) according to the scrolling (or moving).
[0237] Referring to FIG. 9C, the electronic device (101) may display a screen (903) of a content service APP (210) through a display (260). In one embodiment, the screen (903) may be a screen in which a translation (970) of the summary (415) is displayed together with the summary (415) in an area (950) compared to the screen (902). In one embodiment, the electronic device (101) may display the screen (903) based on a user input of selecting a visual object (953) on the screen (902). For example, the electronic device (101) may generate a translation (970) for the summary (415) by inputting the summary (415) into the generative AI module (220) based on a user input of selecting a visual object (953) on the screen (902). For example, the electronic device (101) can display the generated translation (970) along with a summary (415) on the screen (903).
[0238] In one embodiment, the area (950) may display one or more visual objects (963, 971). For example, the visual object (963) may be a visual object indicating a translation of the summary (415). For example, the visual object (971) may be a visual object indicating the language of the summary (415) and the language of the translation (970).
[0239] In one embodiment, the electronic device (101) may generate relationship information between the translation (970) and the summary (415). For example, the electronic device (101) may generate relationship information that maps words in the translation (970) to words in the summary (415). For example, the electronic device (101) may map "4th Industrial Revolution" in the summary (415) to "Fourth Industrial Revolution" in the translation (970). For example, the electronic device (101) may map "4th Industrial Revolution" in the summary (415) to "Fourth Industrial Revolution" in the translation (970) with the highest attention. For example, the electronic device (101) can generate relationship information between the translation (970) and other data (e.g., text data (413), text data (423), summary (425)) based on relationship information that maps words in the translation (970) to words in the summary (415).
[0240] For example, the electronic device (101) may display words of a summary (415) that are temporally and / or semantically mapped to words of text data (423) displayed in an area (910) based on relationship information (430), in an area (950).
[0241] In one embodiment, when text data (423) displayed in an area (910) is scrolled (or moved), the electronic device (101) can scroll (or move) the summary (415) displayed in an area (950) to match the text data (423) displayed in the area (910) according to the scrolling (or moving). In one embodiment, when the summary (415) displayed in an area (950) is scrolled (or moved), the electronic device (101) can scroll (or move) the text data (423) displayed in an area (910) to match the summary (415) displayed in an area (950) according to the scrolling (or moving).
[0242] Referring to the screen (904) of FIG. 9D, the electronic device (101) may receive a user input (980). In one embodiment, the user input (980) may be an input for selecting at least one word (or sentence) (981). In one embodiment, the user input (980) may be a touch input (e.g., a drag input) and / or a hovering input.
[0243] In one embodiment, the electronic device (101) can identify at least one other word (or sentence) (990) corresponding to the selected at least one word (or sentence) (981) based on a user input (980) selecting at least one word (or sentence) (981).
[0244] In one embodiment, the electronic device (101) can identify at least one other word (or sentence) (990) in another content based on a sequence assigned to at least one selected word (or sentence) (981) in one content. In one embodiment, the electronic device (101) can identify at least one other word (or sentence) (990) in another content based on a sequence assigned to at least one selected word (or sentence) (981) in one content, based on relationship information between the translation (970) and the summary (415), and relationship information between the summary (415) and other data (e.g., text data (413), text data (423), summary (425)). For example, referring to FIG. 9d, the electronic device (101) can identify keywords (e.g., “20th century”, “electrical and electronic equipment”, “development”, “computer”, “Internet”, and “emergence”) among words in the text data (423) that have corresponding timestamp information based on keywords (e.g., “Fourth Industrial Revolution”, “business”, “paradigm”, “traditional industries”, “economy and society”, and “speed of digital transformation”) to which sequences are assigned among words included in a sentence (981) selected in the translation (970).
[0245] In one embodiment, the electronic device (101) may highlight and display at least one selected word (or sentence) (981) based on a user input (980) selecting at least one word (or sentence) (981). In one embodiment, the electronic device (101) may highlight and display at least one other word (or sentence) (990) corresponding to the selected at least one word (or sentence) while highlighting and displaying the at least one selected word (or sentence) (981) based on a user input (980) selecting at least one word (or sentence) (981).
[0246] For example, referring to FIG. 9D, the electronic device (101) may highlight and display at least one sentence (981) selected based on a user input (980) and at least one sentence (990) corresponding to the sentence (981). For example, highlighting may include changing characteristics (e.g., font, size, color, visual effects (e.g., typography (e.g., shadow, neon effect))) of the selected sentences (981, 990). For example, highlighting may include changing characteristics of the selected sentences (981, 990) to be different from characteristics of sentences surrounding the selected sentences (981, 990).
[0247] According to an embodiment, the electronic device (101) may highlight and display at least one word (or sentence) (985) corresponding to at least one word (or sentence) (981) selected from the translation (970) within the summary (415) based on a user input (980) selecting at least one word (or sentence) (981).
[0248] In FIG. 9D, a user input (980) for selecting a sentence (981) within a translation (970) is illustrated, but this is merely an example. In some embodiments, the user input may be a user input for selecting a sentence (990) within text data (423). The electronic device (101) may highlight and display at least one sentence (990) selected based on the user input and at least one sentence (981) corresponding to the sentence (990). In some embodiments, the electronic device (101) may highlight and display at least one sentence (990) selected based on the user input and at least one sentence (990) within the translation (970) corresponding to the sentence (990).
[0249] FIG. 10A is a diagram illustrating a relationship between pages and sets of utterances within a document, according to one embodiment. FIG. 10B is a diagram illustrating an example of a screen on which an electronic device displays a page and a set of utterances corresponding to the page, according to one embodiment.
[0250] FIG. 10A may illustrate mapping information between pages in a document (420) and utterance sets in audio data (410). In one embodiment, the mapping information may map utterance set 1 according to received voice input when a user writes note content on page 1. In one embodiment, the mapping information may map utterance set 2 according to received voice input when a user writes note content on page 2. Similarly, in one embodiment, the mapping information may map utterance set N according to received voice input when a user writes note content on page N. Here, N may be a natural number.
[0251] In one embodiment, when displaying the screen of the content service APP (210), the electronic device (101) may refer to mapping information between pages in the document (420) and utterance sets in the audio data (410).
[0252] For example, referring to FIG. 10b, when displaying page 1 in an area (1011) through a screen (1001) of a content service APP (210), the electronic device (101) may display text data (413) based on utterance set 1 corresponding to page 1 and / or a summary (415) in the area (1013). For example, when displaying page 2 in an area (1011) through a screen of a content service APP (210), the electronic device (101) may display text data (413) based on utterance set 2 corresponding to page 2 and / or a summary (415) in the area (1013). For example, when displaying page N in an area (1011) through the screen of a content service APP (210), the electronic device (101) may display text data (413) and / or a summary (415) based on a set of utterances N corresponding to page N in the area (1013). Here, N may be a natural number.
[0253] FIGS. 11A to 11C are diagrams illustrating examples of visual objects indicating synchronization within a document displayed by an electronic device, according to one embodiment.
[0254] In one embodiment, the electronic device (101) may link audio data (410) (or, multiple utterances (411) (or, text data (413) converted from audio data (410) and / or summary (415)) according to voice input to a specific object within a document (420). For example, the electronic device (101) may link a portion of audio data (410) (or, multiple utterances (411) (or, text data (413) converted from audio data (410) and / or summary (415)) according to voice input to a specific object within a document (420). In one embodiment, the specific object may be handwriting, a page, text data (423), summary (425), and / or an image (or video) within the document (420). Linking audio data (410) to a specific object within the document (420) may be a content service. When a specific object is displayed on the screen of the APP (210), it may mean that the associated audio data (410) may be played, or the text data (413) converted from the associated audio data (410) and / or the summary (415) may be displayed. In one embodiment, the playback of the associated audio data (410), or the display of the text data (413) converted from the associated audio data (410), and / or the summary (415) may be performed based on a user's request. However, the present invention is not limited thereto. For example, the playback of the associated audio data (410), or the display of the text data (413) converted from the associated audio data (410), and / or the summary (415) may be performed when a specific object is displayed on the screen of the content service APP (210).
[0255] For example, referring to FIG. 11A, the electronic device (101) may link audio data (410) (or a portion of audio data (410)) to selected text (340) based on a user input of selecting an object (1115) for linking audio data (410) to a specific object in a document (420) among objects included in a UI (1110) on a screen (1101). In one embodiment, the electronic device (101) may display a visual object (1120) indicating that audio data (410) is linked to the selected text (340).
[0256] For example, referring to FIG. 11B, the electronic device (101) may link audio data (410) (or a portion of audio data (410)) to a page based on a user input of selecting an object (1115) for linking audio data (410) to a specific object in a document (420) among objects included in a UI (1110) on a screen (1101). Here, a list of pages may be displayed on the screen (1102) of the content service APP (210) based on selecting a visual object (1130). In one embodiment, the electronic device (101) may display a visual object (1140) indicating that audio data (410) is linked to the selected page.
[0257] For example, referring to FIG. 11C, the electronic device (101) may link audio data (410) (or a portion of audio data (410)) to an image based on a user input of selecting an object (1115) for linking audio data (410) to a specific object in a document (420) among objects included in a UI (1110) on a screen (1101). In one embodiment, the electronic device (101) may display a visual object (1150) indicating that audio data (410) is linked to the selected image.
[0258] FIG. 12 is a flowchart illustrating the operation of an electronic device according to one embodiment.
[0259] FIG. 12 can be described with reference to FIGS. 2 to 11C. The operations of FIG. 12 can be executed by the electronic device (101) (or the processor (120) of the electronic device (101). For example, the operations of FIG. 12 can be performed by the electronic device (101) as the processor (120) executes instructions of a program (140) stored in the memory (130).
[0260] Referring to FIG. 12, in operation 1210, the electronic device (101) may receive an input for selecting an object from first content. In one embodiment, the first content may be at least one of the contents of a document (e.g., document (420) of FIG. 4) included in the note content, a user's text input entered into the document (e.g., handwriting input (e.g., multiple strokes (421) of FIG. 4), text data converted from the text input (e.g., text data (423) of FIG. 4), or a summary of the text data converted from the text input (e.g., summary (425) of FIG. 4), a voice input (e.g., audio data (410) or multiple utterances (411) of FIG. 4), text data converted from the voice input (e.g., text data (413) of FIG. 4), or a summary of the text data converted from the voice input (e.g., summary (415) of FIG. 4).
[0261] For example, the electronic device (101) may receive a user input for selecting at least one word (or keyword) within the text data (423) while displaying a summary (415) that summarizes text data (413) representing audio data (410) and text data (423) according to a split view.
[0262] In operation 1220, the electronic device (101) may identify another object corresponding to the object selected by input from the second content. In one embodiment, the second content may be content other than the first content among the contents included in the note content.
[0263] For example, the electronic device (101) can identify at least one word in the summary (415) corresponding to at least one word (or keyword) in the text data (423) based on the relationship information (430). For example, the electronic device (101) can identify, based on a user input selecting at least one word (or keyword) in the text data (423), an utterance (or at least one word (or keyword) converted from the utterance) in the audio data (410) corresponding to the at least one word (or keyword) selected in the text data (423). For example, the electronic device (101) can identify at least one word in the summary (415) summarized from the utterance (or at least one word (or keyword) converted from the utterance) in the audio data (410).
[0264] In operation 1230, the electronic device (101) may highlight and display objects and other objects. For example, highlighting may include changing characteristics (e.g., font, size, color, visual effects (e.g., typography (e.g., shadow, neon effect))) of the selected objects and other objects. For example, highlighting may include changing characteristics of the selected objects and other objects so that they are different from characteristics of objects surrounding the selected objects and other objects.
[0265] For example, the electronic device (101) may visually highlight at least one word (or keyword) within text data (423). For example, the electronic device (101) may visually highlight at least one word within a summary (415) summarized from utterances within audio data (410) while highlighting at least one word (or keyword) within text data (423) identified by a user input.
[0266] FIG. 13A is a diagram illustrating an example of a screen of note content displaying text data converted from strokes and audio data, according to one embodiment.
[0267] FIG. 13b is a diagram illustrating an example of a progress bar according to one embodiment.
[0268] FIG. 13a and FIG. 13b may be described with reference to the electronic device (101) (or components of the electronic device (101)) of FIG. 2.
[0269] Referring to FIG. 13A, the electronic device (101) can display a screen (1301) of a content service APP (210) through a display (260). For example, in response to an execution request (or a viewing request) for a note (e.g., a file including a document (420) and audio data (410)), the electronic device (101) can display a screen (1301) of the content service APP (210) that displays the note.
[0270] In one embodiment, the screen (1301) may include content (331). In one embodiment, the content (331) displayed on the screen (1301) may be video content. In one embodiment, the content (331) may be included as a video and / or link in a document (420) within the content service APP (210).
[0271] In one embodiment, content (331) may be displayed in one of the segmented regions (710) within screen (1301). For example, content (331) may be displayed based on selecting an icon within a visual object (e.g., 325 in FIG. 3B) that is displayed by an input selecting an affordance (e.g., 315 in FIG. 3A) for adding content within screen (1301). However, the present invention is not limited thereto. For example, content (331) may be displayed as a pop-up screen on screen (1301). For example, content (331) may be included in a screen of another application displayed outside of screen (1301). For example, depending on multi-window, content (331) may be displayed on a screen other than screen (1301).
[0272] In one embodiment, the screen (701) may be a screen divided into a plurality of regions (710, 730). The plurality of regions (710, 730) may be arranged left and right within the screen (701). For example, the region (710) may be arranged to the left of the region (730). The plurality of regions (710, 730) may be arranged so as not to overlap each other within the screen (701). In one embodiment, the size and / or position of the plurality of regions (710, 730) may be adjusted according to a user input. For example, the electronic device (101) may enlarge the region (710) to the right (or enlarge in the direction of the region (730)) or enlarge the region (730) to the left (or enlarge in the direction of the region (710)) based on the user input. For example, the electronic device (101) can arrange multiple areas (710, 730) vertically based on user input.
[0273] In one embodiment, the area (710) may display text data (423) composed of content (331) and a plurality of strokes (421). In one embodiment, the text data (423) may be generated based on the plurality of strokes (421) acquired while playing the content (331). For example, the electronic device (101) may receive the plurality of strokes (421) through handwriting input (e.g., user input for letters and / or symbols) while the user plays the content (331).
[0274] In one embodiment, text data (413) may be displayed in the area (730). In one embodiment, the text data (413) may be a subtitle included in the content (331). In one embodiment, the text data (413) may be text data extracted (or recognized) based on audio data (410) included in the content (331).
[0275] In one embodiment, text data (423) composed of a plurality of strokes (421) displayed in an area (710) and / or text data (413) displayed in an area (730) may correspond to each other. For example, based on the time point at which each of a plurality of sections of content (331) is played, text data (423) composed of a plurality of strokes (421) and / or text data (413) displayed in an area (730) may be temporally and / or semantically mapped. For example, the electronic device (101) may synchronize text data (423) and text data (413) to the time point at which each of a plurality of sections of content (331) is played, based on the relationship information (430). In one embodiment, the electronic device (101) may display a visual indication within the screen (1301) that the text data (423) and the text data (413) are synchronized with the time at which each of the plurality of sections of the content (331) is played. For example, the electronic device (101) may display a visual indication within a progress bar (1330) of the content (331) that the text data (423) and the text data (413) are synchronized with the time at which each of the plurality of sections of the content (331) is played. For example, referring to FIG. 13B, in addition to an indicator (1331) indicating a playback point in a progress bar (1330), the electronic device (101) may display text data (423) and visual indications (1341 to 1350) indicating that the text data (413) are synchronized with the playback point in time of each of a plurality of sections of the content (331), at synchronized points in time. For example, the electronic device (101) may change the display characteristics (e.g., color, size, shape) of the progress bar (1330) so that the display of the progress bar (1330) is distinguished in a playback section of the content (331) in which a user has input handwriting.
[0276] In one embodiment, the electronic device (101) may, based on a user input selecting a visual object for playback of the content (331), sequentially highlight and display a plurality of strokes (421) based on timestamps at which the plurality of strokes (421) were received during a previous playback of the content (331) while simultaneously playing the content (331). In one embodiment, the electronic device (101) may, based on a user input selecting a visual object for playback, sequentially highlight and display words within text data (413) and / or a summary (415) while simultaneously playing the content (331). For example, referring to FIG. 13A, while the electronic device (101) sequentially plays back multiple utterances (411) in the content (331), the electronic device (101) may highlight and display the "sex" (1311) among the "sexual behavior" in the text data (423) in accordance with the timestamp at which some strokes (e.g., the stroke for "ㅅ", the stroke for "ㅓ", and the stroke for "ㅇ") are input.
[0277] In one embodiment, the electronic device (101) may change the screen of the content (331) to a screen corresponding to the timestamp of the selected word based on a user input selecting a specific word within the text data (423). In one embodiment, the electronic device (101) may change the screen of the content (331) to a screen corresponding to the timestamp of the selected word based on a user input selecting a specific word within the text data (413).
[0278] For example, referring to FIG. 13A, in response to a user input selecting “Sex” (1311) among “sexual language” in the text data (423), the electronic device (101) may change the screen of the content (331) to a screen output while strokes for “Sex” (1311) are being input. For example, referring to FIG. 13A, in response to a user input selecting “Sex” (1311) among “sexual language” in the text data (423), the electronic device (101) may display the text data (413) so that words corresponding to the timestamp at which strokes for “Sex” (1311) are input are displayed in the area (730) on the screen of the content (331).
[0279] For example, referring to FIG. 13A, based on a user input selecting “verbal conduct” (1321) within text data (413), the electronic device (101) may highlight (e.g., change the color from gray to black) some strokes (e.g., the stroke for “ㅅ”, the stroke for “ㅓ”, and the stroke for “ㅇ”) among the strokes for “sex” (1311) among the “sexual conduct” within the text data (423), in accordance with the timestamp at which the “verbal conduct” (1321) was played within the content (331). Thereafter, the electronic device (101) may display the next stroke (e.g., the stroke for “ㅇ” among the strokes for “성” (1311)) by highlighting it (e.g., changing the color from gray to black) as the content (331) is played after a user input of selecting “언동” (1321).
[0280] In Fig. 13a, only characters are illustrated, but this is merely an example. According to an embodiment, the electronic device (101) may highlight and display various marks (e.g., highlights, underlines, circles, and asterisks) that can be input based on user input, in accordance with the playback time of the content (331). In addition, the electronic device (101) may highlight marks based on words completed by strokes, in addition to highlighting marks on a stroke-by-stroke basis. For example, when two or more strokes are required to complete a single letter, the electronic device (101) may highlight only a single stroke, or highlight a single letter associated with the single stroke.
[0281] As described above, the electronic device (101) may include a display (260), at least one processor (120) including a processing circuit; and a memory (130) storing instructions and including one or more storage media. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to display, through the display (260), a first text (415) summarizing text data (413) representing audio data (410) and a second text (425) representing the contents of a document (420), according to a split view. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify an input for a word within the second text (425). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to display the at least one word visually highlighted in the first text (415) based on identifying an utterance (411) corresponding to the word identified according to the input among the utterances (413) in the text data (413) representing the audio data (410) and identifying at least one word in the first text (415) summarizing the utterance, while displaying the word in the second text (425) visually highlighted according to the input.
[0282] The second text (425) may be a text summarizing the contents of the document (420). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate mapping information (430) that links words in the document (420) with words in the second text (425). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify, based on the mapping information (430), a word corresponding to the word identified according to the input among the words in the document (420). The above instructions, when executed individually or collectively by the at least one processor (120), may cause the electronic device (101) to identify an utterance corresponding to the word identified based on the mapping information (430) among the utterances (411) in the text data (413).
[0283] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to receive user input for the words in the document (420) while recording the audio data (410). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to obtain timestamp information mapping the utterances (411) in the text data (413) to the words in the document (420) based on a recording time of the audio data (410) and a reception time of the user input for the words. The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify, based on the timestamp information, the utterance corresponding to the word identified based on the mapping information (430) among the utterances (411) in the text data (413).
[0284] The user input for the words in the document (420) may include stroke inputs for entering the words into the document (420).
[0285] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate mapping information (430) that associates utterances (411) in the text data (413) with words in the first text (415). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify, based on the mapping information (430), the at least one word corresponding to the utterance corresponding to the word identified according to the input among the words in the first text (415).
[0286] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate the first text (415) summarizing the text data (413) based on inputting the text data (413) representing the audio data (410) into a generative artificial intelligence (AI) module.
[0287] The document (420) may include a plurality of pages, and the second text (425) may be distributed across the plurality of pages. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to display a portion of the first text (415) corresponding to an utterance obtained while receiving a user input for a portion of the second text (425) while displaying a portion of the second text (425) included in a selected page among the plurality of pages according to the split view through the display (260).
[0288] The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to display a visual object (1140) indicating that the audio data (410) is linked at a location where a portion of the second text (425) is displayed.
[0289] As described above, the electronic device (101) may include a display (260), at least one processor (120) including a processing circuit; and a memory (130) storing instructions and including one or more storage media. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to display, through the display (260), a first text (415) representing audio data (410) and a second text (425) summarizing a document (420), according to a split view. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify an input for a first word within the first text (415). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to display the at least one third word visually highlighted in the second text (425) based on identifying at least one second word corresponding to the first word identified according to the input among words in the document (420) and identifying at least one third word in the second text (425) corresponding to the at least one second word in the document (420), while displaying the first word in the first text (415) visually highlighted according to the input.
[0290] The first text (415) may be a text that summarizes the text data (413) converted from the audio data (410). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate mapping information (430) that links words in the text data (413) with words in the first text (415). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify, based on the mapping information (430), a fourth word corresponding to the first word identified according to the input among the words in the text data (413). The above instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify the at least one second word corresponding to the fourth word identified based on the mapping information (430) among the words in the second text (425).
[0291] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to receive user input for the words in the document (420) while recording the audio data (410). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to obtain timestamp information mapping the words in the text data (413) to the words in the document (420) based on a recording time of the audio data (410) and a reception time of the user input for the words. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify, based on the timestamp information, at least one second word corresponding to the fourth word identified based on the mapping information (430) among the words in the text data (413).
[0292] The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to generate mapping information (430) that associates words in the text data (413) with words in the first text (415). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify, based on the mapping information (430), the at least one word corresponding to the utterance corresponding to the first word identified according to the input among the words in the first text (415).
[0293] As described above, the method may be performed by an electronic device (101) including a display (260). The method may include an operation of displaying, through the display (260), a first text (415) summarizing text data (413) representing audio data (410) and a second text (425) representing the contents of a document (420), according to a split view. The method may include an operation of identifying an input for a word within the second text (425). The method may include an action of displaying the at least one word visually highlighted in the first text (415) while displaying the word visually highlighted in the second text (425) according to the input, based on identifying an utterance (411) corresponding to the word identified according to the input among utterances (413) in the text data (413) representing the audio data (410) and identifying at least one word in the first text (415) summarizing the utterance.
[0294] The second text (425) may be a text summarizing the contents of the document (420). The method may include an operation of generating mapping information (430) that links words in the document (420) and words in the second text (425). The method may include an operation of identifying, based on the mapping information (430), a word corresponding to the word identified according to the input among the words in the document (420). The method may include an operation of identifying an utterance corresponding to the word identified based on the mapping information (430) among the utterances (411) in the text data (413).
[0295] The method may include an operation of receiving user input for the words in the document (420) while recording the audio data (410). The method may include an operation of obtaining timestamp information that maps the utterances (411) in the text data (413) to the words in the document (420) based on a recording time of the audio data (410) and a reception time of the user input for the words. The method may include an operation of identifying, based on the timestamp information, the utterance corresponding to the word identified based on the mapping information (430) among the utterances (411) in the text data (413).
[0296] The method may include an operation of generating mapping information (430) that links utterances (411) in the text data (413) and words in the first text (415). The method may include an operation of identifying, based on the mapping information (430), at least one word corresponding to the utterance corresponding to the word identified according to the input among the words in the first text (415).
[0297] The document (420) may include a plurality of pages, and the second text (425) may be divided and included in the plurality of pages. The method may include an operation of displaying a first partial text corresponding to an utterance obtained while receiving a user input for a portion of the second text (425) among the first text (415) while displaying a portion of the second text (425) included in a selected page among the plurality of pages, according to the split view, through the display (260).
[0298] The method may include an action of displaying a visual object indicating that the audio data (410) is linked at a location where some text of the second text (425) is displayed.
[0299] As described above, the method can be performed in an electronic device (101) including a display (260). The method can include an operation of displaying, through the display (260), a first text (415) representing audio data (410) and a second text (425) summarizing a document (420), according to a split view. The method can include an operation of identifying an input for a first word in the first text (415). The method can include an operation of identifying at least one second word corresponding to the first word identified according to the input among words in the document (420) and identifying at least one third word in the second text (425) corresponding to the at least one second word in the document (420), while displaying the first word in the first text (415) visually highlighted according to the input.
[0300] The first text (415) may be a text that summarizes text data (413) converted from the audio data (410). The method may include an operation of generating mapping information (430) that links words in the text data (413) with words in the first text (415). The method may include an operation of identifying, based on the mapping information (430), a fourth word corresponding to the first word identified according to the input among the words in the text data (413). The method may include an operation of identifying, based on the mapping information (430), at least one second word corresponding to the fourth word identified based on the mapping information (430) among the words in the second text (425).
[0301] A non-transitory computer readable storage medium as described above may store a program including instructions. The instructions, when individually or collectively executed by at least one processor (120) of an electronic device (101) including a display (260), may cause the electronic device (101) to display, through the display (260), a first text (415) summarizing text data (413) representing audio data (410) and a second text (425) representing the content of a document (420), according to a split view. The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to identify an input for a word within the second text (425). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to display the at least one word visually highlighted in the first text (415) based on identifying an utterance (411) corresponding to the word identified according to the input among the utterances (413) in the text data (413) representing the audio data (410) and identifying at least one word in the first text (415) summarizing the utterance, while displaying the word in the second text (425) visually highlighted according to the input.
[0302] As described above, a non-transitory computer-readable recording medium can store a program including instructions. The instructions, when individually or collectively executed by at least one processor (120) of an electronic device (101) including a display (260), can cause the electronic device (101) to display, through the display (260), a first text (415) representing audio data (410) and a second text (425) summarizing a document (420), according to a split view. The instructions, when individually or collectively executed by the at least one processor (120), can cause the electronic device (101) to identify an input for a first word within the first text (415). The instructions, when individually or collectively executed by the at least one processor (120), may cause the electronic device (101) to display the at least one third word visually highlighted in the second text (425) based on identifying at least one second word corresponding to the first word identified according to the input among words in the document (420) and identifying at least one third word in the second text (425) corresponding to the at least one second word in the document (420), while displaying the first word in the first text (415) visually highlighted according to the input.
[0303] As described above, the method may be performed in an electronic device (101) including a display (260). The method may include an operation of displaying, through the display (260), a user interface including a first area representing text data (413) generated from audio data (410) and a second area corresponding to a user input received in relation to the audio data (410). The user input may include a plurality of strokes. The method may include an operation of playing the audio data (410). The method may include an operation of visually highlighting and displaying, in the second area, at least one stroke corresponding to a designated section among the plurality of strokes as a designated section included in the audio data (410) is played.
[0304] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0305] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0306] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0307] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0308] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., by download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0309] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device (101), Display (260), At least one processor (120) comprising a processing circuit; and A memory (130) storing instructions and including one or more storage media, wherein the instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Through the above display (260), according to the split view, the first text (415) summarizing the text data (413) representing the audio data (410) and the second text (425) representing the contents of the document (420) are displayed, Identifying input for words within the above second text (425), Identifying an utterance corresponding to the word identified according to the input among utterances (411) in text data (413) representing the audio data (410) and identifying at least one word in the first text (415) summarizing the utterance, causing the at least one word visually highlighted in the first text (415) to be displayed while displaying the word in the second text (425) visually highlighted according to the input. Electronic devices.
2. In claim 1, The above second text (425) is a text summarizing the above contents of the above document (420), The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Generate mapping information (430) that links words in the above document (420) with words in the second text (425), Based on the above mapping information (430), identify a word corresponding to the word identified according to the input among the words in the document (420), Causing to identify an utterance corresponding to the word identified based on the mapping information (430) among the utterances (411) in the text data (413), Electronic devices.
3. In claim 1 or claim 2, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: While recording the above audio data (410), receiving user input for the above words in the above document (420), Based on the recording time of the audio data (410) and the reception time of the user input for the words, timestamp information is obtained that maps the utterances (411) in the text data (413) and the words in the document (420). Based on the above timestamp information, causing the utterance (411) in the text data (413) to be identified that corresponds to the word identified based on the mapping information (430). Electronic devices.
4. In any one of claims 1 to 3, The user input for the words in the document (420) includes stroke inputs for entering the words into the document (420). Electronic devices.
5. In any one of claims 1 to 4, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Generate mapping information (430) that links utterances (411) in the above text data (413) and words in the first text (415), Based on the above mapping information (430), causing at least one word corresponding to the utterance corresponding to the word identified according to the input among the words in the first text (415), Electronic devices.
6. In any one of claims 1 to 5, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Based on inputting the text data (413) representing the audio data (410) into a generative artificial intelligence (AI) module, causing the first text (415) summarizing the text data (413) to be generated. Electronic devices.
7. In the electronic device (101), Display (260), At least one processor (120) comprising a processing circuit; and A memory (130) storing instructions and including one or more storage media, wherein the instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Through the display (260), a user interface is displayed including a first area representing text data (413) generated from audio data (410) and a second area corresponding to a user input received in relation to the audio data (410), wherein the user input includes a plurality of strokes, Playing the above audio data (410), and Depending on the section being played among the multiple sections included in the above audio data (410): Visually highlighting and displaying at least one word of the text data (413) in the first area, and Causing each of the plurality of strokes to be visually highlighted and displayed in the second area based on the order in which the plurality of strokes are input. Electronic devices.
8. In claim 7, The above plurality of strokes includes a first stroke and a second stroke received later than the first stroke, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Based on the order entered above: The first stroke is visually highlighted before the second stroke, Causing the second stroke to be visually highlighted while the first stroke is visually highlighted, Electronic devices.
9. In claim 7 or claim 8, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: While the playback of the above audio data (410) is maintained, the state of the first stroke is maintained by visually highlighting the first stroke. Electronic devices.
10. In any one of claims 7 to 9, Each of the above plurality of strokes is input while each of the plurality of sections included in the audio data (410) is being recorded. Electronic devices.
11. In any one of claims 7 to 10, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Through the above user interface, receiving the specified input, In response to the above specified input, causing a first text (415) summarizing the text data (413) to be displayed in the first area, Electronic devices.
12. In any one of claims 7 to 11, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Based on inputting the text data (413) representing the audio data (410) into a generative artificial intelligence (AI) module, causing the first text (415) summarizing the text data (413) to be generated. Electronic devices.
13. In any one of claims 7 to 12, At least some of the above plurality of strokes correspond to an input object, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Receives an input for selecting the above input object, Based on receiving the input selecting the input object, causing at least a portion of the first text temporally linked to the input object to be visually highlighted and displayed, Electronic devices.
14. In any one of claims 7 to 13, At least a portion of the above first text is selected according to mapped time information, The above time information includes the time information obtained from the plurality of sections included in the above audio data (410). Electronic devices.
15. In any one of claims 7 to 14, The above instructions, when individually or collectively executed by the at least one processor (120), cause the electronic device (101) to: Receiving an input for selecting at least a portion of the first text, Based on receiving said input selecting at least a portion of said first text, causing at least a portion of said plurality of strokes to be visually highlighted and displayed, said at least portion being temporally linked to said at least portion, Electronic devices.
Citation Information
Patent Citations
Information processing unit and program
JP2021128222A
Handwriting information processing device
JP6810515B2
Automatically creating a mapping between text data and audio data
KR101674851B1
Mobile terminal and control method for the mobile terminal
KR102063766B1
Electronic Device And Method Of Controlling The Same
KR102196671B1