Electronic device and method for providing summarization function
The electronic device uses an AI model to summarize video and text content within web pages, addressing the challenge of generating effective summaries, thereby improving user interaction.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2026-04-02
AI Technical Summary
Existing electronic devices struggle to provide effective summaries of multimedia content, particularly video and text content, within web pages, lacking efficient methods to generate concise summaries based on user input.
The electronic device employs an artificial intelligence model, trained through machine learning, to identify and summarize video and text content within web pages, generating summary information using a generative model that integrates text information and image frames.
The solution enables the device to generate comprehensive and user-requested summaries of web page content, enhancing user interaction by providing concise summaries of multimedia content.
Smart Images

Figure KR2025014680_02042026_PF_FP_ABST
Abstract
Description
Electronic device and method for providing a summary function
[0001] The following descriptions relate to an electronic device and method for providing a summary function.
[0002] An electronic device can provide various data, such as video, sound, or documents, to a user through applications. An electronic device can provide a web page, for example, through an application capable of performing web browser functions. A web page may contain various content. A web page may contain multimedia content, including video and text. An electronic device can display a web page containing various multimedia content through the display of the electronic device.
[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.
[0004] According to one embodiment, the electronic device may include a display, a memory including instructions and one or more storage media, and at least one processor including a processing circuit. The instructions may cause the electronic device to receive an input requesting a summary of said web page while a user interface including at least a portion of a web page is displayed through said display when executed individually or collectively by said at least one processor, and based on said input, to identify video content within said web page, obtain text information regarding said video content and at least one image regarding said video content, and provide summary information generated based on said text information regarding said video content and at least one image regarding said video content through said display.
[0005] According to one embodiment, a method performed by an electronic device may include receiving an input requesting a summary of a web page while a user interface including at least a portion of a web page is displayed through the display; identifying video content within the web page based on the input; acquiring text information regarding the video content and at least one image regarding the video content; and providing summary information generated based on the text information regarding the video content and at least one image regarding the video content through the display.
[0006] According to one embodiment, the electronic device may include a display, a memory including instructions and one or more storage media, and at least one processor including a processing circuit. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause to display a first user interface including text content and video content through the display, receive user input for displaying summary information in relation to the first user interface, obtain subtitle information from the video content, generate the summary information based on the subtitle information regarding the text content and the video content, and display a second user interface including the summary information through the display.
[0007] Figure 1 is a block diagram of an electronic device in a network environment.
[0008] FIG. 2 illustrates an example of the operation of an electronic device for providing summary information for a web page containing video content and text content.
[0009] Figure 3 illustrates an example of a simplified block diagram of an electronic device.
[0010] FIG. 4 illustrates a flowchart regarding the exemplary operation of an electronic device.
[0011] FIG. 5 illustrates an example of the operation of an electronic device for providing summary information.
[0012] FIG. 6 illustrates an example of the operation of an electronic device for acquiring at least one image through subtitle data of video content.
[0013] FIG. 7 illustrates an example of the operation of an electronic device for providing summary information.
[0014] FIG. 8 illustrates a flowchart of an exemplary operation of an electronic device for displaying summary information.
[0015] FIG. 9 illustrates an example of the operation of an electronic device for displaying summary information.
[0016] FIG. 10 illustrates an example of the operation of an electronic device for determining a video among a plurality of videos to provide summary information.
[0017] FIG. 11 illustrates an example of the operation of an electronic device for determining a video among a plurality of videos to provide summary information.
[0018] FIG. 12 illustrates an example of the operation of an electronic device for determining a video among a plurality of videos to provide summary information.
[0019] FIG. 13 illustrates an example of the operation of an electronic device according to the type of playback section of video content.
[0020] FIG. 14a illustrates an example of the operation of an electronic device according to the type of playback section of video content.
[0021] FIG. 14b illustrates an example of the operation of an electronic device according to the state of the electronic device.
[0022] FIG. 15 illustrates a flowchart of an exemplary operation of an electronic device for displaying summary information.
[0023] FIG. 16a illustrates an example of the operation of an electronic device for providing summary information.
[0024] FIG. 16b illustrates an example of the operation of an electronic device for providing summary information.
[0025] FIG. 17 illustrates an example of the operation of an electronic device for providing summary information.
[0026] FIG. 18 is a block diagram showing an integrated intelligence system according to one embodiment.
[0027] Figure 19 is a diagram showing the form in which relationship information between concepts and operations is stored in a database.
[0028] FIG. 20 is a diagram showing a screen in which a user terminal processes voice input received through an intelligent app.
[0029] Figure 21 is a schematic diagram of an exemplary AI system.
[0030] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.
[0031] Figure 1 is a block diagram of an electronic device in a network environment.
[0032] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0033] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0034] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0035] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0036] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0037] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0038] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0039] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0040] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0041] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0042] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0043] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0044] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0045] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0046] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0047] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0048] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0049] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for realizing eMBB, loss coverage (e.g., 164 dB or less) for realizing mMTC, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for realizing URLLC.
[0050] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0051] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0052] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0053] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0054] According to one embodiment, an electronic device (e.g., the electronic device (101) of FIG. 1) can display a web page through a display. The web page may include various multimedia content. For example, the web page may include at least one of text content, video content, image content, and audio content. The electronic device may provide summary information for a web page containing various multimedia content. For example, the electronic device may provide summary information for a web page containing video content and text content. The electronic device may provide summary information for a web page using an artificial intelligence model based on the video content and text content included in the web page. Technical features for providing summary information for a web page using an artificial intelligence model based on the video content and text content included in the web page will be described below.
[0055] The artificial intelligence model of the present disclosure may include a machine learning model (or deep learning model) trained to generate summarized text (e.g., at least one sentence) centered on the topic or main content represented by the input content when content (text, images) is input. The artificial intelligence model of the present disclosure may be trained to classify (or group) the input content or the summary generated from the input content based on at least one of the source of the content, the type of content, or information regarding the place / time / object / person represented by the content. The artificial intelligence model of the present disclosure may be a model trained to generate content (e.g., at least one sentence or image) describing the input content (text, images). The artificial intelligence model of the present disclosure may be a model trained to determine that the input content (text, images) originates from activities belonging to the same single event. Groups or events to be classified are included in a prompt and transmitted to the artificial intelligence model, and the artificial intelligence model may select one of the groups / events included in the prompt to classify the content into the selected group. The artificial intelligence model of the present disclosure may be a model trained to generate an image that can represent the subject or common event of the content using input content (text, image). The artificial intelligence model of the present disclosure may receive a web address or file as input and obtain information contained in the web address or file. The artificial intelligence model may be a model trained to identify the type of video (e.g., event-centered type, object description type, voice information included type). The artificial intelligence model of the present disclosure may be trained or learned to associate the contents of the input content with specific tasks through a sufficiently large amount of training data regarding contexts such as various types of content delivered in a prompt, terminal history information, terminal status information, and current screen information.An artificial intelligence model can output a specific task as a recommended task based on content and context. Learning performed through an electronic device may be initial learning or retraining. The artificial intelligence model of the present disclosure may include various transformer models. The operation of the artificial intelligence model of the present disclosure may include a learning and inference process of finding patterns in data, storing them as a model of generalized rules, and obtaining results by inputting new data into the learned model.
[0056] Although the description of the LLM (large language model) (or LVM (large vision models)) used in this disclosure is provided below, it is obvious that the artificial intelligence neural network of this disclosure may include not only language models but also various foundation models such as code models and image models, as well as other artificial intelligence neural network models.
[0057] The artificial intelligence model mentioned in the present disclosure may refer to an LLM, which is an artificial neural network-based language model that has learned a large amount of text data through prior training. An LLM may contain relatively more parameters (e.g., more than 10 billion) than existing general language models. An LLM may use a transformer artificial neural network structure based on an attention mechanism.
[0058] The attention mechanism is a technique that helps artificial intelligence models focus on important parts within input data. The attention mechanism can be utilized to predict output data by predicting the extent to which parts of time-series input data (e.g., input data such as voice or video, or input data for a specific layer of a neural network) contribute to the intermediate or final output of the neural network. While the recurrent neural network (RNN) structure, which processes each element of a sequence sequentially, suffers from degraded prediction performance when there is information dependency over long time-series distances, the attention mechanism can account for information dependency over long time-series distances by controlling the degree of weighted attention within the overall context (or part thereof) of the input data.
[0059] A transformer can be composed of an encoder-decoder structure. The encoder processes input data to output compressed information (e.g., contextual representation), and the decoder processes the compressed information to output data in token units. Each of the encoder and decoder may include an independent attention network and a cross-attention network connecting the encoder and decoder.
[0060] For example, LLM learning may include pre-training and / or fine-tuning. Pre-training is the process of enabling the LLM to acquire general linguistic knowledge using large amounts of text data, and may include, for example, self-supervised learning that predicts the next word using the previous sequence of words in a sequence of text. Fine-tuning is the process of training the LLM to be suitable for a specific domain (e.g., chatbot, translation, summarization, Q&A) or task; based on the pre-trained model, the LLM can be further supervised (or adaptive) using a dataset tailored to the domain purpose. The LLM can perform tasks using text input containing natural language called a prompt.
[0061] For example, fine-tuning can be omitted during LLM training. Users can control the prompts input to the LLM to improve the performance of desired tasks. Similar to in-context learning or zero-shot / few-shot learning, task examples and / or guides for performing the task can be added to the prompts. Examples of publicly available LLMs include BERT (Bidirectional Encoder Representations from Transformer) and GPT (Generative Pre-trained Transformer).
[0062] The term 'LLM' can refer to the language neural network model itself, but it can also refer to models for LLM-based applications (e.g., chatbots, translation, summarization, text classification, sentence generation). For instance, an LLM-based chatbot or translator like ChatGPT can also be referred to as 'LLM'.
[0063] 'LLM' may include an inference engine using an LLM neural network model. For example, "inputting an input prompt into the LLM" may mean "inputting an input prompt into an inference engine based on the LLM." For example, "the output of the LLM for the input prompt" may refer to the output information of the last neural network layer of the LLM (or output information modified through additional processing) obtained when the input prompt is input into an inference engine based on the LLM.
[0064] FIG. 2 illustrates an example of the operation of an electronic device for providing summary information for a web page containing video content and text content.
[0065] Referring to FIG. 2, the electronic device (200) may correspond to the electronic device (101) of FIG. 1. For example, the electronic device (200) may include at least some or all of the components of the electronic device (101) of FIG. 1.
[0066] In example (210), the electronic device (200) may display a user interface (219) that represents at least part (or all) of a web page through a display (202). For example, at least part (or all) of the web page represented by the user interface (219) may include video content (211) and text content (212). The electronic device (200) may identify the video content (211) and text content (212). While at least part (or all) of the web page containing the video content (211) and text content (212) is displayed through the user interface (219), the electronic device (200) may receive input to request a summary of the web page.
[0067] For example, the user interface (219) may include an object for requesting a summary of a web page. The electronic device (200) may identify an input for requesting a summary of a web page based on an input to the object.
[0068] According to an embodiment, the electronic device (200) may receive input requesting a summary of a screen (e.g., user interface (219) or web page) displayed through the display (202).
[0069] According to one embodiment, the electronic device (200) may perform an operation to provide summary information regarding the contents included in the web page based on an input requesting a summary of the web page. For example, the electronic device (200) may identify video content (211) and text content (212) included in the web page. The electronic device (200) may obtain text information regarding the video content (211) and at least one image regarding the video content (211).
[0070] For example, text information may include text content obtained based on video content (211) and text content (212) displayed within a web page in conjunction with video content (211). For example, an electronic device (200) may obtain text content based on subtitle data of video content (211). For example, an electronic device (200) may obtain text content based on audio data of video content (211).
[0071] For example, at least one image regarding video content (211) may include at least one frame among a plurality of frames of video content (211). For example, at least one image regarding video content (211) may be at least one frame among a plurality of frames of video content (211). For example, at least one image regarding video content (211) may include a thumbnail image regarding video content (211). For example, a thumbnail image regarding video content (211) may be provided by the creator of video content (211). A thumbnail image regarding video content (211) may correspond to an image displayed before playback of video content (211). An electronic device (211) may obtain an image displayed before playback of video content (211) as a thumbnail image.
[0072] According to one embodiment, the electronic device (200) can generate (or obtain) summary information based on text information regarding video content (211) and at least one image regarding video content (211). For example, the electronic device (200) can input text information regarding video content (211) and at least one image regarding video content (211) into an artificial intelligence model. The electronic device (200) can generate (or obtain) summary information about a web page based on the output of the artificial intelligence model.
[0073] For example, the artificial intelligence model may be constructed based on a generative model (or a generative artificial intelligence model). However, it is not limited thereto. For example, the artificial intelligence model may be included in an electronic device (200) or in an external electronic device (e.g., a server) connected to the electronic device (200).
[0074] In example (220), the electronic device (200) may display a user interface (223) for providing generated summary information (224) through a display (202) superimposed on the user interface (219). For example, the summary information (224) may include summary information for video content (211) and / or summary information for text content (212). For example, the summary information (224) may include comprehensive summary information for a web page.
[0075] For convenience of explanation, the following specification describes examples for providing summary information for web pages containing video content, but is not limited thereto. For example, embodiments of the present disclosure may be applied to various user interfaces (or applications) containing multimedia content as well as web pages. For example, within an email application, while video content is displayed through a display (202), the electronic device (200) may identify an input requesting a summary of the screen (or email) displayed through the display (202). Based on the identified input, the processor (201) may obtain information regarding an email containing video content. For example, information regarding an email containing video content may include a subject, sender, recipient, CC recipient, time of transmission, time of reception, a file attached to the email (e.g., a video file), and / or video content in the body of the email. Based on the information regarding the email, the electronic device (200) may obtain summary information regarding the email. The electronic device (200) can provide (or display) summary information about an email in conjunction with the user interface of an email application. For example, if video content is attached within an email, the processor (201) can provide summary information about the video content configured as an attachment. To provide summary information about the attachment, the processor (201) can check whether the received email is spam and perform a security check (e.g., virus check) on the attachment. The processor (201) can perform a security check to provide summary information in advance, even if the user does not execute the attachment within the email, and can provide summary information about the video content attached as an attachment.
[0076] In the specification below, the specific operation of an electronic device (200) for providing summary information about a web page as shown in FIG. 2 will be described later. First, the components of the electronic device (200) according to the embodiments described above will be described later in FIG. 3.
[0077] Figure 3 illustrates an example of a simplified block diagram of an electronic device.
[0078] Referring to FIG. 3, the electronic device (200) may include at least some or all of the components of the electronic device (101) of FIG. 1. For example, the electronic device (200) may correspond to the electronic device (101) of FIG. 1. The electronic device (200) may be a terminal owned by a user. The terminal may include, for example, a personal computer (PC) such as a laptop and a desktop, a smartphone, a smartpad, or a tablet PC. The terminal may include smart accessories such as a smartwatch and / or a head-mounted device (HMD).
[0079] According to one embodiment, the electronic device (200) may include at least one of a processor (201), a display (202), a memory (203), and / or a communication circuit (204). For example, at least some of the processor (201), the display (202), the memory (203), and / or the communication circuit (204) may be omitted according to the embodiment.
[0080] According to one embodiment, the processor (201) may include at least a portion of the processor (120) of FIG. 1 or correspond to at least a portion of the processor (120). For example, the processor (201) may include one or more processors including an application processor (AP) and / or a communication processor (CP). For example, the processor (201) may be implemented as a single chip, such as a system on chip (SoC), or as multiple chips. For example, the processor (201) may be implemented as a single integrated circuit or as multiple integrated circuits. For example, the processor (201) may be distributedly arranged within an electronic device (200).
[0081] The processor (201) may be operatively coupled with or connected with the display (202), memory (203), and communication circuit (204). For example, the processor (201) being operatively coupled with other components may mean that the processor (201) can control other components. The processor (201) can control the display (202), memory (203), and / or communication circuit (204).
[0082] According to one embodiment, a display (202) of an electronic device (200) can output visualized information (e.g., a screen) to a user. For example, the display (202) can be controlled by a controller, such as a GPU (graphic processing unit), to output visualized information to a user. The display (202) may include a liquid crystal display (LCD), a plasma display panel (PDP), and / or one or more light emitting diodes (LEDs). The LEDs may include organic LEDs (OLEDs). The display (202) may include a flat panel display (FPD) and / or electronic paper. The embodiments are not limited thereto, and the display (202) may have at least a partially curved shape or a deformable shape. A display (202) having a deformable shape may be referred to as a flexible display.
[0083] The display (202) of FIG. 3 may be an example of the display module (160) of FIG. 1. A display (202) that supports touch functions may be referred to as a touch screen. The display (202) may further include a structure capable of detecting input using a stylus pen, such as an electro-magnetic resonance (EMR) or an active electrostatic solution (AES). For example, at least some of the events of the present disclosure (e.g., inputs for requesting a summary of a web page) may be initiated using a stylus pen or a user's finger. For example, a user may specify a part of the content displayed on the display (202) that they wish to summarize through input based on a stylus pen (e.g., a drawing that surrounds at least a part of the desired area). For example, if a user uses a stylus pen to draw a specific shape (e.g., a circle) around content (e.g., an image, text, video, file, and / or handwriting input) and handwrites the word “summary” around it, the electronic device (200) can identify this as a request to summarize the content specified by the handwriting. For example, the user may handwrite other commands in addition to “summary” (e.g., specifying the language to summarize, requesting a search, requesting a translation).
[0084] According to one embodiment, the memory (203) of the electronic device (200) may include a circuit and / or a storage medium for storing data and / or instructions that are input and / or output to the processor (201). The memory (203) may include, for example, volatile memory such as random-access memory (RAM) and / or non-volatile memory such as read-only memory (ROM). Non-volatile memory may be referred to as storage. Volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). Non-volatile memory may include, for example, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disk, solid state drive (SSD), and embedded multi-media card (eMMC).
[0085] According to one embodiment, the memory (203) may include at least a portion of the memory (130) of FIG. 1 or correspond to at least a portion of the memory (130) of FIG. 1. For example, the memory (203) may be implemented as a single chip or as a plurality of chips. For example, the memory (203) may be implemented as a single integrated circuit or as a plurality of integrated circuits. For example, the memory (203) may be distributedly arranged within an electronic device (200).
[0086] According to one embodiment, a processor (201) of an electronic device (200) may execute instructions in a memory (203) within the electronic device (200) to perform a function and / or operation indicated by said instructions. For example, if the electronic device (200) includes at least one processor, said at least one processor may be configured to execute said instructions collectively or individually.
[0087] For example, memory (203) may include at least one artificial intelligence model. Memory (203) may store instructions regarding at least one artificial intelligence model. For example, at least one artificial intelligence model may include an artificial intelligence model, which will be described below. The artificial intelligence model may be a language-based artificial intelligence model. According to an embodiment, at least one artificial intelligence model may be included in a chip (e.g., NPU) distinct from memory (203). For example, at least one artificial intelligence model may be implemented as (on-device artificial intelligence) hardware (e.g., an artificial intelligence chip) included inside a separate device or as an artificial intelligence model included in an external server.
[0088] According to one embodiment, the communication circuit (204) can be used for various radio access technologies (RAT). For example, the communication circuit (204) can be used to perform Bluetooth communication, wireless local area network (WLAN) communication, or ultra-wideband (UWB) communication. For example, the communication circuit (204) can be used to perform cellular communication. For example, the processor (201) can establish a connection with an external electronic device (e.g., a server) through the communication circuit (204). For example, the communication circuit (204) can correspond to the communication module (190) of FIG. 1.
[0089] FIG. 4 illustrates a flowchart regarding an exemplary operation of an electronic device. In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0090] Referring to FIG. 4, in operation 410, the processor (201) may receive input to request a summary of a web page. For example, the processor (201) may receive input to request a summary of a web page while a user interface containing at least a portion of a web page is displayed through the display (202).
[0091] According to one embodiment, the processor (201) may display a user interface (e.g., the user interface (219) of FIG. 2) that includes at least a portion (or all) of a web page. For example, depending on the size of the display (202), the area of the web page displayed on the display (202) may change. For example, depending on the scale of the web page (e.g., magnification or reduction), the area of the web page displayed on the display (202) may change.
[0092] According to one embodiment, the processor (201) may receive an input for requesting a summary of a web page. The input for requesting a summary of a web page may be configured in various ways.
[0093] For example, a user interface including at least a portion of a web page may include an object for requesting a summary of the web page. The processor (201) may receive an input for requesting a summary of the web page based on an input to said object.
[0094] For example, the electronic device (200) may include a physical button for requesting a summary of a web page. The electronic device (200) may receive an input for requesting a summary of a web page based on input to the physical button while the user interface is displayed through the display (202).
[0095] For example, the electronic device (200) may receive input to request a summary of a web page via a voice-based intelligent assistant. The electronic device (200) may invoke the voice-based intelligent assistant based on input to a specific physical button and / or input to a specific area of the touchscreen. As an example, the electronic device (200) may execute an interactive application for the use of an artificial intelligence model based on input to a specific physical button and / or input to a specific area of the touchscreen.
[0096] For example, the electronic device (200) can receive input to request a summary of a web page based on a specified gesture input while the user interface is displayed through the display (202).
[0097] According to an embodiment, the electronic device (200) may receive input requesting a summary of a screen (e.g., a user interface or a web page) displayed through the display (202).
[0098] In operation 420, the processor (201) can identify video content within a web page. For example, the processor (201) can identify video content within a web page based on input requesting a summary of the web page.
[0099] According to one embodiment, the processor (201) can identify that video content is included within the web page. For example, the video content may be provided by a service provider for the web page. According to one embodiment, the web page may include various multimedia content. For example, the web page may include at least one of text content, video content, image content, and audio content.
[0100] According to one embodiment, the processor (201) can identify markup language information (e.g., HTML (hypertext mark-up language)) regarding a web page. The markup language information regarding a web page may include a plurality of tags. The processor (201) identifies a tag regarding video (e.g., among the plurality of tags) <video>Based on identifying ), it is possible to identify that video content is included within a web page.
[0101] In operation 430, the processor (201) can obtain text information regarding video content and at least one image regarding video content.
[0102] According to one embodiment, text information regarding video content may include at least one of a first text content obtained based on video content and a second text content displayed within a web page in conjunction with video content.
[0103] For example, the first text content can be obtained based on the subtitle data of the video content. The processor (201) can obtain the subtitle data of the video content. For example, the processor (201) can obtain the subtitle data of the video content provided by a service provider for a web page. The processor (201) can obtain the subtitle data of the video content through an API (application programming interface) (or OpenAPI) provided by the service provider. For example, the processor (201) can obtain the subtitle data of the video content using WebVTT (web video text tracks) information.
[0104] For example, subtitle data may include at least one of a plurality of start times for displaying subtitles within video content, a plurality of display durations corresponding to the plurality of start times, or a plurality of characters (e.g., character blocks) corresponding to the plurality of display durations. As an example, at least part of the subtitle data may be generated by a service provider or a content creator (e.g., a video content creator or a subtitle creator). Specific examples of subtitle data will be described later in FIG. 6.
[0105] According to one embodiment, the processor (201) can identify (or extract, acquire) audio data for video content. The processor (201) can acquire text information regarding video content based on the audio data. For example, the processor (201) can perform a speech-to-text (STT) operation based on the audio data. Based on the STT operation, the processor (201) can acquire subtitle data of the video content through the audio data. The processor (201) can acquire a first text content using the subtitle data.
[0106] According to one embodiment, the processor (201) can obtain text information regarding video content based on audio data using an artificial intelligence model (e.g., a generative artificial intelligence model). The artificial intelligence model may be included in an electronic device (200) or in an external electronic device (e.g., a server).
[0107] For example, the processor (201) can identify at least one image of the video content when the video content does not contain subtitle data or when multiple frames of the video content do not contain text. The processor (201) can obtain a description of at least one image based on inputting at least one image into an artificial intelligence model. Based on the description of at least one image, the processor (201) can obtain a first text content.
[0108] According to one embodiment, the processor (201) can identify that the amount of characters (e.g., number of characters, number of sentences, size of the area occupied by the characters) according to the second text content is less than a reference. For example, the processor (201) can identify the first text content as text information based on identifying that the number of characters according to the second text content is less than a reference number (e.g., 50 characters). For example, the processor (201) can identify the first text content as text information if the number of characters of the second text content is not sufficient to provide summary information. According to an embodiment, the processor (201) can identify both the first text content and the second text content as text information if the number of characters of each of the first text content and the second text content is not sufficient to provide summary information.
[0109] According to one embodiment, the processor (201) can identify the similarity between the first text content and the second text content. The processor (201) can identify the first text content as text information based on whether the identified similarity exceeds a reference similarity. The processor (201) can identify the first text content and the second text content as text information based on whether the identified similarity is less than or equal to a reference similarity. For example, by identifying the similarity between the first text content and the second text content, the processor (201) can identify whether the second text content, which is displayed in conjunction with video data, corresponds to the subtitle data (or voice data) of the video content. The processor (201) can identify the first text content as text information if the second text content, which is displayed in conjunction with video data, substantially corresponds to the subtitle data (or voice data) of the video content. The processor (201) can identify the first text content and the second text content as text information if the second text content displayed in association with the video data does not correspond to the subtitle data (or voice data) of the video content.
[0110] According to one embodiment, the processor (201) can identify at least one visual object representing at least one character included in a plurality of frames of video content. Based on at least one character identified according to at least one visual object, the processor (201) can obtain text information regarding the video content (e.g., first text content). For example, the processor (201) can identify characters included in a plurality of frames of video content. Based on the identified characters, the processor (201) can obtain text information regarding the video content. The processor (201) can obtain text information not only for subtitles but also for characters displayed within a plurality of frames. According to an embodiment, the processor (201) can identify at least one frame containing at least one character. The processor (201) can use at least one frame to obtain summary information along with the text information.
[0111] According to an embodiment, the processor (201) may obtain at least some (or all) of text information regarding video content before receiving an input to request a summary of the web page, based on the fact that a user interface including at least a portion of the web page is provided through the display (202). The processor (201) may also obtain at least some of the text information regarding video content in advance, even before receiving an input to request a summary of the web page. For example, the processor (201) may obtain at least some (or all) of the text information regarding video content (e.g., subtitle data) before receiving an input to request a summary of the web page, if the video is included within the web page or if the video is included on at least a portion of the web page displayed through the display (202).
[0112] According to one embodiment, the processor (201) can acquire at least one image of video content. For example, the processor (201) can acquire at least one image of video content based on text information (e.g., first text content) of video content.
[0113] For example, the processor (201) can obtain text information based on subtitle data. The subtitle data may include a plurality of start times, a plurality of display periods, and a plurality of character blocks for displaying subtitles within video content.
[0114] The processor (201) can identify a reference number of character blocks among a plurality of character blocks based on a plurality of display periods. The processor (201) can identify a reference number of character blocks whose display period is longer than that of other character blocks. The processor (201) can acquire at least one image of video content based on at least one frame of video data corresponding to the reference number of character blocks.
[0115] According to one embodiment, the processor (201) can identify the sum of playback segments of video content corresponding to text information. The processor (201) can identify whether the sum of playback segments of video content corresponding to text information exceeds a reference playback segment. For example, the processor (201) can identify a plurality of frames corresponding to text information based on identifying that the sum of playback segments of video content corresponding to text information exceeds a reference playback segment. The processor (201) can identify at least one frame among the plurality of frames as at least one image. For example, the processor (201) can identify at least one frame (or at least one other frame) within the entire playback segment of video content based on identifying that the sum of playback segments of video content corresponding to text information is less than or equal to a reference playback segment. The processor (201) can identify the identified at least one frame as at least one image. For example, if the processor (201) is not sufficient to acquire at least one image, the sum of the playback intervals of the video content may be less than or equal to a reference playback interval. Accordingly, the processor (201) can identify at least one frame within the entire playback interval of the video content. For example, the processor (201) can identify at least one frame within the entire playback interval of the video content that includes at least one of a frame containing text, a frame containing a person (or a main character of the video content), or a frame displayed during a time interval exceeding a reference time interval. The processor (201) may acquire at least one identified frame as at least one image. According to an embodiment, the processor (201) can acquire at least one frame acquired based on subtitle data and at least one other frame acquired within the entire playback interval of the video content.The processor (201) can identify the at least one frame and the at least one other frame as at least one image for obtaining summary information.
[0116] According to an embodiment, the processor (201) can identify (or acquire) audio data of video content. The processor (201) can identify voice data among the audio data of video content. The processor (201) can identify a plurality of frames regarding voice data. The processor (201) can identify at least one frame among the plurality of frames regarding voice data identified according to the display period of each of the plurality of frames. The processor (201) can acquire at least one image based on at least one frame. For example, the processor (201) can acquire at least one frame as at least one image.
[0117] In operation 440, the processor (201) may provide summary information generated based on text information regarding video content and at least one image regarding video content through the display (202). According to an embodiment, the processor (201) may output a voice signal for the summary information using a speaker.
[0118] According to one embodiment, summary information may be composed based on a plurality of sentences. The summary information may include a first section comprising at least one sentence regarding video content among the plurality of sentences, and a second section comprising the remaining sentences regarding second text content among the plurality of sentences. According to an embodiment, the processor (201) may determine the display method of the first section and the second section. For example, the processor (201) may determine the display method of the first section and the second section based on identification information for the web page (e.g., a uniform resource locator (URL), characteristics of the web page, information provided on the web page). For example, the display method of the first section and the second section may include a display order and / or layout.
[0119] According to one embodiment, the processor (201) can identify at least one character block (e.g., [music]) containing a designated identification symbol for indicating a sound effect among a plurality of character blocks included in subtitle data. Information regarding sound effects according to at least one character block may be included in summary information. For example, the processor (201) may provide at least one of the title, mood, rhythm, and / or lyrics of music played through video content as summary information.
[0120] FIG. 5 illustrates an example of the operation of an electronic device for providing summary information.
[0121] Referring to FIG. 5, the processor (201) may display at least a portion of a web page (510) through a user interface (500). The web page (510) may include an area (511) for displaying video content (520) and an area (512) for displaying other content (e.g., a second text content (532)) in conjunction with the video content (520). For example, the video content (520) may be displayed within the area (511). Within the area (512), text (513) indicating the title of the video content (520) and text (514) indicating a description of the video content (520) may be displayed. Although not illustrated, text indicating at least one user's reaction (e.g., a comment) to the video content (520) may also be displayed within the area (512).
[0122] According to one embodiment, the processor (201) can identify video content (520) and second text content (532) within a web page (510). For example, the processor (201) can identify video content (520) and second text content (532) based on markup language (e.g., html (hypertext mark-up language)) information regarding the web page (510).
[0123] According to one embodiment, the processor (201) can identify a first text content (531) regarding video content (520). The processor (201) can obtain subtitle data of the video content (520). The processor (201) can identify the first text content (531) based on the subtitle data. For example, the processor (201) can identify at least some of the characters included in the subtitle data as the first text content (531).
[0124] For example, the processor (201) can obtain subtitle data provided by a service provider regarding a web page (510). For example, the processor (201) can obtain subtitle data of video content (520) through an API (or OpenAPI). For example, the processor (201) can obtain subtitle data of video content (520) using WebVTT (web video text tracks) information. For example, the processor (201) can obtain subtitle data based on metadata of the web page (510) (or video content (520)).
[0125] According to an embodiment, if subtitle data is not provided by the service provider of the web page (510), the processor (201) can identify audio data of the video content (520). The processor (201) can acquire (or extract) audio data based on the video content (520). For example, the audio data can be identified based on at least one of music without lyrics, music with lyrics, voice corresponding to narration, and / or speech of a character in the video content (520). For example, the audio data can be acquired based on voice information included in the video content (520). The voice information can be acquired from voice corresponding to narration and / or speech of a character in the video content (520). The processor (201) can acquire audio data corresponding to playback segments containing voice information.
[0126] The processor (201) can perform STT operations based on audio data. Based on the STT operations, the processor (201) can obtain subtitle data (e.g., characters, text data, text data associated with time, timestamps, subtitle effects) of the video content (520) through the audio data. The processor (201) can obtain a first text content using the subtitle data.
[0127] According to one embodiment, the processor (201) can acquire at least one image (540) regarding video content (520). For example, the processor (201) can acquire at least one image (540) based on at least one frame among a plurality of frames of video content (520).
[0128] For example, at least one image (540) may include at least one frame among a plurality of frames of video content (520). For example, at least one image (540) may include a thumbnail image of video content (520). For example, at least one image (540) may include an image obtained by decoding compressed information related to video frames in video content (520), an image predicted from the decoded image, or an image predicted using the decoded image and the predicted image. For example, at least one image (540) may include a key frame of video content (520). For example, the key frame may be set by the provider or producer of video content (520). For example, it may be set within a frequently viewed section within the playback section of video content (520). For example, a key frame may be included in metadata regarding a web page (510) (or video content (520)). The processor (201) may identify (or obtain) the key frame based on the metadata regarding the web page (510) (or video content (520)).
[0129] According to one embodiment, when a first text content is obtained based on audio data, the processor (201) can identify playback segments containing voice information (e.g., voice information used in STT) among the audio data. Within the playback segments containing voice information, the processor (201) can identify a reference number of frames whose display time is longer than other frames. The processor (201) can identify the reference number of frames as at least one image (540).
[0130] According to one embodiment, the processor (201) may input text information (530) including at least one of a first text content (531) and / or a second text content (532) and at least one image (540) into the artificial intelligence model (550). According to an embodiment, the processor (201) may also input information for requesting address information and a summary of a web page (510) into the artificial intelligence model (550).
[0131] For example, the processor (201) can generate a prompt for text information (530) and at least one image (540) and input the generated prompt into an artificial intelligence model (550). The prompt can be configured as shown in the table below.
[0132] Assume you are a functional expert to summarize webpage contents.Your task is to read <input> of webpage contents and provide summary of it.Steps:1. Read and understand text and images in <input> thoroughly.2. The images of <input> are frames extracted through sampling from the video.3. Think about what text and images in <input> want to say comprehensively.4. Summarize the whole contents of <input> into a brief title and five sentences in {source_language}.5. Each sentence should be under 20 words for summary.6. Avoid to use any other information out of <input> when making a summary. <input> ===Your answer should be in JSON format with following keys:1. "title" keyA brief title2. "summary" keyArray of 5 sentences3. Order of keys in the response must be "title", "summary"4. Arrange JSON <output>:
[0133] The prompts according to Table 1 are exemplary and are not limited thereto. According to one embodiment, the artificial intelligence model (550) may be configured to generate summary information (560) based on the prompt (e.g., the prompt in Table 1).
[0134] According to one embodiment, the artificial intelligence model (550) may be configured within an electronic device (200). According to one embodiment, the artificial intelligence model (550) may be configured within an external electronic device (e.g., a server) that is distinct from the electronic device (200). According to an embodiment, at least a portion of the artificial intelligence model (550) may be configured within the electronic device (200), and the remainder of the artificial intelligence model (550) may be configured within an external electronic device (e.g., a server).
[0135] According to one embodiment, the artificial intelligence model (550) may be constructed based on a generative model. For example, the generative model may include a generative model comprising a plurality of parameters related to a neural network having a structure based on an encoder and a decoder, such as a transformer. In one embodiment, the generative model may include a bidirectional model based on learning about an encoder (e.g., BERT (bidirectional encoder representations from transformers)), or an auto-encoding model (e.g., a diffusion model). In one embodiment, the generative model may include an auto-regressor model based on learning about a decoder (e.g., GPT (generative pre-trained transformer)). In one embodiment, the generative model may include a sequence-to-sequence model based on learning about an encoder and a decoder (e.g., stable diffusion, DALL-E 2). In one embodiment, the generative model may include a large language model (LLM) for processing natural language based on massive parameters, but is not limited thereto. The generative model may include parameters for driving neural networks such as a convolutional neural network (CNN), a recurrent neural network (RNN), a feedforward neural network (FNN), and / or a long short-term memory (LSTM).
[0136] According to one embodiment, the summary information (560) may include a plurality of sentences representing the summary and / or summary image(s). For example, the summary image(s) may be generated based on text information regarding the web page (510) and / or at least one image (540) regarding the web page (510). For example, the summary image(s) may include still images (or images) and / or videos.
[0137] According to one embodiment, the summary information (560) may be provided based on content (e.g., video content (520), first text content (531), or second text content (532)) regarding the web page (510). The processor (201) may also provide information about the content used to obtain a plurality of sections included in the summary information (560).
[0138] For example, the summary information (560) may include a first section containing at least one sentence regarding video content (520) and a second section containing at least one sentence regarding second text content (532). The first section may include summary information regarding video content (520). The second section may include summary information regarding second text content (532) indicated in association with video content. According to an embodiment, the plurality of sections may further include a section containing summary information regarding at least one user reaction (e.g., comment) to video content (520).
[0139] According to one embodiment, the processor (201) may provide together the content used for acquiring each of the plurality of sections. For example, the processor (201) may provide information indicating that the first section of the summary information (560) was acquired based on video content (520). The processor (201) may provide information indicating that the second section of the summary information (560) was acquired based on second text content (532). According to an embodiment, if an attachment is included within the web page (510), the plurality of sections may include a section for indicating summary information about the attachment. The processor (201) may provide information indicating that the section for indicating summary information about the attachment was acquired (or generated) based on the attachment. For example, if the attachment is composed of an image format, the processor (201) may acquire summary information about at least one object (e.g., person, object, background) included in the attachment and provide a section for indicating summary information about the attachment. The processor (201) may also provide information indicating that the section was obtained based on an attachment in an image format.
[0140] According to one embodiment, the summary information (560) may include not only the summary information obtained based on the text information (530) and at least one image (540), but also other content regarding the content included in the web page (510). The other content may include additional information related to the content included in the web page (510). The processor (201) may obtain the other content based on searching for additional information related to the content included in the web page (510).
[0141] For example, if the video content (520) represents a review of the product, content related to the product (e.g., release date, release country, advertisement, and / or sales volume) that is not included in the web page (510) can be obtained (or retrieved). The processor (201) can obtain (or generate) summary information (560) based on the content included in the web page (510) and the content related to the product.
[0142] For example, if the video content (520) represents a music video, the processor (201) can obtain (or search) content related to the music (e.g., release date, other music in the album containing the music, issues related to the music, artist information, recent activities of the artist) that is not included in the web page (510). The processor (201) can obtain (or generate) summary information (560) based on the content included in the web page (510) and the content related to the music.
[0143] According to one embodiment, when content not included in the web page (510) is used to obtain summary information (560), the source of the content not included in the web page (510) may be provided along with the summary information (560).
[0144] According to an embodiment, the processor (201) may store a plurality of sentences and / or summary image(s) representing a summary included in the summary information in the memory (203) of the electronic device (200) based on user input. According to an embodiment, the plurality of sentences and / or summary image(s) representing a summary included in the summary information may be edited according to user input. According to an embodiment, the processor (201) may transmit a plurality of sentences and / or summary image(s) representing a summary included in the summary information to an external electronic device based on user input.
[0145] According to one embodiment, if the video content (520) included in the web page (510) includes copyright restrictions, the processor (201) may provide information indicating that the video content (520) includes copyright restrictions, along with summary information (560). According to an embodiment, if the video content (520) included in the web page (510) includes copyright restrictions, the video content (520) may not be used to obtain (or generate) summary information (560).
[0146] According to one embodiment, the processor (201) may provide summary information (560) as a file (e.g., audio file, text file, video file). The above-described embodiments are described as the electronic device (200) providing summary information (560) for a web page (510) containing video content (520), but are not limited thereto. According to an embodiment, the processor (201) may obtain summary information for video content obtained through the camera of the electronic device (200) in a manner similar to that of the above-described embodiments.
[0147] According to one embodiment, the processor (201) can obtain summary information for audio content (e.g., podcast) containing text content as well as video content (520) in a manner similar to the above-described embodiment.
[0148] FIG. 6 illustrates an example of the operation of an electronic device for acquiring at least one image through subtitle data of video content.
[0149] Referring to FIG. 6, the processor (201) can obtain subtitle data (610) regarding video content (520). For example, the processor (201) can obtain subtitle data (610) of video content (520) provided by a service provider for a web page. The processor (201) can obtain subtitle data (610) of video content (520) through an API (application programming interface) (or OpenAPI) provided by a service provider. As an example, the processor (201) can obtain subtitle data (610) of video content (520) using WebVTT (web video text tracks) information.
[0150] For example, subtitle data (610) may include a plurality of start times for displaying subtitles within video content, a plurality of display durations corresponding to the plurality of start times, and a plurality of character blocks corresponding to the plurality of display durations.
[0151] For example, each of the multiple start times within the subtitle data (610) can be expressed as "text start". Each of the multiple display periods can be expressed as "dur". Each of the multiple character blocks is <text start=""" ” dur=""">and< / text> It can be expressed between.
[0152] For example, among a plurality of character blocks, at least one character block (e.g., [music]) containing a designated identifier (e.g., square brackets) for representing a sound effect can be identified.
[0153] For example, the start time of the character block (611) may be 0.53 [seconds] and the display period may be 2.209 [seconds]. The start time of the character block (612) may be 4.16 [seconds] and the display period may be 6.25 [seconds]. The start time of the character block (613) may be 5.17 [seconds] and the display period may be 5.24 [seconds]. The start time of the character block (614) may be 12.799 [seconds] and the display period may be 7.201 [seconds]. The start time of the character block (615) may be 20 [seconds] and the display period may be 3.199 [seconds].
[0154] According to one embodiment, the processor (201) can identify a plurality of character blocks included in subtitle data (610). The processor (201) can obtain a first text content (531) using the plurality of character blocks. The first text content (531) can be configured based on the plurality of character blocks of the subtitle data (610). An example of the first text content (531) can be configured as shown in the table below.
[0155] [Music] A boat floating on the blue sea [Music] Nature is beautiful. There is no traffic congestion here....
[0156] According to one embodiment, the processor (201) can identify frames of video content (520) based on a plurality of start times. For example, the processor (201) can identify a frame (621) with a start time (e.g., 0.53 [seconds]) corresponding to a character block (611). The processor (201) can identify a frame (622) with a start time (e.g., 4.16 [seconds]) corresponding to a character block (612). The processor (201) can identify a frame (623) with a start time (e.g., 5.17 [seconds]) corresponding to a character block (613). The processor (201) can identify a frame (624) with a start time (e.g., 12.799 [seconds]) corresponding to a character block (614). The processor (201) can identify a frame (625) of a start time (e.g., 20 [seconds]) corresponding to a character block (615).
[0157] The processor (201) can identify frames of video content (520) based on multiple start times and identify at least one frame among the identified frames. For example, the processor (201) can identify a reference number of character blocks (e.g., up to 10) in order of the longest display period. The processor (201) can identify a reference number of frames corresponding to each of the reference number of character blocks. For example, the reference number may be changed depending on the performance of the video content and / or the artificial intelligence model (550).
[0158] For example, the processor (201) can identify two frames to identify at least one image (540) to be input into an artificial intelligence model (550). The processor (201) can identify frames of video content (520) (e.g., frames (621) to (625)) based on multiple start times. The processor (201) can identify two character blocks with the longest display periods. The processor (201) can identify character block (612) and character block (614). The processor (201) can identify a frame (622) corresponding to character block (612). The processor (201) can identify a frame (624) corresponding to character block (614). The processor (201) can identify frame (622) and frame (622) as at least one image (540) to be input into an artificial intelligence model (550).
[0159] According to one embodiment, the processor (201) can obtain text information regarding video content (520). For example, the processor (201) can obtain text information (530) using a first text content (531) and a second text content (532). For example, the second text content (532) may include text indicating the title of the video content (520) displayed on the web page (510), a description of the video content (520), and a user's reaction (e.g., comments) regarding the video content (520). The processor (201) can obtain text information (530) regarding the video content (520) by combining the first text content (531) and the second text content (532).
[0160] According to the above-described embodiment, the processor (201) can provide comprehensive summary information (560) regarding the content included in the web page (510). Accordingly, since the summary information (560) is obtained using not only the summary information regarding the description of the video content (520) but also at least one image (540) obtained from the video content (520) and the first text content (531) obtained through the video content (520), advanced summary information can be provided.
[0161] FIG. 7 illustrates an example of the operation of an electronic device for providing summary information.
[0162] Referring to FIG. 7, the processor (201) can determine a display method for a plurality of sections included in summary information (560) based on the content included in the web page (510). For example, section (701) may include summary information for video content (520). Section (702) may include summary information for second text content (532). Section (703) may include summary information for additional information related to the content included in the web page (510). The summary information included in sections (701) through (703) may be configured based on multimedia content. For example, each of sections (701) through (703) may include at least one of text, images, videos, and / or links.
[0163] According to one embodiment, the processor (201) can identify the main content of the web page (510) based on the proportion that each of the contents included in the web page (510) occupies within the web page (510). For example, the processor (201) can identify the video content (520) as the main content of the web page (510) if the area occupied by the video content (520) in the web page (510) is the largest. For example, the processor (201) can identify the second text content (532) as the main content of the web page (510) if the area occupied by the second text content (532) in the web page (510) is the largest.
[0164] According to an embodiment, the processor (201) may determine at least one content to be used to obtain summary information based on identification information for the web page (510) (e.g., URL, characteristics of the web page (510), information provided on the web page (510). For example, if the web page (510) is provided by a media company, the processor (201) may identify a second text content (532) as the main content.
[0165] For example, based on identifying that the similarity between the subtitle data of the second text content (532) and the video content (520) is greater than or equal to a reference similarity, the processor (201) can identify the second text content (532) as the main content.
[0166] According to one embodiment, the processor (201) can obtain (or generate) summary information (560) centered on the main content by applying weights to the main content of the web page (510). For example, within an interface for displaying the summary information (560), the processor (201) can set the proportion of sections related to the main content to be large and the proportion of sections related to the remaining content that is distinct from the main content to be small.
[0167] According to one embodiment, the processor (201) can determine the display method (e.g., display order, or display method) of a plurality of sections. For example, within an interface for displaying summary information (560), the processor (201) may place a section related to the main content at the front and a section related to the remaining content, which is distinct from the main content, at the back of the interface.
[0168] According to an embodiment, if the second text content (532) simply represents voice information of the video content (520), the processor (201) can obtain (or generate) summary information (560) using only the second text content (532).
[0169] FIG. 8 illustrates a flowchart of an exemplary operation of an electronic device for displaying summary information. In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0170] In operation 801, the processor (201) may display a first user interface including text content and video content. The processor (201) may identify that the first user interface includes multimedia content.
[0171] According to one embodiment, the processor (201) can identify text content and video content within the first user interface. For example, the processor (201) can determine what type of content (e.g., text content, video content) is included in the first user interface. For example, the text content may correspond to the second text content (532) of FIG. 5. For example, the video content may correspond to the video content (520) of FIG. 5.
[0172] In operation 802, the processor (201) may receive user input for displaying summary information (560). For example, the processor (201) may receive user input for displaying summary information (560) in relation to the first user interface.
[0173] For example, user input for displaying summary information (560) may include at least one of input to a physical button of the electronic device (200), gesture input, and input to an object for requesting a summary.
[0174] In operation 803, the processor (201) can obtain subtitle information (or subtitle data) from video content. For example, the processor (201) can obtain subtitle information (or subtitle data) of video content through an API (or OpenAPI) provided by a service provider. For example, the processor (201) can obtain subtitle information (or subtitle data) of video content using WebVTT information. For example, the processor (201) can obtain audio data from video content. The processor (201) can identify voice information within the audio data based on STT techniques. Based on the voice information, the processor (201) can obtain subtitle information (or subtitle data) of video content.
[0175] For example, subtitle information (or subtitle data) may include at least one of a plurality of start times for displaying subtitles within video content, a plurality of display durations corresponding to the plurality of start times, or a plurality of characters (e.g., character blocks) corresponding to the plurality of display durations.
[0176] In operation 804, the processor (201) can generate summary information (560) based on subtitle information (or subtitle data) regarding text content and video content. For example, the processor (201) can generate summary information (560) based on subtitle information regarding video content. For example, the processor (201) can input at least one of text content or subtitle information (or subtitle data) into an artificial intelligence model (550). The processor (201) can obtain summary information (560) based on the output of the artificial intelligence model (550).
[0177] In operation 805, the processor (201) may display a second user interface containing summary information (560). For example, the processor (201) may display a second user interface containing summary information (560) through a display (202).
[0178] For example, the processor (201) can change the user interface displayed on the screen of the display (202) from the first user interface to the second user interface. For example, the processor (201) can switch the first user interface to the second user interface.
[0179] According to one embodiment, the processor (201) may display a second user interface together with a first user interface. For example, the second user interface may be displayed superimposed on the first user interface. For example, the processor (201) may display the first user interface and the second user interface on a display (202) via a split screen.
[0180] According to operations 801 to 805, the processor (201) can obtain summary information (560) based on subtitle information (or subtitle data) of text content and video content within the first user interface and display the summary information (560) through the second user interface.
[0181] FIG. 9 illustrates an example of the operation of an electronic device for displaying summary information.
[0182] Referring to FIG. 9, the processor (201) can display a screen (910). The processor (201) can display the screen (910) through a display (202). The processor (201) can display a user interface (911) through the screen (910). For example, the user interface (911) may be used to display at least a portion of a web page, but is not limited thereto.
[0183] For example, the user interface (911) may include video content (912) and text content (913). For example, the text content (913) may include a description of the video content (912).
[0184] According to one embodiment, the user interface (911) may include a control area (914) that includes objects for controlling an application regarding the user interface (911). For example, the control area (914) may include an object (915) for performing a function according to an artificial intelligence model (550). Based on input for the object (915), the processor (201) may change the screen displayed through the display (202) of the electronic device (200) from the screen (910) to the screen (920).
[0185] On the screen (920), the processor (201) may display an object (921) for representing functions according to the artificial intelligence model (550) based on input to the object (915). For example, the object (921) may include an object (922) for performing a summary function and an object (923) for performing a translation function. For example, the object (915) may cause a display of the object (922) for performing a summary function. The processor (201) may display the object (922) based on input to the object (915). For example, the processor (201) may provide an option to specify the type of language as a submenu of the object (922) for performing a summary function, and may receive a request to display summary information in the specified language.
[0186] According to one embodiment, the processor (201) can change the screen displayed through the display (202) of the electronic device (200) from screen (920) to screen (930) based on input for an object (922).
[0187] On the screen (930), the processor (201) may display a user interface (931) for providing summary information. For example, the processor (201) may display the user interface (931) overlaid on a user interface (911) that includes video content (912) and text content (913). The processor (201) may display the user interface (931) over the user interface (911). For example, the user interface (931) may be displayed to provide summary information (560).
[0188] According to one embodiment, the processor (201) can identify content within the user interface (911) based on input to the object (922). The processor (201) can identify video content (912) and text content (913) based on input to the object (922).
[0189] For example, the processor (201) may obtain at least one image regarding the video content (912) and text information regarding the video content (912) based on the video content (912) and text content (913). For example, the text information regarding the video content (912) may include subtitle data of the video content (520) and the text content and text content (913). For example, at least one image may correspond to at least one frame among a plurality of frames of the video content (912).
[0190] For example, the processor (201) can input at least one image of the video content (912) and text information of the video content (912) into the artificial intelligence model (550). The processor (201) can obtain summary information (560) based on the output of the artificial intelligence model (550).
[0191] For example, the processor (201) may provide summary information (560) through the user interface (931). For example, the processor (201) may display the summary information (560) within the user interface (931). For example, the summary information (560) may include a title (936) and summary content (937). The title (936) may correspond to the title of the video content (912).
[0192] According to one embodiment, the user interface (931) may include an object (932) for providing information about a summary function currently being performed, an object (933) for setting a summary method, an object (934) for displaying a previous page among the pages of summary information (560), and an object (935) for displaying a next page among the pages of summary information (560). If the summary information (560) is not all displayed within the user interface (931), the processor (201) may configure multiple pages of summary information (560). The processor (201) may display an object (934) and an object (935) for moving to multiple pages of summary information (560).
[0193] According to an embodiment, the processor (201) may identify in advance the contents contained within the user interface (911) in response to an input regarding an object (915) displayed on the screen (920). According to an embodiment, the processor (201) may perform at least one of acquiring audio data according to video content (912) and / or STT in response to an input regarding an object (915) displayed on the screen (920). The processor (201) may perform at least some of the operations for acquiring summary information (560) in the background in advance in response to an input regarding the object (915).
[0194] FIG. 10 illustrates an example of the operation of an electronic device for determining a video among a plurality of videos to provide summary information.
[0195] Referring to FIG. 10, a web page (1000) may include a plurality of videos. For example, the web page (1000) may include video content (1001), video content (1002), and video content (1003).
[0196] According to one embodiment, the processor (201) can determine video content for providing summary information (560) when a web page (1000) includes a plurality of video contents. To determine video content for providing summary information (560), the processor (201) can identify a plurality of video contents included in the web page (1000).
[0197] For example, the processor (201) can identify markup language information (1010) (e.g., HTML (hypertext mark-up language)) regarding a web page (1000). Based on the markup language information (1010) regarding the web page (1000), the processor (201) can identify multiple video contents. The processor (201) includes tags (e.g., that indicate video contents within the markup language information (1010) <video>Based on identifying whether ) is included, it can identify whether video content is included in the web page (1000). The processor (201) can identify attribute information (1011), attribute information (1012), and attribute information (1013) containing tags indicating video content within the markup language information (1010). Based on identifying attribute information (1011), attribute information (1012), and attribute information (1013), the processor (201) can identify that three video contents are included in the web page (1000). For example, the processor (201) can identify that video content (1001), video content (1002), and video content (1003) are included in the web page (1000).
[0198] According to one embodiment, the processor (201) may determine the video content among the plurality of video contents included in the web page (1000) to provide summary information (560). For example, the processor (201) may display all of the plurality of video contents within the screen of a display (202) corresponding to at least a part of the web page (1000). The processor (201) may identify a plurality of areas for displaying the plurality of video contents. For example, the plurality of areas may include the area occupied by each of the plurality of video contents within the web page (1000).
[0199] For example, the processor (201) can identify the size of a first area where video content (1001) is displayed based on attribute information (1011). For example, based on attribute information (1011), the width and height of the first area where video content (1001) is displayed can be identified. For example, the processor (201) can identify the size of a second area where video content (1002) is displayed based on attribute information (1012). For example, based on attribute information (1012), the width and height of the second area where video content (1002) is displayed can be identified. For example, the processor (201) can identify the size of a third area where video content (1003) is displayed based on attribute information (1013). For example, based on attribute information (1013), the width and height of the third area where the video content (1003) is displayed can be identified.
[0200] According to one embodiment, the processor (201) may determine a video content for acquiring at least one image among a plurality of video contents based on the size of each of a plurality of regions. For example, the processor (201) may determine a video content (1001) corresponding to the region having the largest size among the first to third regions as a video content for providing summary information (560). For example, the processor (201) may arrange a plurality of video contents in order of the size of the regions where the video content is displayed (or occupied) and determine a video content (1001) displayed in the largest region as a video content for providing summary information (560).
[0201] According to one embodiment, when the size of the area where each of the plurality of video contents is displayed (or occupied) is the same, the processor (201) may determine the currently focused video content as the video content for providing summary information (560). For example, the currently focused video content may be identified via a mouse, a stylus, or touch. For example, the video content where the mouse cursor is located may be identified as the currently focused video content. For example, the video content where a touch input occurs may be identified as the currently focused video content. For example, the video content located in the direction pointed by the stylus may be identified as the currently focused video content.
[0202] According to one embodiment, when the size of the area where each of the plurality of video contents is displayed (or occupied) is the same, the processor (201) may determine the video content that has a playback history as the video content for providing summary information (560).
[0203] According to one embodiment, the processor (201) can determine a video content to provide summary information (560) among a plurality of video contents according to a first criterion to a fourth criterion.
[0204] For example, the first criterion may be the size of the area occupied by the video content. Based on the first criterion, the processor (201) may determine the video content having the largest area in which each of the multiple video contents is displayed (or occupied) as the video content for providing summary information (560).
[0205] For example, the second criterion may be whether or not it is focused. If the processor (201) has not determined video content to provide summary information (560) according to the first criterion, it may determine video content to provide summary information (560) according to the second criterion. According to the second criterion, the processor (201) may determine the currently focused video content as video content to provide summary information (560).
[0206] For example, the third criterion may be a playback history. If the processor (201) has not determined video content to provide summary information (560) according to the first criterion and the second criterion, it may determine video content to provide summary information (560) according to the third criterion. According to the third criterion, the processor (201) may determine video content with a playback history as video content to provide summary information (560).
[0207] For example, the fourth criterion may be the size of the area overlapping with the designated area set based on the center of the screen. If the processor (201) has not determined the video content for providing summary information (560) according to the first through third criteria, it may determine the video content for providing summary information (560) according to the fourth criterion. According to the fourth criterion, the processor (201) may determine the video content for providing summary information (560) that has the largest size of the area overlapping with the designated area set based on the center of the screen. Specific examples of the fourth criterion will be described later in FIG. 11.
[0208] The aforementioned first to fourth criteria are exemplary and may be changed, and the order of judgment regarding whether the first to fourth criteria are satisfied may be changed.
[0209] FIG. 11 illustrates an example of the operation of an electronic device for determining a video among a plurality of videos to provide summary information.
[0210] Referring to FIG. 11, the web page (1000) may correspond to the web page (1000) illustrated in FIG. 10. The operations described in FIG. 11 may correspond to the fourth criterion described in FIG. 10.
[0211] A web page (1000) may include multiple video contents. For example, a web page (1000) may include video content (1001), video content (1002), and video content (1003).
[0212] Video content (1001), video content (1002), and video content (1003) included in the web page (1000) can all be displayed through the display (202). The processor (201) can set an area (1110) based on the center of the screen of the display (202). The processor (201) can identify the video content (1001) with the largest area overlapping with the area (1110). The processor (201) can determine the video content (1001) as the video content to provide summary information (560).
[0213] According to an embodiment, the processor (201) may arrange a plurality of video contents in order of the size of the regions that overlap with the region (1110), and determine the video content (1001) having the largest overlap region as the video content for providing summary information (560).
[0214] FIG. 12 illustrates an example of the operation of an electronic device for determining a video among a plurality of videos to provide summary information.
[0215] Referring to FIG. 12, the processor (201) may not be able to display the entire web page (1200) through the display (202). The processor (201) may display at least a portion of the web page (1200). For example, the web page (1000) may include an area (1210) where text content is placed and an area (1220) where video content is placed. The processor (201) may obtain summary information (560) based on the area (1210) and area (1220) of the web page (1200) that is displayed within the screen displayed through the display (202).
[0216] According to one embodiment, the processor (201) can identify video content placed in area (1220) even when the area (1220) where video content is displayed is not displayed through a screen displayed via a display (202). The processor (201) can obtain summary information (560) based on text content (1210) placed in area (1210), video content placed in area (1220), and area (1220).
[0217] According to one embodiment, the processor (201) can identify text content placed in an area (1210) even when the area (1210) where text content is displayed is not displayed through a screen displayed via a display (202). The processor (201) can obtain summary information (560) based on the text content (1210) placed in the area (1210), video content placed in the area (1220), and the area (1220).
[0218] According to the above-described embodiment, the processor (201) can identify the video content and text content included in the web page (1200) even if one of the video content and text content included in the web page (1200) is not displayed through the screen of the display (202). The processor (201) can provide summary information based on the video content and text content included in the web page (1200).
[0219] FIG. 13 illustrates an example of the operation of an electronic device according to the type of playback section of video content.
[0220] Referring to FIG. 13, the processor (201) can identify the type of playback section of the video content (1300). For example, the processor (201) can identify a playback section in which a scene depicting an event and dialogue between characters are displayed as a first type. The processor (201) can identify a playback section in which voice for describing a designated object (e.g., place, object, animal) is output as a second type. The processor (201) can identify a playback section in which only background music without voice (or lyrics) is output and the designated object is not displayed as a third type. However, it is not limited thereto. Depending on the embodiment, the type of playback section of the video content (1300) can be set in various ways. For example, a playback section in which a person is displayed can be identified as a first type, and a playback section in which an animal (or object) is displayed can be identified as a second type.
[0221] According to an embodiment, the processor (201) can identify the type of a playback section of video content (1300) based on various methods. For example, the processor (201) can identify at least one playback section and identify the type for each of the at least one playback section based on inputting the video content (1300) into an artificial intelligence model. For example, the processor (201) can identify the type of a playback section of video content (1300) based on at least one of metadata regarding the video content (1300), text content included in a web page, or information obtained (or retrieved) based on information from a web page.
[0222] For example, the processor (201) may divide the entire playback section of the video content (1300) into a first type of playback section (1301), a second type of playback section (1302), and a third type of playback section. The processor (201) may obtain summary information for each of the first type of playback section (1301), the second type of playback section (1302), and the third type of playback section (1303). For example, the summary information (560) may include summary information for the first type of playback section (1301), summary information for the second type of playback section (1302), and summary information for the third type of playback section (1303).
[0223] For example, the processor (201) can obtain first summary information for the first type of playback section (1301) based on inputting at least one image and text information for the first type of playback section (1301) into the artificial intelligence model (550).
[0224] For example, the processor (201) can obtain second summary information for the second type of playback section (1302) based on inputting at least one image and text information for the second type of playback section (1302) into the artificial intelligence model (550).
[0225] For example, the processor (201) can obtain third summary information for the third type of playback section (1303) based on inputting at least one image and text information for the third type of playback section (1303) into the artificial intelligence model (550).
[0226] Although not illustrated, according to the embodiment, the artificial intelligence models for obtaining summary information for a first type of playback section (1301), summary information for a second type of playback section (1302), and summary information for a third type of playback section (1303) may be configured differently. For example, the artificial intelligence model (550) may include a first artificial intelligence model for a first type of playback section (1301), a second artificial intelligence model for a second type of playback section (1302), and a third artificial intelligence model for a third type of playback section (1303). The processor (201) may obtain first summary information using the first artificial intelligence model. The processor (201) may obtain second summary information using the second artificial intelligence model. The processor (201) may obtain third summary information using the third artificial intelligence model.
[0227] According to one embodiment, the processor (201) may set the data for obtaining summary information differently depending on the type of playback section. For example, the processor (201) may obtain the first summary information and the second summary information based on voice information (or voice data). The processor (201) may obtain the third summary information based on a main frame (or at least one frame) within the playback section (1303) of the third type.
[0228] The first summary information and the second summary information may include a description of the event that occurred. The third summary information may include a description of the main frame. The processor (201) may obtain summary information (560) by combining the first summary information, the second summary information, and the third summary information. The processor (201) may also provide information indicating the type of playback section together with the first summary information, the second summary information, and / or the third summary information.
[0229] According to an embodiment, the processor (201) may distinguish the type of video content (1300). For example, the processor (201) may identify the type of video content (1300) as a first type based on identifying that the playback section (1301) of a first type is greater than or equal to a specified ratio. For example, the processor (201) may identify the type of video content (1300) as a second type based on identifying that the playback section (1302) of a second type is greater than or equal to a specified ratio. For example, the processor (201) may identify the type of video content (1300) as a third type based on identifying that the playback section (1303) of a third type is greater than or equal to a specified ratio. For example, the processor (201) may use subtitle data and audio data to obtain summary information (560) based on identifying the type of video content (1300) as a first type or a second type. For example, the processor (201) may use a key frame of the video content (1300) to obtain summary information (560) based on identifying the type of video content (1300) as a third type.
[0230] FIG. 14a illustrates an example of the operation of an electronic device according to the type of playback section of video content.
[0231] Referring to FIG. 14a, the web page (1400) may include video content (1401) to be summarized and text content (1402) to be summarized. The web page (1400) may include not only video content (1401) and text content (1402), but also advertising content.
[0232] According to one embodiment, the processor (201) can identify advertising content (1411), advertising content (1412), and advertising content (1413) based on configuration information and / or tag information of the web page (1400). The processor (201) can exclude advertising content (1411), advertising content (1412), and advertising content (1413) from being summarized.
[0233] According to one embodiment, some frames of the video content (1401) may contain advertisements. The processor (201) may exclude some frames containing advertisements from the summary target. According to one embodiment, when the video content (1401) is played, the advertisement video may be played first. The processor (201) may exclude the advertisement video that is played before the video content (1401) from the summary target.
[0234] For example, the processor (201) can identify whether an advertisement video is included in the video content (1401) based on data that has been pre-trained through other video content related to the video content (1401), and / or the type of the user's subscription account (e.g., free or paid).
[0235] For example, the processor (201) can identify whether an advertisement video is included in the video content (1401) by playing the video content (1401) in the background. For example, the processor (201) can identify whether an advertisement video is included in the video content (1401) based on identifying whether a video of a product description distinct from the subject of the video content (1401) is played. According to an embodiment, the processor (201) can identify that an advertisement video is included in the video content (1401) based on identifying that a designated identifier for an advertisement is displayed upon playing the video content (1401). For example, a designated identifier for an advertisement may include “skip”, “AD skip”, “skip after 5 seconds”, “sponsor”, and / or “play after 10 seconds”. According to an embodiment, when the processor (201) pre-plays the video content (1401) in the background, it may play the video content (1401) at a playback speed faster than the default playback speed.
[0236] FIG. 14b illustrates an example of the operation of an electronic device according to the state of the electronic device.
[0237] Referring to FIG. 14b, the electronic device (200) may be foldable. For example, the electronic device (200) may include a first housing (1451), a second housing (1452), and a third housing (1453). Depending on the positional relationship of the first housing (1451), the second housing (1452), and the third housing (1453), the electronic device (200) may have three states. For example, the states of the electronic device (200) may include state (1431), state (1432), and state (1433). State (1431) may be referred to as a fully folded state. State (1432) may be referred to as a partially folded state. State (1433) may be referred to as a fully unfolded state.
[0238] For example, the electronic device (200) may include a display (1442) disposed on a first surface of the first housing (1451), a first surface of the second housing (1452), and a first surface of the third housing (1453). The electronic device (200) may include a display (1441) disposed on a second surface of the third housing (1453). The second surface may be opposite to the first surface. For example, the display (1441) disposed on the second surface of the third housing (1453) may be referred to as a cover display. The display (1442) may be referred to as a main display. According to an embodiment, the display (1441) may be disposed on a second surface of the second housing (1452). As the display (1441) is disposed on the second surface of the second housing (1452), the folding state of the electronic device (200) may be changed.
[0239] For example, the display (1442) may be configured as a flexible display. The display (1442) may include a first display area (1461) corresponding to a first surface of the first housing (1451), a second display area (1462) corresponding to a first surface of the second housing (1452), and a third display area (1463) corresponding to a first surface of the third housing (1453).
[0240] In state (1431), the processor (201) may display a user interface (1471) for displaying summary information on a display (1442). While the user interface (1471) is displayed, the processor (201) may identify that the state of the electronic device (200) changes from state (1431) to state (1432). For example, the processor (201) may identify that the state of the electronic device (200) changes from state (1431) to state (1432) based on identifying that the third housing (1453) of the electronic device (200) is rotated relative to the first housing (1451) and the second housing (1452).
[0241] In state (1432), the processor (201) may display a user interface (1472) for displaying summary information in a third display area (1463) of the display (1442). The user interface (1472) may correspond to the user interface (1471) displayed in state (1431). While the user interface (1472) is displayed, the processor (201) may identify that the state of the electronic device (200) changes from state (1432) to state (1433). For example, the processor (201) may identify that the state of the electronic device (200) changes from state (1432) to state (1433) based on identifying that the first housing (1451) of the electronic device (200) is rotated relative to the second housing (1452).
[0242] In state (1433), the processor (201) may display a user interface (1473) for displaying video content (or a web page) in the first display area (1461) and the second display area (1462) of the display (1442). The processor (201) may display a user interface (1474) for displaying summary information in the third display area (1463) of the display (1442).
[0243] For example, an object (1482) representing a navigation bar may be displayed in an area (1481) for displaying video content displayed in the user interface (1473). The processor (201) may display a representative image (or main image) within a time interval according to playback timing that changes according to input for the object (1482) in an area (1483) within the user interface (1474).
[0244] According to an embodiment, the processor (201) can change the image displayed in the area (1483) based on an input to the area (1483) (e.g., a swipe input). The processor (201) can identify a time interval corresponding to the changed image and play video content according to the identified time interval.
[0245] According to the above-described embodiment, the processor (201) can synchronize and display an image displayed as the playback timing and summary information of the video content.
[0246] FIG. 15 illustrates a flowchart of an exemplary operation of an electronic device for displaying summary information. In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.
[0247] Referring to FIG. 15, in operation 1510, the processor (201) can display at least a portion of the web page through the display (202). For example, the processor (201) can display at least a portion of the web page through a user interface.
[0248] In operation 1520, the processor (201) can identify whether text content and video content are included within the main body of a web page. For example, the processor (201) can identify whether text content and video content are included within the main body of a web page by using the markup language file of the web page.
[0249] According to one embodiment, the processor (201) can identify that the web page contains video content by using the markup language file of the web page. The processor (201) can identify video content within at least a portion of the display of the wearable device (410) among the video content included in the web page.
[0250] In operation 1530, if text content and video content are included within the main body of a web page, the processor (201) can obtain subtitle information for the video content. For example, the processor (201) can identify that text content and video content are included within the main body of a web page. Based on identifying that text content and video content are included within the main body of a web page, the processor (201) can obtain subtitle information for the video content.
[0251] According to one embodiment, the processor (201) can obtain identification information of video content by using a markup language file of a web page. For example, the processor (201) can receive subtitle information for video content from an external electronic device for providing video content through an application programming interface (API) that includes identification information. For example, the processor (201) can request subtitle information for video content from an external electronic device for providing video content through an API that includes identification information. Based on the request, the processor (201) can receive subtitle information for video content from an external electronic device for providing video content.
[0252] In operation 1540, the processor (201) can obtain first summary information corresponding to video content by using subtitle information.
[0253] According to one embodiment, when a trained model (e.g., an artificial intelligence model (550)) for obtaining summary information is included in memory (203), the processor (201) can obtain (or generate) first summary information based on subtitle information by using the trained model included in memory (203).
[0254] According to one embodiment, if a trained model for obtaining summary information is included in a server, the processor (201) may transmit subtitle information to the server including the trained model. Based on transmitting subtitle information to the server, the processor (201) may request first summary information corresponding to video content from the server. The processor (201) may receive the first summary information from the server.
[0255] According to an embodiment, the processor (201) can identify from subtitle information at least one playback point associated with at least one text block configured to be displayed in the video content. Before receiving input for playback of the video content, the processor (201) can acquire at least one image of the video content based on at least one playback point. The processor (201) can acquire first summary information based on the subtitle information and at least one image.
[0256] In operation 1550, the processor (201) can obtain second summary information corresponding to the text content.
[0257] According to one embodiment, when a trained model (e.g., an artificial intelligence model (550)) for obtaining summary information is included in memory (203), the processor (201) can obtain (or generate) second summary information based on text content by using the trained model included in memory (203).
[0258] According to one embodiment, if a trained model for obtaining summary information is included in a server, the processor (201) may transmit text content to the server including the trained model. Based on transmitting the text content to the server, the processor (201) may request second summary information corresponding to the text content from the server. The processor (201) may receive the second summary information from the server.
[0259] According to one embodiment, the processor (201) can obtain multimedia content included in a web page. The multimedia content may not include video content. The multimedia content may include text content. The processor (201) can remove advertising content from the multimedia content. Based on the multimedia content from which advertising content has been removed, the processor (201) can obtain second summary information.
[0260] In operation 1560, the processor (201) may provide first summary information and second summary information. For example, the processor (201) may provide the first summary information and second summary information through a display (202). For example, the processor (201) may display a first visual object for representing the first summary information and a second visual object for representing the second summary information in a superimposed manner on at least a portion of the display of a web page. The first visual object may be displayed in a superimposed manner on video content. The second visual object may be displayed in a superimposed manner on text content. According to an embodiment, the processor (201) may also provide the first summary information and second summary information through voice signals.
[0261] According to one embodiment, the processor (201) may obtain a third summary information based on the first summary information and the second summary information. The processor (201) may provide the third summary information. For example, the processor (201) may identify the similarity between the first summary information and the second summary information. The processor (201) may identify that the similarity between the first summary information and the second summary information exceeds a reference similarity. The processor (201) may obtain the third summary information based on identifying that the identified similarity exceeds the reference similarity. The processor (201) may provide the obtained third summary information.
[0262] According to an embodiment, the processor (201) may request subtitle information for the video content from an external electronic device for providing video content. Based on the request, the processor (201) may receive information from the external electronic device indicating that subtitle information for the video content is not provided. If subtitle information for the video content is not provided, the processor (201) may obtain identification information for the video content by using a markup language file of a web page. To obtain the first summary information, the processor (201) may determine a prompt for a trained model based on the identification information. Based on the prompt, the processor (201) may obtain the first summary information using the trained model.
[0263] According to an embodiment, if subtitle information for video content is not provided, the processor (201) can obtain audio data of the video content based on the video content. The processor (201) can identify voice data included in the audio data. The processor (201) can obtain first summary information using text information identified based on the voice data.
[0264] In operation 1570, the processor (201) may obtain second summary information corresponding to text content if video content is not included within the main body of the web page. Operation 1570 may correspond to operation 1550.
[0265] In operation 1580, the processor (201) may provide second summary information. If only text content is included among video content and text content within the main body of a web page, the processor (201) may provide only second summary information corresponding to the text content.
[0266] FIG. 16a illustrates an example of the operation of an electronic device for providing summary information.
[0267] Referring to FIG. 16a, in state (1610), the processor (201) can display a screen (1611). The screen (1611) may correspond to a display of at least a portion of a web page.
[0268] According to one embodiment, the processor (201) can identify that video content (1612) and text content (1613) are included within a web page. The processor (201) can identify that video content (1612) and text content (1613) are included within a web page by using the markup language file of the web page.
[0269] According to one embodiment, the processor (201) can obtain first summary information corresponding to video content (1612). For example, the processor (201) can obtain identification information of the video content (1612) by using a markup language file of a web page. For example, the processor (201) can request subtitle information for the video content (1612) from an external electronic device for providing the video content (1612) through an API containing identification information. Based on the request, the processor (201) can receive subtitle information for the video content (1612) from the external electronic device for providing the video content (1612). The processor (201) can obtain first summary information based on the subtitle information. In one example, the processor (201) can obtain (or generate) first summary information based on the subtitle information by using a trained model (e.g., an artificial intelligence model (550)).
[0270] According to one embodiment, the processor (201) can obtain second summary information corresponding to text content. For example, the processor (201) can obtain (or generate) second summary information based on text content by using a trained model (e.g., an artificial intelligence model (550)).
[0271] According to one embodiment, the processor (201) can identify the similarity between the first summary information and the second summary information. The processor (201) can identify that the similarity between the first summary information and the second summary information exceeds a reference similarity. Based on identifying that the similarity between the first summary information and the second summary information exceeds a reference similarity, the processor (201) can obtain a third summary information. Based on obtaining the third summary information, the processor (201) can change the state of the electronic device (200) from state (1610) to state (1620).
[0272] In state (1620), the processor (201) may display a visual object (1621) representing third summary information within the screen (1611). For example, the visual object (1621) may be displayed overlaid on at least a portion of the display of the web page.
[0273] In FIG. 16a, an example is described in which first summary information and second summary information are obtained, and third summary information obtained based on the first summary information and second summary information is displayed on the screen (1611), but is not limited thereto. According to an embodiment, the processor (201) may display a visual object representing the first summary information and a visual object representing the second summary information on the screen (1611). An embodiment in which a visual object representing the first summary information and a visual object representing the second summary information are displayed will be described in FIG. 16b.
[0274] FIG. 16b illustrates an example of the operation of an electronic device for providing summary information.
[0275] Referring to FIG. 16a, in state (1630), the processor (201) can display a screen (1631). The screen (1631) may correspond to a display of at least a portion of a web page.
[0276] According to one embodiment, the processor (201) can identify that video content (1632) and text content (1633) are included within a web page. The processor (201) can identify that video content (1632) and text content (1633) are included within a web page by using the markup language file of the web page.
[0277] According to one embodiment, the processor (201) can obtain first summary information corresponding to video content (1632). For example, the processor (201) can obtain identification information of the video content (1632) by using a markup language file of a web page. For example, the processor (201) can request subtitle information for the video content (1632) from an external electronic device for providing the video content (1632) through an API containing identification information. Based on the request, the processor (201) can receive subtitle information for the video content (1632) from the external electronic device for providing the video content (1632). The processor (201) can obtain first summary information based on the subtitle information. In one example, the processor (201) can obtain (or generate) first summary information based on the subtitle information by using a trained model (e.g., an artificial intelligence model (550)).
[0278] According to one embodiment, the processor (201) can obtain second summary information corresponding to text content. For example, the processor (201) can obtain (or generate) second summary information based on text content by using a trained model (e.g., an artificial intelligence model (550)).
[0279] In state (1640), the processor (201) may display a visual object (1641) representing first summary information and a visual object (1642) representing second summary information on the screen (1631). For example, the processor (201) may display the visual object (1641) and the visual object (1642) superimposed on at least a portion of the display of the web page.
[0280] According to one embodiment, the processor (201) may obtain a third summary information based on the first summary information and the second summary information. The processor (201) may provide the third summary information. For example, the processor (201) may identify the similarity between the first summary information and the second summary information. The processor (201) may identify that the similarity between the first summary information and the second summary information is less than or equal to a reference similarity. Based on identifying that the identified similarity is less than or equal to a reference similarity, the processor (201) may bypass obtaining the third summary information. The processor (201) may provide both the first summary information and the second summary information. The processor (201) may provide both the first summary information and the second summary information through the visual object (1641) and the visual object (1642).
[0281] FIG. 17 illustrates an example of the operation of an electronic device for providing summary information.
[0282] Referring to FIG. 17, in state (1710), the processor (201) can identify an input (1750) for changing the display of at least a portion of the web page while at least a portion of the web page is displayed. For example, the input (1750) may correspond to an input for scrolling the web page (or screen). The processor (201) can change the display of at least a portion of the web page based on the input (1750).
[0283] According to one embodiment, the processor (201) can display a screen (1711) corresponding to a changed display through the display (202) based on a change in the display of at least a portion of a web page.
[0284] The screen (1711) (or altered display) may include video content (1712). The processor (201) can identify that video content (1712) is included within the screen (1711) (or altered display) by using a markup language file of a web page. The processor (201) can identify that the screen (1711) (or altered display) is maintained for a reference time (e.g., 3 seconds). Based on identifying that the screen (1711) (or altered display) is maintained for a reference time (e.g., 3 seconds), the processor (201) can change the state of the electronic device (200) from state (1710) to state (1720).
[0285] In state (1720), the processor (201) may display a visual object (1721) for requesting summary information about video content (1712). For example, the processor (201) may display a visual object (1721) for requesting summary information about video content (1712) on the screen (1711). The processor (201) may display the visual object (1721) for requesting summary information about video content (1712) superimposed on the changed display.
[0286] According to one embodiment, the processor (201) can identify input for a visual object (1721). Based on the input for the visual object (1721), the processor (201) can provide summary information for video content (1712).
[0287] According to an embodiment, the screen (1711) may include not only video content (1712) but also text content (1713) regarding the video content (1712). The processor (201) may provide summary information based on the video content (1712) and the text content (1713). For example, the processor (201) may provide first summary information corresponding to the video content (1712) and second summary information corresponding to the text content (1713). For example, the processor (201) may obtain third summary information based on the first summary information corresponding to the video content (1712) and the second summary information corresponding to the text content (1713). The processor (201) may provide third summary information.
[0288] FIG. 18 is a block diagram showing an integrated intelligence system according to one embodiment.
[0289] Referring to FIG. 18, an integrated intelligent system (10) of one embodiment may include a user terminal (1800), an intelligent server (1900), and a service server (2000).
[0290] A user terminal (1800) of one embodiment (e.g., the electronic device (101) of FIG. 1) may be a terminal device (or electronic device) capable of connecting to the Internet, and may be, for example, a mobile phone, a smartphone, a PDA (personal digital assistant), a laptop computer, a TV, a home appliance, a wearable device, an HMD, or a smart speaker.
[0291] According to one embodiment, the user terminal (1800) may include a communication interface (1810), a microphone (1820), a speaker (1830), a display (1840), a memory (1850), and a processor (1860). The listed components may be operatively or electrically connected to each other.
[0292] According to one embodiment, the communication interface (1810) may be configured to be connected to an external device to transmit and receive data. According to one embodiment, the microphone (1820) may receive sound (e.g., user speech) and convert it into an electrical signal. According to one embodiment, the speaker (1830) may output the electrical signal as sound (e.g., voice). According to one embodiment, the display (1840) may be configured to display an image or video. According to one embodiment, the display (1840) may display a graphic user interface (GUI) of an app (or application program) being executed.
[0293] A display (1840) of one embodiment may be configured to display an image or video. A display (1840) of one embodiment may display a graphic user interface (GUI) of an app (or application program) being executed. A display (1840) of one embodiment may receive touch input through a touch sensor. For example, the display (1840) may receive text input through a touch sensor in an on-screen keyboard area displayed within the display (1840).
[0294] According to one embodiment, memory (1850) may store a client module (1851), a software development kit (SDK) (1853), and a plurality of apps (1855). The client module (1851) and the SDK (1853) may form a framework (or solution program) for performing general-purpose functions. Additionally, the client module (1851) or the SDK (1853) may form a framework for processing user input (e.g., voice input, text input, touch input).
[0295] According to one embodiment, the memory (1850) may be a program for performing a specified function, wherein the plurality of apps (1855) are programs. According to one embodiment, the plurality of apps (1855) may include a first app (1855_1) and a second app (1855_3). According to one embodiment, each of the plurality of apps (1855) may include a plurality of operations for performing a specified function. For example, the plurality of apps (1855) may include at least one of an alarm app, a message app, and a schedule app. According to one embodiment, the plurality of apps (1855) may be executed by a processor (1860) to sequentially execute at least some of the plurality of operations.
[0296] According to one embodiment, the processor (1860) can control the overall operation of the user terminal (1800). For example, the processor (1860) can perform specified operations by being electrically connected to a communication interface (1810), a microphone (1820), a speaker (1830), a display (1840), and a memory (1850).
[0297] According to one embodiment, the processor (1860) may also perform a specified function by executing a program stored in the memory (1850). For example, the processor (1860) may execute at least one of the client module (1851) or the SDK (1853) to perform the following operations for processing user input. The processor (1860) may, for example, control the operation of a plurality of apps (1855) through the SDK (1853). The following operations described as the operation of the client module (1851) or the SDK (1853) may be operations performed by the execution of the processor (1860).
[0298] According to one embodiment, the client module (1851) can receive user input. For example, the client module (1851) can generate a voice signal corresponding to a user utterance detected through the microphone (1820). Alternatively, the client module (1851) can receive touch input detected through the display (1840). Alternatively, the client module (1851) can receive text input detected through a keyboard or a virtual keyboard. In addition, various forms of user input can be received through an input module included in the user terminal (1800) or an input module connected to the user terminal (1800). The client module (1851) can transmit the received user input to an intelligent server (1900). According to one embodiment, the client module (1851) can transmit status information of the user terminal (1800) to the intelligent server (1900) along with the received user input. The above state information may be, for example, the execution state information of an app.
[0299] According to one embodiment, the client module (1851) can receive a result corresponding to the received user input. For example, the client module (1851) can receive a result corresponding to the user input from an intelligent server (1900). The client module (1851) can display the received result on a display (1840). Additionally, the client module (1851) can output the received result as audio through a speaker (1830).
[0300] According to one embodiment, the client module (1851) may receive a plan corresponding to the received user input. The client module (1851) may display the results of executing multiple actions of the app according to the plan on the display (1840). For example, the client module (1851) may sequentially display the results of executing multiple actions on the display and output audio through the speaker (1830). In another example, the user terminal (1800) may display only some of the results of executing multiple actions (e.g., the result of the last action) on the display and output audio through the speaker (1830).
[0301] According to one embodiment, a client module (1851) may receive a request from an intelligent server (1900) to obtain information necessary to produce a result corresponding to user input. The information necessary to produce the result may be, for example, status information of a user terminal (1800). According to one embodiment, the client module (1851) may transmit the necessary information to the intelligent server (1900) in response to the request.
[0302] According to one embodiment, the client module (1851) can transmit result information of executing a plurality of operations according to a plan to the intelligent server (1900). The intelligent server (1900) can confirm that the received user input has been correctly processed through the result information.
[0303] According to one embodiment, the client module (1851) may include a voice recognition module. According to one embodiment, the client module (1851) may recognize voice input that performs a limited function through the voice recognition module. For example, the client module (1851) may perform an intelligent app for processing voice input to perform an organic action through a specified input (e.g., Wake Up!).
[0304] According to one embodiment, an intelligent server (1900) can receive information related to user voice input from a user terminal (1800) via a communication network. According to one embodiment, the intelligent server (1900) can convert the data related to the received voice input into text data. According to one embodiment, the intelligent server (1900) can generate a plan for performing a task corresponding to the user voice input based on the text data.
[0305] According to one embodiment, a plan may be generated by an artificial intelligence (AI) system. The AI system may be a rule-based system or a neural network-based system (e.g., a feedforward neural network (FNN) or a recurrent neural network (RNN)). Alternatively, it may be a combination of the foregoing or an AI system different therefrom. According to one embodiment, the plan may be selected from a set of predefined plans or may be generated in real time in response to a user request. For example, the AI system may select at least one plan from a plurality of predefined plans.
[0306] According to one embodiment, the intelligent server (1900) may transmit the result calculated according to the generated plan to the user terminal (1800) or transmit the generated plan to the user terminal (1800). According to one embodiment, the user terminal (1800) may display the result calculated according to the plan on a display. According to one embodiment, the user terminal (1800) may display the result of executing an operation according to the plan on a display.
[0307] An intelligent server (1900) of one embodiment may include a front end (1910), a natural language platform (1920), a capsule database (1930), an execution engine (1940), an end user interface (1950), a management platform (1960), a big data platform (1970), and an analytic platform (1980).
[0308] According to one embodiment, the front end (1910) can receive user input received from the user terminal (1800). The front end (1910) can transmit a response corresponding to the user input.
[0309] According to one embodiment, the natural language platform (1920) may include an automatic speech recognition module (ASR module) (1921), a natural language understanding module (NLU module) (1923), a planner module (1925), a natural language generator module (NLG module) (1927), and a text to speech module (TTS module) (1929).
[0310] According to one embodiment, the automatic speech recognition module (1921) can convert voice input received from the user terminal (1800) into text data. According to one embodiment, the natural language understanding module (1923) can identify the user's intent using the text data of the voice input. For example, the natural language understanding module (1923) can identify the user's intent by performing a syntactic analysis or a semantic analysis on the user input in the form of text data. According to one embodiment, the natural language understanding module (1923) can identify the meaning of a word extracted from the user input using linguistic features (e.g., grammatical elements) of a morpheme or phrase, and determine the user's intent by matching the identified meaning of the word to the intent. The natural language understanding module (1923) can acquire intent information corresponding to the user's utterance. The intent information may be information indicating the user's intent determined by interpreting the text data. Intention information may include information indicating an action or function that the user intends to execute using the device.
[0311] According to one embodiment, the planner module (1925) can generate a plan using the intent and parameters determined by the natural language understanding module (1923). According to one embodiment, the planner module (1925) can determine multiple domains necessary to perform a task based on the determined intent. The planner module (1925) can determine multiple actions included in each of the multiple domains determined based on the intent. According to one embodiment, the planner module (1925) can determine parameters necessary to execute the determined multiple actions or result values output by the execution of the multiple actions. The parameters and result values may be defined as concepts associated with a specified format (or class). Accordingly, the plan may include multiple actions and multiple concepts determined by the user's intent. The planner module (1925) can determine the relationship between the multiple actions and the multiple concepts in a stepwise (or hierarchical) manner. For example, the planner module (1925) can determine the execution order of multiple actions determined based on the user's intentions based on multiple concepts. In other words, the planner module (1925) can determine the execution order of multiple actions based on parameters required for the execution of multiple actions and results output by the execution of multiple actions. Accordingly, the planner module (1925) can generate a plan that includes association information (e.g., ontology) between multiple actions and multiple concepts. The planner module (1925) can generate the plan using information stored in a capsule database (1930) in which a set of relationships between concepts and actions is stored.
[0312] According to one embodiment, the natural language generation module (1927) can change specified information into a text form. The information changed into a text form may be in the form of a natural language utterance. The text-to-speech conversion module (1929) of one embodiment can change information in a text form into information in a speech form.
[0313] According to one embodiment, the capsule database (1930) may store information regarding the relationships between multiple concepts and actions corresponding to multiple domains. For example, the capsule database (1930) may store multiple capsules including multiple action objects (or action information) and concept objects (or concept information) of a plan. According to one embodiment, the capsule database (1930) may store the multiple capsules in the form of a CAN (concept action network). According to one embodiment, the multiple capsules may be stored in a function registry included in the capsule database (1930).
[0314] According to one embodiment, the capsule database (1930) may include a strategy registry in which strategy information necessary for determining a plan corresponding to a voice input is stored. The strategy information may include reference information for determining one plan when there are multiple plans corresponding to the user input. According to one embodiment, the capsule database (1930) may include a follow-up registry in which information of a follow-up action is stored for suggesting a follow-up action to the user in a specified situation. The follow-up action may include, for example, a follow-up utterance. According to one embodiment, the capsule database (1930) may include a layout registry in which layout information of information output through the user terminal (1800) is stored. According to one embodiment, the capsule database (1930) may include a vocabulary registry in which vocabulary information included in the capsule information is stored. According to one embodiment, the capsule database (1930) may include a dialog registry in which information about a conversation (or interaction) with a user is stored.
[0315] According to one embodiment, the capsule database (1930) can update stored objects through a developer tool. The developer tool may include, for example, a function editor for updating action objects or concept objects. The developer tool may include a vocabulary editor for updating vocabulary. The developer tool may include a strategy editor for creating and registering strategies for determining plans. The developer tool may include a dialogue editor for creating conversations with the user. The developer tool may include a follow-up editor for editing follow-up utterances that activate follow-up goals and provide hints. The follow-up goals may be determined based on currently set goals, user preferences, or environmental conditions.
[0316] According to one embodiment, the capsule database (1930) may also be implemented within the user terminal (1800). In other words, the user terminal (1800) may include a capsule database (1930) that stores information for determining an action corresponding to voice input.
[0317] According to one embodiment, the execution engine (1940) can produce a result using the generated plan. According to one embodiment, the end user interface (1950) can transmit the produced result to the user terminal (1800). Accordingly, the user terminal (1800) can receive the result and provide the received result to the user. According to one embodiment, the management platform (1960) can manage information used in the intelligent server (1900). According to one embodiment, the big data platform (1970) can collect user data. According to one embodiment, the analysis platform (1980) can manage the quality of service (QoS) of the intelligent server (1900). For example, the analysis platform (1980) can manage the components and processing speed (or efficiency) of the intelligent server (1900).
[0318] According to one embodiment, the service server (2000) may provide a designated service (e.g., food ordering or hotel reservation) to the user terminal (1800). According to one embodiment, the service server (2000) may be a server operated by a third party. For example, the service server (2000) may include a first service server (2001), a second service server (2003), and a third service server (2005) operated by different third parties. According to one embodiment, the service server (2000) may provide information to the intelligent server (1900) for generating a plan corresponding to the received voice input. The provided information may be stored, for example, in a capsule database (1930). Additionally, the service server (2000) may provide result information according to the plan to the intelligent server (1900).
[0319] In the integrated intelligent system (10) described above, the user terminal (1800) can provide various intelligent services to the user in response to user input. The user input may include, for example, input via a physical button, touch input, or voice input.
[0320] According to one embodiment, the user terminal (1800) may provide a voice recognition service through an intelligent app (or voice recognition app) stored internally. In this case, for example, the user terminal (1800) may recognize a user utterance or voice input received through the microphone and provide a service to the user corresponding to the recognized voice input.
[0321] According to one embodiment, the user terminal (1800) may perform a specified action, either alone or together with the intelligent server and / or service server, based on the received voice input. For example, the user terminal (1800) may execute an app corresponding to the received voice input and perform a specified action through the executed app.
[0322] According to one embodiment, when a user terminal (1800) provides services together with an intelligent server (1900) and / or a service server, the user terminal can detect user speech using the microphone (1820) and generate a signal (or voice data) corresponding to the detected user speech. The user terminal can transmit the voice data to the intelligent server (1900) using a communication interface (1810).
[0323] According to one embodiment, an intelligent server (1900) may generate, in response to a voice input received from a user terminal (1800), a plan for performing a task corresponding to the voice input, or a result of performing an operation according to the plan. The plan may include, for example, a plurality of operations for performing a task corresponding to the user's voice input, and a plurality of concepts related to the plurality of operations. The concepts may define parameters input to the execution of the plurality of operations or result values output by the execution of the plurality of operations. The plan may include association information between the plurality of operations and the plurality of concepts.
[0324] A user terminal (1800) of one embodiment can receive the response using a communication interface (1810). The user terminal (1800) can output a voice signal generated inside the user terminal (1800) to the outside using the speaker (1830), or output an image generated inside the user terminal (1800) to the outside using the display (1840).
[0325] Figure 19 is a diagram showing the form in which relationship information between concepts and operations is stored in a database.
[0326] The capsule database (e.g., the capsule database (1930) of FIG. 18) of the intelligent server (e.g., the intelligent server (1900) of FIG. 18) can store multiple capsules in the form of a CAN (concept action network) (2200). The capsule database can store actions for processing tasks corresponding to user voice input, and parameters required for said actions, in the form of a CAN (concept action network). The CAN may represent an organic relationship between an action and a concept that defines the parameters required to perform said actions.
[0327] The above capsule database may store multiple capsules (e.g., Capsule A (2201), Capsule B (2204)) corresponding to each of a plurality of domains (e.g., applications). According to one embodiment, one capsule (e.g., Capsule A (2201)) may correspond to one domain (e.g., applications). Additionally, one capsule may correspond to at least one service provider (e.g., CP 1 (2202), CP 2 (2203), CP 3 (2206), or CP 4 (2205)) for performing functions of the domain associated with the capsule. According to one embodiment, one capsule may include at least one operation (2210) and at least one concept (2220) for performing a specified function.
[0328] According to one embodiment, a natural language platform (e.g., the natural language platform (1920) of FIG. 18) can generate a plan for performing a task corresponding to a received voice input using capsules stored in a capsule database. For example, a planner module of the natural language platform (e.g., the planner module (1925) of FIG. 18) can generate a plan using capsules stored in a capsule database. For example, a plan (2207) can be generated using the actions (2311, 2313) and concepts (2312, 2314) of Capsule A (2201) and the actions (2341) and concepts (2342) of Capsule B (2204).
[0329] FIG. 20 is a diagram showing a screen in which a user terminal processes voice input received through an intelligent app.
[0330] A user terminal (2000) (e.g., user terminal (1800) of FIG. 18) can run an intelligent app to process user input through an intelligent server (e.g., intelligent server (1900) of FIG. 18).
[0331] According to one embodiment, on a screen 2010, the user terminal (2000) may execute an intelligent app for processing voice input when it recognizes a specified voice input (e.g., "Summarize it!") or receives input via a hardware key (e.g., a dedicated hardware key). The user terminal (2000) may execute the intelligent app while, for example, a web page containing video content is displayed. According to one embodiment, the user terminal (2000) may display an object (e.g., an icon) (2011) corresponding to the intelligent app on a display (e.g., the display (1840) of FIG. 18). According to one embodiment, the user terminal (2000) may receive voice input by user speech. For example, the user terminal (2000) may receive voice input saying "Summarize it!". According to one embodiment, the user terminal (2000) can display a UI (user interface) (2013) (e.g., an input window) of an intelligent app on a display that displays text data of the received voice input.
[0332] According to one embodiment, on a screen of 2020, the user terminal (2000) can display a result corresponding to the received voice input on the display. For example, the user terminal (2000) can receive a plan corresponding to the received user input and display summary information obtained according to the plan on the display.
[0333] Some of the operations described above may be executed (or performed) through an artificial intelligence (AI) system described with reference to FIG. 21.
[0334] Figure 21 is a schematic diagram of an exemplary AI system.
[0335] Referring to FIG. 21, the AI system (2100) may include an input / output interface (2110), an AI framework (2120), a generative AI model (2130), and / or a knowledge repository (2190).
[0336] The input / output interface (2110) can receive input. The input may include user input and / or data acquired or generated by the electronic device. The data may include images, videos, and / or sensor data generated by at least one processor of the electronic device (e.g., processor (201)) (e.g., illuminance data around the electronic device acquired from a sensor or sensor hub (e.g., auxiliary processor), attitude data (or orientation data) of the electronic device, temperature inside the electronic device (e.g., display (202)) or temperature of the processor (201)), size information of the display area of the display (202), and / or images acquired through an image sensor of the electronic device (e.g., included in the camera module (210)). The user input may include natural language, touch data acquired through a touch circuit included in the display panel (e.g., used to identify input from a finger and / or stylus), images displayed (and / or to be displayed) on the display panel, and / or videos. As a non-limiting example, the user input may be received by an input / output interface (2110) along with context information. The context information may be described as additional information obtained in relation to the user input. The context information may be related to the state at the time the user input is received (e.g., the state of the electronic device and / or the state of the surroundings of the electronic device (e.g., user state)). For example, the context information may include information about one or more software applications executed within the electronic device at the time the user input is received. For example, the context information may include information about the location of the electronic device (or the location of the user of the electronic device) at the time the user input is received. For example, the user input may be integrated with the context information.For example, a user input incorporating the above situation information into the above input can be received by the input / output interface (2110).
[0337] The input / output interface (2110) may transmit (or provide) an output. The output may include a result (or result information) generated or obtained by the AI system (2100) based on at least part of the input. The format of the output may vary. For example, the output may include natural language. For example, the output may include content (e.g., media content and / or multimedia content). For example, the output may include actions related to the user of the electronic device. For example, the output may have a format according to the user settings of the electronic device.
[0338] The input / output interface (2110) can be described as a user question / response interface (2110).
[0339] The AI framework (2120) can be used to obtain information (or data) about the input from the input / output interface (2110) and to control one or more components related to the AI system (2100) using the obtained information.
[0340] For example, a prompt design component (2121) within an AI framework (2120) can generate or obtain prompts for a generative AI model (2130) (e.g., including a large language model (LLM) or a large multimodal model (LMM)) using the acquired information. For example, the prompt design component (2121) may be described as an AI component that uses a learning algorithm and / or a neural network to provide prompts that are enhanced over time. For example, the prompt design component (2121) can generate or obtain prompts by accessing a knowledge component (e.g., a knowledge repository (2190)) containing user preference data, a prompt library, and / or prompt examples using the acquired information. The generated prompts may be provided to the generative AI model (2130) (e.g., including an LLM or LMM).
[0341] For example, an API / plugin management component (2122) within the AI framework (2120) may be used to support communication for additional information requested (or induced) in relation to the prompt provided (or to be provided) to the generative AI model (2130). For example, the API / plugin management component (2122) may be used to create or establish a channel for communication with various data sources (e.g., knowledge repository (2190)). For example, the API / plugin management component (2122) may support access to at least some of the data sources. For example, the API / plugin management component (2122) may be used to request another component (e.g., application / service component (2180)) that performs feedback (or response) according to the prompt. As a non-limiting example, information obtained (or generated) through the API / plugin management component (2122) may be provided to the prompt design component (2121) for generating a prompt. As a non-limiting example, information obtained (or generated) through the API / plugin management component (2122) may be provided to the generative AI model (2130).
[0342] For example, an improvement component (2123) within the AI framework (2120) can at least partially tune (or adjust) (or change) the result (e.g., content) obtained (or output) from the generative AI model (2130). For example, the improvement component (2123) can determine or verify whether the content obtained from the generative AI model (2130) is related to the input. For example, the improvement component (2123) can determine or verify whether the content obtained from the generative AI model (2130) contains biased content. For example, the improvement component (2123) can determine or verify whether the content obtained from the generative AI model (2130) contains harmful content. For example, the improvement component (2123) can support or assist in performing additional processing to improve the content obtained from the generative AI model (2130). For example, the improvement component (2123) may support providing a hint to the user to improve the content.
[0343] A generative AI model (2130) can be described as an artificial intelligence neural network that generates feedback in response to a prompt. For example, the feedback may include additional data and / or information relative to the prompt, but relative to the prompt. For example, the feedback may include new content relative to the prompt. For example, the generative AI model (2130) may include a model that generates images and / or a model that generates language. For example, the model that generates images may include a generative adversarial network (GAN) and / or a variational autoencoder (VAE). For example, the model that generates images may include a diffusion-based generative model (e.g., a transformer VAE). For example, the model that generates language may include CHAT-GPT 3 and / or CHAT-GPT 4. For example, the generative AI model (2130) may include an LMM that generates the feedback by recognizing text, images, and / or speech.
[0344] As an example without limitation, the AI framework (2120) and / or generative AI model (2130) may be included within an AI module (e.g., including a processing circuit) within the electronic device. For example, the AI module may be operatively coupled with at least one processor of the electronic device. For example, the AI module may be operatively coupled with a display driving circuit of the electronic device. For example, the AI module may be operatively coupled with a sensor hub of the electronic device for one or more sensors within the electronic device.
[0345] According to one embodiment, the electronic device may include a display, a communication circuit, instructions, a memory including one or more storage media, and at least one processor including a processing circuit. The instructions may cause the electronic device to, when executed individually or collectively by the at least one processor, display at least a portion of a web page through the display, identify whether text content and video content are included in the main body of the web page while the display is being displayed, and if the video content and the text content are included in the main body of the web page, obtain subtitle information for the video content, obtain first summary information corresponding to the video content using the subtitle information, obtain second summary information corresponding to the text content, and provide the first summary information and the second summary information.
[0346] For example, when the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause a first visual object for representing the first summary information and a second visual object for representing the second summary information to be superimposed on the display of the at least part of the web page.
[0347] For example, when the above instructions are executed individually or collectively by the at least one processor, the electronic device may obtain a third summary information based on the first summary information and the second summary information and provide the third summary information.
[0348] For example, the above instructions may cause the electronic device to identify the similarity between the first summary information and the second summary information when executed individually or collectively by the at least one processor, identify that the similarity between the first summary information and the second summary information exceeds a reference similarity, and obtain the third summary information based on the similarity exceeding the reference similarity.
[0349] For example, when the above instructions are executed individually or collectively by the at least one processor, if the video content is not included within the main body of the web page, the electronic device may obtain the second summary information corresponding to the text content and provide the second summary information.
[0350] For example, when the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify, from the subtitle information, at least one playback point associated with at least one text block configured to be displayed in the video content, and, before an input for playback of the video content is received, to acquire at least one image of the video content based on the at least one playback point, and to acquire the first summary information based on the subtitle information and the at least one image.
[0351] For example, the above instructions may cause the electronic device to receive the subtitle information for the video content from an external electronic device for providing the video content, through an API (application programming interface) containing the identification information, for requesting the subtitle information, by using the markup language file of the web page to identify that the web page contains video content, to identify among the video content the video content within the display of the at least part of the web page, to obtain identification information of the video content using the markup language file of the web page.
[0352] For example, when the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to request subtitle information for the video content from an external electronic device for providing the video content, receive information from the external electronic device indicating that the subtitle information for the video content is not provided based on the request, obtain identification information for the video content using a markup language file of the web page based on the receipt of the information, determine a prompt for a trained model to obtain the first summary information based on the identification information for the video content, and obtain the first summary information using the trained model based on the prompt.
[0353] For example, the above instructions may cause the electronic device to request subtitle information for the video content from an external electronic device for providing the video content when executed individually or collectively by the at least one processor, and to receive information from the external electronic device indicating that the subtitle information for the video content is not provided based on the request, and in response to receiving the information, to obtain audio data of the video content based on the video content, identify voice data included in the audio data, and obtain the first summary information using text information identified based on the voice data.
[0354] For example, the above instructions may cause the electronic device to display a visual object superimposed on the changed display to request summary information about the other video content, based on identifying that the other video content is included within the changed display based on the change of the display of the at least part of the web page when executed individually or collectively by the at least one processor, and based on identifying that the changed display is maintained for a reference time.
[0355] For example, when the above instructions are executed individually or collectively by the at least one processor, they can acquire multimedia content included in the web page and remove advertising content from the multimedia content. The multimedia content from which the advertising content has been removed may include the text content. When the above instructions are executed individually or collectively by the at least one processor, they can cause the electronic device to acquire the second summary information based on the multimedia content from which the advertising content has been removed.
[0356] For example, the above instructions may cause the electronic device to request the first summary information corresponding to the video content and the second summary information corresponding to the text content based on transmitting the subtitle information and the text content to a server including a trained model when executed individually or collectively by the at least one processor, and to receive the first summary information and the second summary information from the server based on the request.
[0357] For example, when the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to obtain the first summary information based on the subtitle information using the trained model contained in the memory, and to obtain the second summary information based on the text content using the trained model.
[0358] According to one embodiment, a method performed by an electronic device may include: displaying at least a portion of a web page through a display of the electronic device; identifying whether text content and video content are included within the main body of the web page while the display is being shown; obtaining subtitle information for the video content when the video content and the text content are included within the main body of the web page; obtaining first summary information corresponding to the video content using the subtitle information; obtaining second summary information corresponding to the text content; and providing the first summary information and the second summary information.
[0359] For example, the above method may include the operation of displaying a first visual object for representing the first summary information and a second visual object for representing the second summary information in a superposition on at least a portion of the display of the web page.
[0360] For example, the above method may include the operation of obtaining a third summary information based on the first summary information and the second summary information, and the operation of providing the third summary information.
[0361] For example, the above method may include an operation of identifying the similarity between the first summary information and the second summary information, an operation of identifying that the similarity between the first summary information and the second summary information exceeds a reference similarity, and an operation of obtaining the third summary information based on the similarity exceeding the reference similarity.
[0362] For example, the above method may include, when the video content is not included within the main body of the web page, the operation of obtaining the second summary information corresponding to the text content, and the operation of providing the second summary information.
[0363] For example, the above method may include the operation of identifying, from the subtitle information, at least one playback point associated with at least one text block configured to be displayed in the video content; the operation of acquiring at least one image of the video content based on the at least one playback point before an input for playback of the video content is received; and the operation of acquiring the first summary information based on the subtitle information and the at least one image.
[0364] According to one embodiment, the electronic device may include a display, a communication circuit, a memory including instructions and one or more storage media, and at least one processor including a processing circuit. The instructions may cause the electronic device to display at least a portion of a web page through the display when executed individually or collectively by the at least one processor, identify that video content is included in the main body of the web page while the display is being displayed, obtain subtitle information regarding the video content and at least one image regarding the video content, obtain summary information regarding the video content based on the subtitle information and the at least one image, and provide the summary information.
[0365] According to one embodiment, the electronic device may include a display, a memory including instructions and one or more storage media, and at least one processor including a processing circuit. The instructions may cause the electronic device to receive an input requesting a summary of said web page while a user interface including at least a portion of a web page is displayed through said display when executed individually or collectively by said at least one processor, and based on said input, to identify video content within said web page, obtain text information regarding said video content and at least one image regarding said video content, and provide summary information generated based on said text information regarding said video content and at least one image regarding said video content through said display.
[0366] For example, text information regarding the video content may include at least one of a first text content obtained based on the video content and a second text content displayed within the web page in conjunction with the video content.
[0367] For example, the summary information may be constructed based on a plurality of sentences. The summary information may include a first section comprising at least one sentence regarding the video content among the plurality of sentences, and a second section comprising the remaining sentences regarding the second text content among the plurality of sentences.
[0368] For example, the above instructions may cause the electronic device to determine the display method of the first section and the second section based on identification information for the web page when executed individually or collectively by the at least one processor.
[0369] For example, the above instructions may cause the electronic device to identify the first text content as the text information based on identifying that the number of characters according to the second text content is less than a reference number when executed individually or collectively by the at least one processor.
[0370] For example, the above instructions may cause the electronic device to identify the similarity between the first text content and the second text content when executed individually or collectively by the at least one processor, and to identify the first text content as the text information based on whether the identified similarity exceeds a reference similarity (or whether the identified similarity is greater than or equal to the reference similarity).
[0371] For example, the above instructions may cause the electronic device to acquire the first text data based on subtitle data regarding the video content when executed individually or collectively by the at least one processor. The subtitle data may include a plurality of start times for displaying the subtitle within the video content, a plurality of display durations corresponding to the plurality of start times, and a plurality of character blocks corresponding to the plurality of display durations.
[0372] For example, when the above instructions are executed individually or collectively by the at least one processor, the electronic device may be configured to identify a reference number of character blocks among a plurality of character blocks based on the plurality of display periods, and to acquire at least one image of the video content based on at least one frame of the video data corresponding to the reference number of character blocks.
[0373] For example, the plurality of character blocks may include at least one character block containing a designated identification symbol for indicating a sound effect. The summary information may include information regarding sound effects according to the at least one character block.
[0374] For example, the above instructions may cause the electronic device to identify audio data for the video content and, based on the audio data, obtain text information regarding the video content when executed individually or collectively by the at least one processor.
[0375] For example, the above instructions may cause the electronic device to identify a plurality of frames of the video content relating to voice data among the audio data when executed individually or collectively by the at least one processor, and to acquire the at least one image based on at least one frame among the identified plurality of frames according to the display period of each of the plurality of frames.
[0376] For example, the above instructions may cause the electronic device to identify a plurality of frames corresponding to the text information and identify at least one frame among the plurality of frames as the at least one image, based on identifying that the sum of the playback intervals of the video content corresponding to the text information, when executed individually or collectively by the at least one processor, exceeds a reference playback interval.
[0377] For example, the above instructions may cause the electronic device to identify at least one other frame within the entire playback interval of the video content and identify the at least one other frame as the at least one image, based on identifying that the sum of the playback intervals of the video content corresponding to the text information, when executed individually or collectively by the at least one processor, is less than or equal to a reference playback interval.
[0378] For example, the above instructions may cause the electronic device to identify at least one visual object representing at least one character included in a plurality of frames of the video content when executed individually or collectively by the at least one processor, and to obtain text information regarding the video content based on the at least one character identified according to the at least one visual object.
[0379] For example, the above instructions may cause the electronic device to acquire at least some of the text information regarding the video content before the input is received, based on the user interface including at least a portion of the web page being displayed through the display when executed individually or collectively by the at least one processor.
[0380] For example, the above instructions may cause the electronic device to identify a plurality of video contents included in the web page when executed individually or collectively by the at least one processor, identify a plurality of regions for displaying the plurality of video contents within the web page, and determine the video content for acquiring the at least one image among the plurality of video contents based on the size of each of the plurality of regions.
[0381] For example, the above instructions may cause the electronic device to identify a plurality of video contents included in the web page when executed individually or collectively by the at least one processor, display an object in the user interface for identifying the video content for acquiring the at least one image among the plurality of video contents, and identify the video content for acquiring the at least one image based on inputs associated with the object.
[0382] According to one embodiment, a method performed by an electronic device may include receiving an input requesting a summary of a web page while a user interface including at least a portion of a web page is displayed through a display of the electronic device; identifying video content within the web page based on the input; acquiring text information regarding the video content and at least one image regarding the video content; and providing summary information generated based on the text information regarding the video content and at least one image regarding the video content through the display.
[0383] For example, text information regarding the video content may include at least one of a first text content obtained based on the video content and a second text content displayed within the web page in conjunction with the video content.
[0384] For example, the summary information may be constructed based on a plurality of sentences. The summary information may include a first section comprising at least one sentence regarding the video content among the plurality of sentences, and a second section comprising the remaining sentences regarding the second text content among the plurality of sentences.
[0385] For example, the above method may include an operation to determine the display method of the first section and the second section based on identification information for the web page.
[0386] For example, the above method may include an operation of identifying the first text content as the text information based on identifying that the number of characters according to the second text content is less than a reference number.
[0387] For example, the above method may include an operation of identifying a similarity between the first text content and the second text content, and an operation of identifying the first text content as text information based on the fact that the identified similarity exceeds a reference similarity.
[0388] For example, the above method may include an operation of obtaining the first text data based on subtitle data regarding the video content. The subtitle data may include a plurality of start times for displaying the subtitle within the video content, a plurality of display durations corresponding to the plurality of start times, and a plurality of character blocks corresponding to the plurality of display durations.
[0389] For example, the above method may include the operation of identifying a reference number of character blocks among a plurality of character blocks based on the plurality of display periods, and the operation of acquiring at least one image of the video content based on at least one frame of the video data corresponding to the reference number of character blocks.
[0390] For example, the plurality of character blocks may include at least one character block containing a designated identification symbol for indicating a sound effect. The summary information may include information regarding sound effects according to the at least one character block.
[0391] For example, the above method may include an operation of identifying audio data for the video content, and an operation of obtaining text information regarding the video content based on the audio data.
[0392] For example, the above method may include an operation of identifying a plurality of frames of the video content relating to voice data among the audio data, and an operation of acquiring the at least one image based on at least one frame among the plurality of frames identified according to the display period of each of the plurality of frames.
[0393] For example, the above method may include an operation of identifying a plurality of frames corresponding to the text information based on identifying that the sum of the playback intervals of the video content corresponding to the text information exceeds a reference playback interval, and an operation of identifying at least one frame among the plurality of frames as the at least one image.
[0394] For example, the above method may include the operation of identifying at least one other frame within the entire playback interval of the video content, based on identifying that the sum of the playback intervals of the video content corresponding to the text information is less than or equal to a reference playback interval, and the operation of identifying the at least one other frame as the at least one image.
[0395] For example, the above method may include an operation of identifying at least one visual object representing at least one character included in a plurality of frames of the video content, and an operation of obtaining text information regarding the video content based on the at least one character identified according to the at least one visual object.
[0396] For example, the above method may include an operation of obtaining at least some of the text information regarding the video content before the input is received, based on the user interface including at least a portion of the web page being displayed through the display.
[0397] For example, the above method may include the operation of identifying a plurality of video contents included in the web page, the operation of identifying a plurality of regions for displaying the plurality of video contents within the web page, and the operation of determining a video content for acquiring at least one image among the plurality of video contents based on the size of each of the plurality of regions.
[0398] For example, the above method may include an operation of identifying a plurality of video contents included in the web page, an operation of displaying an object in the user interface for identifying the video content for acquiring at least one image among the plurality of video contents, and an operation of identifying the video content for acquiring at least one image based on an input related to the object.
[0399] According to one embodiment, the electronic device may include a display, a memory including instructions and one or more storage media, and at least one processor including a processing circuit. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause to display a first user interface including text content and video content through the display, receive user input for displaying summary information in relation to the first user interface, obtain subtitle information from the video content, generate the summary information based on the subtitle information regarding the text content and the video content, and display a second user interface including the summary information through the display.
[0400] For example, the above instructions may cause the electronic device to acquire multimedia content regarding at least one of the text content or the video content when executed individually or collectively by the at least one processor, and to generate summary information based on the text content, the subtitle information, and the multimedia content.
[0401] According to the embodiments described above, the electronic device can provide summary information for a web page containing video content and text content. The electronic device can provide advanced summary information using video subtitle data and at least one frame. The electronic device can provide summary information for video content even when the user has not played the video content.
[0402] The electronic device according to the embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the aforementioned devices.
[0403] The embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, each of phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish a component from another component and do not limit the components in any other aspect (e.g., importance or order). Where any component (e.g., the first) is referred to as "coupled" or "connected" to another component (e.g., the second), with or without the terms "functionally" or "communicationally," it means that said component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0404] In one embodiment of this document, the term “module” used may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0405] One embodiment of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.
[0406] According to one embodiment, the method according to the embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., CD-ROM (compact disc read-only memory)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0407] According to one embodiment, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to one embodiment, one or more of the components or operations among the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to one embodiment, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.< / video> < / output> < / video>
Claims
1. In an electronic device, display; Communication circuit; Memory comprising instructions and one or more storage media; and It includes at least one processor including a processing circuit, and When the above instructions are executed individually or collectively by the at least one processor, Displaying at least a portion of a web page through the display, While the above indication is displayed, determine whether text content and video content are included within the main body of the web page, and If the above video content and the above text content are included within the above main body of the web page: Obtain subtitle information for the above video content, and Using the above subtitle information, first summary information corresponding to the above video content is obtained, and Obtaining second summary information corresponding to the above text content, and Causing the electronic device to provide the above first summary information and the above second summary information, Electronic device.
2. In claim 1, when the instructions are executed individually or collectively by the at least one processor, The electronic device causing the first visual object for displaying the first summary information and the second visual object for displaying the second summary information to be superimposed on the display of at least a portion of the web page. Electronic device.
3. In claim 1, when the instructions are executed individually or collectively by the at least one processor, Based on the first summary information and the second summary information above, obtain third summary information, and Causing the electronic device to provide the above third summary information, Electronic device.
4. In claim 3, when the instructions are executed individually or collectively by the at least one processor, Identifying the similarity between the first summary information and the second summary information, Identifying that the similarity between the first summary information and the second summary information exceeds the reference similarity, Causing the electronic device to obtain the third summary information based on the similarity exceeding the reference similarity above, Electronic device.
5. In claim 1, when the instructions are executed individually or collectively by the at least one processor, If the above video content is not included within the main body of the above web page: Obtain the second summary information corresponding to the text content above, and Causing the electronic device to provide the above second summary information, Electronic device.
6. In claim 1, when the instructions are executed individually or collectively by the at least one processor, From the subtitle information above, identify at least one playback point associated with at least one text block configured to be displayed in the video content, and Before receiving an input for playback of the video content, based on the at least one playback point, at least one image relating to the video content is obtained, and Causing the electronic device to obtain the first summary information based on the subtitle information and the at least one image, Electronic device.
7. In claim 1, when the instructions are executed individually or collectively by the at least one processor, Using the markup language file of the above web page, identify that the web page includes video content, and Among the above video contents, identify the video content within the display of at least a portion of the web page, and Using the markup language file of the above web page, identification information of the above video content is obtained, and Causing the electronic device to receive the subtitle information for the video content from an external electronic device for providing the video content through an API (application programming interface) including the identification information for requesting the subtitle information. Electronic device.
8. In claim 1, when the instructions are executed individually or collectively by the at least one processor, Request the subtitle information for the video content from an external electronic device for providing the video content, and Based on the above request, receiving information from the external electronic device indicating that the subtitle information for the video content is not provided, and Based on the reception of the above information, identification information for the video content is obtained using the markup language file of the web page, and Based on the identification information for the video content, a prompt for a trained model is determined to obtain the first summary information, and Causing the electronic device to obtain the first summary information using the trained model based on the above prompt, Electronic device.
9. In claim 1, when the instructions are executed individually or collectively by the at least one processor, Request the subtitle information for the video content from an external electronic device for providing the video content, and Based on the above request, receiving information from the external electronic device indicating that the subtitle information for the video content is not provided, and In response to receiving the above information, based on the video content, audio data of the video content is obtained, and Identify voice data included in the above audio data, and Using text information identified based on the voice data above, causing the electronic device to obtain the first summary information, Electronic device.
10. In claim 1, when the instructions are executed individually or collectively by the at least one processor, Based on the change of at least some of the above-mentioned display of the web page, identifying that other video content is included within the changed display, and Based on identifying that the above-mentioned changed display is maintained for a reference time, the electronic device causes to display a visual object for requesting summary information about the other video content superimposed on the above-mentioned changed display, Electronic device.
11. In claim 1, when the instructions are executed individually or collectively by the at least one processor, Acquire multimedia content included in the above web page, and Ad content is removed from the multimedia content, and the multimedia content from which the ad content has been removed includes the text content. Causing the electronic device to obtain the second summary information based on the multimedia content from which the advertisement content has been removed, Electronic device.
12. In claim 1, when the instructions are executed individually or collectively by the at least one processor, Based on transmitting the above subtitle information and the above text content to a server including a trained model, the first summary information corresponding to the video content and the second summary information corresponding to the text content are requested, and Causing the electronic device to receive the first summary information and the second summary information from the server based on the above request, Electronic device.
13. In claim 1, when the instructions are executed individually or collectively by the at least one processor, Using the trained model included in the memory above, the first summary information is obtained based on the subtitle information, and Using the above-mentioned trained model, the electronic device causes to obtain the above-mentioned second summary information based on the above-mentioned text content, Electronic device.
14. In a method performed by an electronic device, The operation of displaying at least a portion of a web page through the display of the electronic device; An action of identifying whether text content and video content are included within the main body of the web page while the above indication is displayed; If the above video content and the above text content are included within the above main body of the web page: An operation to obtain subtitle information for the above video content; The operation of obtaining first summary information corresponding to the video content using the subtitle information above; The operation of obtaining second summary information corresponding to the above text content; and The operation of providing the first summary information and the second summary information, method.
15. A non-transient computer-readable storage medium storing one or more programs, wherein the one or more programs, when executed by at least one processor of an electronic device having a display and communication circuit, Displaying at least a portion of a web page through the display, While the above indication is displayed, determine whether text content and video content are included within the main body of the web page, and If the above video content and the above text content are included within the above main body of the web page: Obtain subtitle information for the above video content, and Using the above subtitle information, first summary information corresponding to the above video content is obtained, and Obtaining second summary information corresponding to the above text content, and Instructions that cause the electronic device to provide the first summary information and the second summary information, Non-transient computer-readable storage media.
Citation Information
Patent Citations
Recording and playback system
JP7137815B2
Apparatus and method for detecting advertisment of moving-picture, and compter-readable storage storing compter program controlling the apparatus
KR100707189B1
Method and apparatus for generating summarized data, and a server for the same
KR101956373B1
System and method selecting and managing moving image
KR1020180062005A
Moving Picture Summary Play Device, Moving Picture Summary Providing Server and Methods Thereof
KR102055766B1