Electronic device and method for generating summary data
The electronic device uses AI models to classify and generate summary data from user logging data, addressing the need for efficient summarization by creating personalized summaries based on user inputs.
Patent Information
- Application Number
- PCT/KR2025/006295
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-22
- Filing Date
- 2025-05-09
- Publication Date
- 2025-12-26
AI Technical Summary
There is a growing need to efficiently summarize users' logging data using artificial intelligence models to provide diverse information, as existing rule-based smart systems are being replaced by deep learning-based AI systems.
An electronic device employs a first AI model to classify logging data into categories and generate first summary data, which is then selected and processed using a second AI model to create second summary data based on user input, leveraging generative AI models for efficient data summarization.
The approach allows for the generation of personalized and efficient summary data tailored to user requests, enhancing the ability to summarize daily life activities and experiences.
Smart Images

Figure KR2025006295_26122025_PF_FP_ABST
Abstract
Description
Electronic device and method for generating summary data
[0001] The present disclosure relates to an electronic device and method for generating summary data.
[0002] Recently, artificial intelligence systems that achieve human-level intelligence are being utilized in various fields. Unlike existing rule-based smart systems, AI systems are machines that learn, make decisions, and become intelligent on their own. As AI systems become more used, their recognition rates improve and their ability to understand user preferences more accurately is increasing. As a result, existing rule-based smart systems are gradually being replaced by deep learning-based AI systems.
[0003] Artificial intelligence technology consists of machine learning (e.g., deep learning) and elemental technologies that utilize machine learning.
[0004] Machine learning is an algorithm technology that classifies / learns the characteristics of input data on its own, and element technology is a technology that imitates the functions of the human brain, such as cognition and judgment, by utilizing machine learning algorithms such as deep learning, and is composed of technical fields such as linguistic understanding, visual understanding, inference / prediction, knowledge representation, and motion control.
[0005] The various fields in which artificial intelligence technology is applied are as follows. Linguistic understanding refers to the technology that recognizes, applies, and processes human language / text, including natural language processing, machine translation, dialogue systems, question-answering, and speech recognition / synthesis. Visual understanding refers to the technology that recognizes and processes objects similar to human vision, including object recognition, object tracking, image search, person recognition, scene understanding, spatial understanding, and image enhancement. Inference prediction refers to the technology that logically infers and predicts information by judging it, including knowledge / probability-based inference, optimization prediction, preference-based planning, and recommendations. Knowledge representation refers to the technology that automatically processes human experience information into knowledge data, including knowledge construction (data creation / classification) and knowledge management (data utilization). Motion control refers to the technology that controls the movement of autonomous vehicles and robots, including movement control (navigation, collision, driving), and manipulation control (behavior control).
[0006] Meanwhile, technologies for generating and providing logging data that records users' daily lives are advancing. In particular, recent electronic devices or programs can summarize and provide data using summary models acquired through artificial intelligence learning.
[0007] Accordingly, there is a growing need to provide users with diverse information through the ability to efficiently summarize users' logging data using an artificial intelligence model.
[0008] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.
[0009] According to one embodiment, a method for generating summary data by an electronic device may be provided, including: obtaining a plurality of logging data including a plurality of audio data by recording sounds around the electronic device; obtaining a plurality of first summary data classified according to a plurality of categories by analyzing the plurality of logging data; receiving a user input requesting generation of second summary data; selecting at least one first summary data related to the user input from among the plurality of first summary data; generating a second input prompt for generation of the second summary data based on the user input and the selected at least one first summary data; and obtaining the second summary data by applying the second input prompt to a second artificial intelligence model.
[0010] According to one embodiment, an electronic device for generating summary data may be provided, comprising: a camera; a microphone; a memory for storing commands; and one or more processors; wherein the commands, when executed by the one or more processors, cause the electronic device to: record sounds surrounding the electronic device through the microphone, thereby obtaining a plurality of logging data including a plurality of audio data; analyze the plurality of logging data, thereby obtaining a plurality of first summary data classified according to a plurality of categories; receive a user input requesting generation of second summary data; select at least one first summary data related to the user input from among the plurality of first summary data; generate a second input prompt for generation of the second summary data based on the user input and the selected at least one first summary data; and obtain the second summary data by applying the second input prompt to a second artificial intelligence model.
[0011] According to one embodiment, a computer-readable recording medium having recorded thereon a program for executing a method, the method comprising: obtaining a plurality of logging data including a plurality of audio data by recording sounds around the electronic device; obtaining a plurality of first summary data classified according to a plurality of categories by analyzing the plurality of logging data; receiving a user input requesting generation of second summary data; selecting at least one first summary data related to the user input from among the plurality of first summary data; generating a second input prompt for generation of the second summary data based on the user input and the selected at least one first summary data; and obtaining the second summary data by applying the second input prompt to a second artificial intelligence model.
[0012] According to one embodiment, a method for an electronic device to obtain summary data may be provided, including: obtaining a plurality of user logging data; analyzing the plurality of logging data to obtain a plurality of first summary data classified according to a plurality of categories; receiving a user input requesting generation of second summary data; selecting at least one first summary data related to the user input from among the plurality of first summary data; transmitting the user input and the selected first summary data to another electronic device; and receiving, from the other electronic device, the second summary data generated based on the user input and the selected at least one first summary data, wherein the second summary data is output from a generative artificial intelligence model by applying a second input prompt generated by the other electronic device to the generative artificial intelligence model.
[0013] According to one embodiment, an electronic device for generating summary data may be provided, comprising: a camera; a microphone; a memory for storing commands; and one or more processors; wherein the commands, when executed by the one or more processors, cause the electronic device to: obtain a plurality of first summary data classified according to a plurality of categories; receive a user input requesting generation of second summary data; select at least one first summary data related to the user input from among the plurality of first summary data; transmit the user input and the selected first summary data to another electronic device; and receive, from the other electronic device, the second summary data generated based on the user input and the selected at least one first summary data, wherein the second summary data is output from a generative artificial intelligence model by applying a second input prompt generated by the other electronic device to the generative artificial intelligence model.
[0014] According to one embodiment, a computer-readable recording medium having recorded thereon a program for executing the following operations: obtaining a plurality of user logging data; analyzing the plurality of logging data to obtain a plurality of first summary data classified according to a plurality of categories; receiving a user input requesting generation of second summary data; selecting at least one first summary data related to the user input from among the plurality of first summary data; transmitting the user input and the selected first summary data to another electronic device; and receiving, from the other electronic device, the second summary data generated based on the user input and the selected at least one first summary data, wherein the second summary data is output from a generative artificial intelligence model by applying a second input prompt generated by the other electronic device to the generative artificial intelligence model.
[0015] FIG. 1 is a diagram illustrating an overview of an electronic device according to one embodiment of the present invention for generating summary data based on logging data.
[0016] FIG. 2 is a flowchart of a method for an electronic device to generate summary data based on logging data according to one embodiment.
[0017] FIG. 3 is a flowchart of a method for an electronic device according to one embodiment to obtain a plurality of first summary data classified according to a plurality of categories by analyzing a plurality of logging data.
[0018] FIG. 4 is a diagram illustrating an example of a first artificial intelligence model generating first summary data based on input values according to one embodiment.
[0019] FIG. 5 is a diagram illustrating an example of first summary data classified according to a plurality of categories according to one embodiment.
[0020] FIG. 6 is a flowchart of a method for an electronic device to select first summary data related to a user input for generating second summary data according to one embodiment.
[0021] FIG. 7 is a diagram illustrating an example in which first summary data related to user input is selected from first summary data classified according to a plurality of categories according to one embodiment.
[0022] FIG. 8 is a diagram illustrating an example of a second artificial intelligence model generating second summary data based on input values according to one embodiment.
[0023] FIG. 9A is a diagram illustrating an example of first summary data according to one embodiment.
[0024] FIG. 9b is a diagram illustrating an example of first summary data according to one embodiment.
[0025] FIG. 9c is a diagram illustrating an example of first summary data selected based on user input and second summary data generated based on the selected first summary data according to one embodiment.
[0026] FIG. 10 is a diagram illustrating an example of second summary data generated based on first summary data from a plurality of electronic devices according to one embodiment.
[0027] FIG. 11 is a diagram illustrating an example of second summary data generated according to user input related to a specific situation according to one embodiment.
[0028] FIG. 12 is a diagram illustrating an example of second summary data generated according to user input according to one embodiment.
[0029] FIG. 13 is a diagram showing an example of first summary data displayed on a screen of an electronic device when the electronic device is a watch-type electronic device according to one embodiment.
[0030] FIG. 14 is a diagram illustrating an example of displaying various types of summary data on a screen of an electronic device according to one embodiment.
[0031] FIG. 15 is a diagram illustrating an example of first summary data stored in an electronic device and another electronic device according to one embodiment.
[0032] FIG. 16 is a flowchart of a method for generating second summary data by an electronic device in conjunction with another electronic device according to one embodiment.
[0033] FIG. 17 is a flowchart of a method for generating second summary data by an electronic device in conjunction with another electronic device according to one embodiment.
[0034] FIG. 18 is a block diagram of an electronic device within a network environment, according to one embodiment.
[0035] Figure 19 is a block diagram of an electronic device according to one embodiment.
[0036] FIG. 20 is a diagram illustrating a system including a generative artificial intelligence model according to one embodiment.
[0037] Below, embodiments of the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In addition, for the purpose of clearly explaining the present disclosure in the drawings, parts irrelevant to the description are omitted, and similar parts are designated with similar reference numerals throughout the specification.
[0038] The terms used in this disclosure are described as currently common terms, taking into account the functions mentioned herein. However, these terms may mean various other terms depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Therefore, the terms used in this disclosure should not be interpreted solely based on their names, but rather based on the meanings of the terms and the overall content of this disclosure.
[0039] Additionally, while terms such as first, second, etc. may be used to describe various components, the components should not be limited by these terms. These terms are used to distinguish one component from another.
[0040] Throughout the specification, when a part is said to be "connected" to another part, this includes not only the cases where the parts are "directly connected" but also the cases where the parts are "electrically connected" with other elements intervening. Furthermore, when a part is said to "include" a component, this does not exclude other components, but rather includes other components, unless otherwise stated.
[0041] The phrases “in one embodiment” and the like appearing in various places throughout this disclosure do not necessarily all refer to the same embodiment.
[0042] An embodiment of the present disclosure may be represented by functional block configurations and various processing steps. Some or all of these functional blocks may be implemented by various hardware and / or software configurations that perform specific functions. For example, the functional blocks of the present disclosure may be implemented by one or more microprocessors or by circuit configurations for a given function. Furthermore, for example, the functional blocks of the present disclosure may be implemented in various programming or scripting languages. The functional blocks may be implemented by algorithms that execute on one or more processors. Furthermore, the present disclosure may employ conventional techniques for electronic configuration, signal processing, and / or data processing. Terms such as "mechanism," "element," "means," and "configuration" may be used broadly and are not limited to mechanical and physical configurations.
[0043] Additionally, the connecting lines or connecting members between components depicted in the drawings are merely exemplary representations of functional connections and / or physical or circuit connections. In an actual device, connections between components may be represented by various functional connections, physical connections, or circuit connections that may be replaced or added.
[0044] According to one embodiment, the logging data may be data recorded or generated in relation to the user's daily activities and actions. The logging data may include, but is not limited to, media data such as image data (e.g., still image data, video data), audio data, and / or text data. For example, the logging data may include images captured of the surroundings of the electronic device, audio data recorded of the surroundings of the electronic device, and biometric information of the user of the electronic device. For example, the logging data may include, but is not limited to, location data of the electronic device indicating the user's movement path and / or places visited by the user, data regarding the user's activities (e.g., walking, running, driving, etc.) collected through the electronic device, data regarding programs or applications used by the user, biometric data of the user (e.g., heart rate, sleep patterns, amount of exercise, etc.), and / or sensing data regarding the surrounding environment of the electronic device carried by the user (e.g., ambient temperature, noise level, lighting conditions, etc.). Additionally, for example, the logging data may include information about functions executed on the electronic device (1000) and / or usage history of devices within the electronic device (1000).
[0045] According to one embodiment, the first AI model may be an AI model trained to generate first summary data based on logging data. For example, the first AI model may be a generative AI model trained to generate first summary data based on logging data and / or analysis results of the logging data.
[0046] According to one embodiment, the first summary data may be summary data generated based on logging data and classified by categories. The first summary data may be used to generate second summary data according to a user input request.
[0047] According to one embodiment, the first input prompt may be an input prompt input to the first artificial intelligence model. If the first artificial intelligence model is a generative artificial intelligence model, the input prompt input to the generative artificial intelligence model may be generated by various types of commands. For example, the first input prompt may include at least text in a natural language format, but is not limited thereto. For example, the first input prompt may include content regarding how the first artificial intelligence model generates the first summary data, and may include elements for determining the style, topic, and / or length of the first summary data generated by the first artificial intelligence model.
[0048] The second summary data according to one embodiment may be summary data generated based on summary data selected from among the first summary data according to a predetermined criterion.
[0049] According to one embodiment, the second AI model may be an AI model trained to generate second summary data based on the first summary data. For example, the second AI model may be a generative AI model for generating second summary data based on first summary data selected according to a user input. For example, the second AI model may be the same as or different from the first AI model. If the second AI model is the same as the first AI model, the electronic device may obtain the first summary data and the second summary data using a single AI model.
[0050] According to one embodiment, the second input prompt may be an input prompt input to the second artificial intelligence model. When the second artificial intelligence model is a generative artificial intelligence model, the input prompt input to the generative artificial intelligence model may be generated by various types of commands. For example, the second input prompt may include at least text in a natural language format, but is not limited thereto. For example, the second input prompt may include content regarding how the second artificial intelligence model generates the second summary data, and may include elements for determining the style, topic, and / or length of the second summary data generated by the second artificial intelligence model.
[0051] The present disclosure will be described in detail with reference to the attached drawings below.
[0052] FIG. 1 is a diagram illustrating an overview of an electronic device according to one embodiment of the present invention for generating summary data based on logging data.
[0053] Referring to FIG. 1, an electronic device (1000) according to one embodiment can obtain second summary data (16) by sequentially summarizing a plurality of logging data (10) about a user.
[0054] An electronic device (1000) according to one embodiment may acquire a plurality of logging data (10) related to the daily life of a user of the electronic device (1000), and may acquire a plurality of first summary data (12) from the logging data (10) using a first artificial intelligence model (11) trained to generate summary data. In this case, the plurality of first summary data (12) may be classified according to a plurality of categories. According to one embodiment, the electronic device (1000) may directly generate the first summary data, or may receive first summary data generated by another electronic device (not shown) from the other electronic device.
[0055] According to one embodiment, the electronic device (1000) may receive a user input (13) for generating summary data and select first summary data (14) related to the user input from among a plurality of first summary data based on the intention of the user input. In addition, the electronic device (1000) may obtain second summary data (16) from the selected first summary data (14) using a second artificial intelligence model (15) trained for generating summary data. For example, the electronic device (1000) may select the first summary data (14) and obtain the second summary data (16) when a user input (13) such as “What did you do on the way to your recent travel destination?” is received.
[0056] An electronic device (1000) according to one embodiment can generate and store a plurality of first summary data generated primarily based on logging data in advance according to a category, and when a user requests generation of second summary data, selects some of the plurality of first summary data to generate second summary data, thereby generating summary data related to the user's daily life more efficiently.
[0057] An electronic device (1000) according to one embodiment may be, but is not limited to, a smartphone, a tablet PC, a PC, a smart TV, a mobile phone, a personal digital assistant (PDA), a laptop, a media player, a micro server, a global positioning system (GPS) device, an e-book reader, a digital broadcasting terminal, a navigation device, a kiosk, an MP3 player, a digital camera, a home appliance, and other mobile or non-mobile computing devices. In addition, the electronic device (1000) may be a wearable device such as a watch, glasses, a hair band, an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) pin, and a ring, all of which have communication and data processing functions. However, the electronic device (1000) is not limited thereto, and may include any type of device capable of displaying a message acquired through a network.
[0058] FIG. 2 is a flowchart of a method for an electronic device to generate summary data based on logging data according to one embodiment.
[0059] In operation 210, the electronic device (1000) may obtain a plurality of logging data including a plurality of images and a plurality of voice data. The logging data may be data recorded or generated in relation to the user's daily activities and actions. According to one embodiment, the electronic device (1000) may obtain images obtained by photographing the surroundings of the electronic device (1000) using a camera of the electronic device (1000) and voice data recorded through a microphone of the electronic device (1000), and may obtain a plurality of logging data including the obtained images and voice data.
[0060] According to one embodiment, the electronic device (1000) can obtain logging data according to a preset cycle. In this case, the cycle at which the electronic device (1000) generates logging data can be preset.
[0061] According to one embodiment, the electronic device (1000) may acquire logging data when a specific event occurs. In this case, the specific event may include, but is not limited to, an event in which the location of the electronic device (1000) changes by a predetermined threshold or more, an event in which a predetermined application or function is executed within the electronic device (1000), and / or an event in which sensing data detected by a sensor of the electronic device (1000) changes by a threshold or more.
[0062] According to one embodiment, the logging data may include media data such as image data (e.g., still image data, video data), audio data, and / or text data. For example, the logging data may include images captured of the surroundings of the electronic device (1000) and audio data recorded of the surroundings of the electronic device (1000). However, the type of logging data is not limited thereto. For example, the logging data may include biometric information of the user sensed by a biometric sensor of the electronic device (1000). For example, the logging data may include location data of the electronic device (1000) indicating the user's movement path and / or places visited by the user, and the location data may be acquired by a GPS sensor within the electronic device (1000). For example, the logging data may include data about the user's activity (e.g., walking, running, driving, etc.) collected through the electronic device (1000), and the electronic device (1000) may obtain data about the user's activity by analyzing sensing data obtained through a GPS sensor, a motion sensor, a gyro sensor, and / or a geomagnetic sensor within the electronic device (1000). For example, the logging data may include data about a program or application used by the user, and the electronic device (1000) may obtain data about the program or application used by the user based on the execution history of the program or application executed within the electronic device (1000). For example, the logging data may include sensing data about the user's biometric data (e.g., heart rate, sleep pattern, amount of exercise, etc.) and / or the surrounding environment of the electronic device (1000) carried by the user (e.g., ambient temperature, noise level, lighting condition, etc.).For example, if the logging data includes the user's biometric data (e.g., heart rate, sleep pattern, amount of exercise, etc.), the first summary data and the second summary data may include information about the user's emotions (e.g., excitement, joy, sadness) identified by analyzing the user's biometric data. For example, if the summary data is generated based on the logging data including image data and biometric information collected over a certain period of time, the electronic device (1000) may include information about the user's emotions related to the image data in the summary data while generating the summary data based on the image data.
[0063] In one embodiment, the electronic device (1000) may receive logging data from another electronic device (not shown). In this case, the other electronic device (not shown) may include, but is not limited to, a wearable electronic device worn by a user and / or a sensing device installed within a space designated by the user.
[0064] In one embodiment, the electronic device (1000) may limit the collection of logging data. For example, if the remaining battery level of the electronic device (1000) is below a threshold, and if the privacy of the logging data is an issue, the electronic device (1000) may at least temporarily suspend the collection of logging data. For example, the electronic device (1000) may collect logging data using some sensors when the remaining battery level is below a threshold. For example, the electronic device (1000) may constantly use a microphone with relatively low power consumption to detect the voice of the user of the electronic device (1000), and may operate a camera to capture an image when the user's voice is detected.
[0065] For example, if the privacy of logging data is a concern, the electronic device (1000) may delete at least a portion of the logging data that may pose a privacy issue. For example, if the electronic device (1000) continuously records, the electronic device may delete all recorded voices except for the voices of conversations in which the user participated or the voices of the user talking to himself.
[0066] Although the electronic device (1000) has been described above as acquiring multiple logging data including multiple images and multiple audio data, the present invention is not limited thereto. For example, if the electronic device (1000) does not include a display or has a low-resolution display, the electronic device (1000) may acquire multiple logging data including multiple audio data but not including image data.
[0067] In operation 220, the electronic device (1000) can obtain a plurality of first summary data classified according to a plurality of categories by analyzing a plurality of logging data.
[0068] According to one embodiment, the electronic device (1000) may obtain a plurality of first summary data from a plurality of logging data. The first summary data may be summary data generated based on the logging data and classified according to categories. For example, the first summary data may include at least a portion of the logging data and / or data generated as a result of analysis of the logging data. For example, when the logging data includes voice data, the first summary data may include at least a portion of text converted from the voice data and / or summary text summarizing the text converted from the voice data. For example, when the logging data includes image data (e.g., still images and / or moving images), the first summary data may include at least a portion of the image data, an image generated by summarizing the image data, and / or text describing the image data. For example, when the logging data includes biometric data of a user (e.g., heart rate, sleep patterns, amount of exercise, etc.), the first summary data may include at least a portion of the biometric data of the user and / or text regarding the user's emotion (e.g., excitement, joy, sadness) identified by analyzing the biometric data of the user. However, examples of logging data and first summary data are not limited thereto.
[0069] According to one embodiment, the electronic device (1000) may acquire a plurality of first summary data by inputting a plurality of logging data and / or a result of a primary analysis of the plurality of logging data into at least one first artificial intelligence model trained to generate summary data. In this case, the input data input to the first artificial intelligence model may have a preset format supported by the first artificial intelligence model. For example, when the first artificial intelligence model is a generative artificial intelligence model, the input value input to the first artificial intelligence model may include, but is not limited to, a first input prompt in a natural language format.
[0070] According to one embodiment, the electronic device (1000) may generate a first input prompt to be input into a first artificial intelligence model, and input the first input prompt and a plurality of logging data into the first artificial intelligence model. The first input prompt may include, for example, information on attributes of first summary data to be generated by the first artificial intelligence model. The attributes of the first summary data may include preset information related to the type and number of media data to be included in the first summary data, a period corresponding to the first summary data, and a method of generating the first summary data. In addition, for example, the first input prompt may include elements for determining the style, topic, and / or length of the first summary data generated by the first artificial intelligence model. In this case, for example, the first artificial intelligence model may be a generative artificial intelligence model trained to generate summary data based on the logging data.
[0071] According to one embodiment, the electronic device (1000) can obtain first summary data by primarily analyzing logging data and inputting the analysis result of the logging data into a first artificial intelligence model. In this case, according to one embodiment, the electronic device (1000) can select at least one artificial intelligence model from among a plurality of artificial intelligence models for analyzing the logging data according to the type of the logging data, and can primarily obtain an analysis result of the logging data using the selected at least one artificial intelligence model. In addition, the electronic device (1000) can obtain first summary data by inputting predetermined input data including the analysis result of the logging data into the first artificial intelligence model.
[0072] For example, the electronic device (1000) can obtain image data tagged with an object identification value by inputting logging data, which is image data (e.g., still image data and / or video data), into an artificial intelligence model (e.g., vision model) trained for object recognition. In this case, for example, the image data tagged with an object identification value can be used to generate first summary data.
[0073] For example, the electronic device (1000) can obtain an output value representing the meaning of text data by inputting text data into an artificial intelligence model trained for natural language interpretation (e.g., a natural language understanding (NLU) model, a large language model (LLM)). In this case, the text data input into the artificial intelligence model trained for natural language interpretation may include logging data generated in text format and / or text data generated from voice data. The text data generated from the voice data may be generated from the voice data using, for example, automatic speech recognition (ASR) technology and / or speech-to-text (STT) technology. According to one embodiment, summarized text data may be obtained from the text data input into the artificial intelligence model trained for natural language interpretation. According to one embodiment, the voice data may be analyzed using an audio analysis technology (e.g., ASR technology) and used to distinguish speakers, identify the emotions of the speakers, and / or understand the content of the conversation.
[0074] For example, the electronic device (1000) may input logging data into a trained multi-modal artificial intelligence model to acquire characteristics of media data. In this case, for example, the electronic device (1000) may input logging data including at least text data, voice data, and / or image data into the multi-modal artificial intelligence model, thereby identifying characteristics of the text data, voice data, and / or image data.
[0075] According to one embodiment, the electronic device (1000) may obtain first summary data of logging data by generating a first input prompt based on an output value of an artificial intelligence model and inputting the generated first input prompt into a first artificial intelligence model. In this case, the electronic device (1000) may generate a first input prompt to be input into the first artificial intelligence model, and input the first input prompt and the analysis result of the logging data into the first artificial intelligence model. The first input prompt may include, for example, information on attributes of the first summary data to be generated by the first artificial intelligence model. The attributes of the first summary data may include preset information related to the type and number of media data to be included in the first summary data, a period corresponding to the first summary data, and a method of generating the first summary data. In this case, for example, the first artificial intelligence model may be a generative artificial intelligence model trained to generate summary data based on the analysis result of the logging data. For example, the analysis result of the logging data may include features of the logging data (e.g., objects identified from images, input intent and content of text, input intent and content of voice, etc.). For example, the first artificial intelligence model may be an artificial intelligence model trained to generate summary data classified according to predetermined categories.
[0076] According to one embodiment, the first summary data may be classified into a plurality of categories. The first summary data obtained from the first artificial intelligence model may be summary data classified into a plurality of categories. For example, the first summary data output from the first artificial intelligence model may include classification values for corresponding categories. The classified categories of the first summary data may be determined based on, for example, the type and / or identification value of an object in an image included in the logging data, the meaning of voice data and / or text data, but are not limited thereto. In addition, the plurality of categories for the first summary data may include, but are not limited to, at least one of a person, a place, a time, an emotion, stress, health, or a conversation topic.
[0077] According to one embodiment, the electronic device (1000) may generate the first summary data when a specific event occurs. In this case, the specific event may include, but is not limited to, an event in which the location of the electronic device (1000) changes by a predetermined threshold or more, an event in which the electronic device (1000) remains at a predetermined location for a predetermined threshold or more, an event in which logging data exceeding a predetermined threshold is acquired for a predetermined period of time, an event in which a conversation partner changes, an event in which a predetermined application or function is executed within the electronic device (1000), and / or an event in which sensing data detected through a sensor of the electronic device (1000) changes by a threshold or more.
[0078] According to one embodiment, the electronic device (1000) may determine the generation time of the first summary data based on the conversation content between the user and the other party. For example, the electronic device (1000) may determine a situation in which the user's conversation partner has significantly changed as the generation time of the first summary data. For example, if the user and A finish their conversation and start talking to B, the time when the conversation partner changes from A to B may be determined as the summary time. Alternatively, for example, the electronic device (1000) may detect a conversation indicating a farewell during a conversation between the user and A and determine the summary time. For example, if B joins a conversation between the user and A, the electronic device (1000) may determine that the entire conversation has ended and determine the time when B joins as the summary time. The electronic device may distinguish between the conversation between the user and A and the conversation between the user and A and B to generate the first summary data.
[0079] Additionally, the importance of the first summary data may be determined, for example, based on criteria specified by the user or based on AI judgment. For example, the order in which the first summary data is displayed, the highlighting of the first summary data, and / or the priority of the first summary data used when generating the second summary data may be determined based on the importance.
[0080] A specific example of how an electronic device (1000) according to one embodiment generates first summary data will be described in more detail later in FIG. 3. In addition, an example of first summary data classified according to a plurality of categories according to one embodiment will be described in more detail later in FIG. 5.
[0081] Although the above description describes that the electronic device (1000) acquires logging data and generates first summary data based on the acquired logging data, the present invention is not limited thereto. For example, the electronic device (1000) may collect user logging data from another electronic device. In addition, for example, the electronic device (1000) may also receive first summary data generated by a first artificial intelligence model stored in another electronic device (not shown) and / or a server (not shown) from another electronic device (not shown) and / or a server (not shown). For example, the first summary data classified into the person category called "Younghee" may include at least a portion of audio data of "Younghee" and / or video data of "Younghee" collected and created through a smartwatch. The video data may include photos and videos taken by a mobile device or photos and videos collected by another device with a camera. The first summary data of the person category called "Younghee" may be classified by type and by the device from which it was collected. Additionally, each first summary data may include an identifier of the device that created the first summary data.
[0082] In operation 230, the electronic device (1000) may receive a user input requesting the generation of second summary data. According to one embodiment, the user input requesting the generation of the second summary data may include a natural language input from the user requesting the generation of summary data related to the user's logging. For example, the user input requesting the generation of the second summary data may include a voice input and / or a text input from the user. For example, the user input requesting the generation of the second summary data may include an input directly requesting the generation of the second summary data and / or an input implicitly requesting the generation of the second summary data. For example, the user input requesting the generation of the second summary data may include an input including natural language such as, “Summarize the conversation I had with Younghee on the morning of February 4th,” “What did we do on the way to our recent travel destination?” and “What did Jiwoo like to do in Central Park last time?”
[0083] In operation 240, the electronic device (1000) can select at least one first summary data related to the user input from among a plurality of first summary data.
[0084] According to one embodiment, the electronic device (1000) can identify a category related to a user input among categories of first summary data, and select first summary data corresponding to the identified category among a plurality of first summary data.
[0085] According to one embodiment, the electronic device (1000) may interpret a user input requesting generation of second summary data, and select at least one category related to the user input from among a plurality of categories of first summary data based on the interpretation of the user input.
[0086] According to one embodiment, when a user input requesting the generation of second summary data is a natural language input, the electronic device (1000) may obtain information indicating the meaning of the user input using an artificial intelligence model for natural language interpretation. For example, the electronic device (1000) may input the user input into an NLU (natural language understanding) model and obtain an intent and parameters output from the NLU model. The intent is information determined by interpreting text using the NLU model and may indicate the user's input intention. The parameters may include words related to the details intended by the user. The parameters are information related to the intent, and a single intent may correspond to multiple parameters. For example, the electronic device (1000) may input the user input into an LLM (large language model) model and obtain output data related to the meaning of the text output from the LLM model. The output data output from the LLM model may include, for example, natural language representing the meaning of the text and / or category values corresponding to the meaning of the text, but is not limited thereto. Additionally, for example, the LLM model may be an artificial intelligence model stored within the electronic device (1000) or an artificial intelligence model stored on a server (not shown).
[0087] According to one embodiment, when information indicating the meaning of a user input is obtained using an artificial intelligence model for natural language interpretation, the electronic device (1000) may select a category corresponding to the meaning of the user input from among a plurality of categories of first summary data. For example, the electronic device (1000) may select a category corresponding to at least one of an intent or parameter output from the NLU model. In addition, the electronic device (1000) may select first summary data corresponding to the selected category. For example, the electronic device (1000) may select at least one category related to the meaning of the text based on output data related to the meaning of the text output from the LLM model, and select first summary data corresponding to the selected category.
[0088] At operation 250, the electronic device (1000) may generate a second input prompt for generating second summary data.
[0089] According to one embodiment, a second input prompt for generating second summary data may be generated based on a user input and selected first summary data, so that second summary data may be generated. The electronic device (1000) may generate a second input prompt to be input to a second artificial intelligence model trained for generating second summary data, so that the second summary data may be generated based on at least one of the user input or information indicating the meaning of the user input. The second artificial intelligence model may be an artificial intelligence model trained for generating the second summary data based on the first summary data. For example, the second artificial intelligence model may be a generative artificial intelligence model for generating the second summary data based on the first summary data selected according to the user input. For example, the second artificial intelligence model may be the same as or different from the first artificial intelligence model. For example, when the second artificial intelligence model is the same as the first artificial intelligence model, the electronic device (1000) may obtain the first summary data and / or the second summary data using one artificial intelligence model. Additionally, for example, when the electronic device (1000) generates first summary data and second summary data, the electronic device (1000) can use one artificial intelligence model to generate the first summary data and the second summary data.
[0090] According to one embodiment, the second input prompt may be an input prompt input to the second artificial intelligence model. If the second artificial intelligence model is a generative artificial intelligence model, the second input prompt may include at least text in natural language. For example, the second input prompt may include information regarding how the second artificial intelligence model generates the second summary data, and may include elements for determining the style, topic, and / or length of the second summary data generated by the second artificial intelligence model.
[0091] According to one embodiment, the electronic device (1000) may generate a second input prompt using a template of an input prompt. For example, when the electronic device (1000) has been staying at one place and then moved, the template “Summarize the content of the video or audio data collected from <place name> and display <tag>” may be used to generate a second input prompt “Summarize the content of the video or audio data collected from Central Park and display location information.” In this case, for example, the second input prompt may be generated by inputting “Central Park” as <place name> in the template and “location information” as <tag> in the template.
[0092] According to one embodiment, the electronic device (1000) may generate a second input prompt based on a user input requesting generation of second summary data. For example, based on a user input such as “What did Jiwoo do at Central Park last time?”, the electronic device (1000) may select a template such as “Summarize what activities were performed in the summary data of <period> classified as <person name> in <place name> among the summary data” and generate a second input prompt such as “Summarize what activities were performed in the summary data classified as Jiwoo in Central Park up to one month ago among the summary data” based on the selected template. In this case, for example, the second input prompt may be generated by inputting “Central Park” as <place name> in the template, “Jiwoo” as <person name> in the template, and “one month ago” as <period> in the template.
[0093] In operation 260, the electronic device (1000) may obtain second summary data by applying a second input prompt to a second artificial intelligence model. According to one embodiment, the electronic device (1000) may input the second input prompt and the selected first summary data to the second artificial intelligence model and obtain second summary data output from the second artificial intelligence model. Examples of second summary data obtained from the second artificial intelligence model will be described in more detail with reference to FIGS. 12 to 14 .
[0094] FIG. 3 is a flowchart of a method for an electronic device according to one embodiment to obtain a plurality of first summary data classified according to a plurality of categories by analyzing a plurality of logging data.
[0095] Actions 310 to 360 of FIG. 3 may correspond to action 220 of FIG. 2.
[0096] In operation 310, the electronic device (1000) may recognize objects in a plurality of images using an artificial intelligence model for object recognition. According to one embodiment, the electronic device (1000) may obtain image data tagged with an object identification value by inputting logging data, which is image data (e.g., still image data and / or video data), into an artificial intelligence model (e.g., vision model) trained for object recognition. The artificial intelligence model trained for object recognition may include, but is not limited to, an artificial intelligence model based on convolutional neural networks (CNN) and an artificial intelligence model based on region-based convolution neural networks (R-CNN). In this case, in order to recognize objects in the image data, images of people stored in the electronic device (1000) and / or images of people registered in a server (not shown) by a user of the electronic device (1000) may be used. For example, the artificial intelligence model may obtain a user's name corresponding to an object identified in the image data by comparing an object identified in the image data with an image of a person.
[0097] In operation 320, the electronic device (1000) may interpret the meaning of a plurality of speech data using an artificial intelligence model for natural language interpretation. According to one embodiment, the electronic device (1000) may obtain an output value representing the meaning of the text data by inputting the text data into an artificial intelligence model trained for natural language interpretation. The artificial intelligence model for natural language interpretation may include, but is not limited to, a natural language understanding (NLU) model and a large language model (LLM) model, for example. In this case, the text data input into the artificial intelligence model trained for natural language interpretation may include text data generated from speech data. The text data generated from the speech data may be generated from the speech data using, for example, automatic speech recognition (ASR) technology and / or speech-to-text (STT) technology. According to one embodiment, for example, the artificial intelligence model trained for natural language interpretation may be text data summarized from the input text data. According to one embodiment, in operation 320, the electronic device (1000) may input logging data generated in text format into an artificial intelligence model trained for natural language interpretation.
[0098] In operation 330, the electronic device (1000) can recognize objects within a plurality of images and interpret the meaning of a plurality of voice data using a multi-modal artificial intelligence model. The electronic device (1000) can input logging data into a trained multi-modal artificial intelligence model to acquire features of media data. In this case, for example, the electronic device (1000) can identify features of text data, voice data, and / or image data by inputting logging data including at least text data, voice data, and / or image data into the multi-modal artificial intelligence model.
[0099] Although FIG. 3 describes that the electronic device (1000) performs operations 310, 320, and 330 together, the present invention is not limited thereto. The electronic device (1000) may perform at least one of operations 310, 320, or 330. For example, the electronic device (1000) may perform operations 310 and 320 and not perform operation 330. For example, the electronic device (1000) may perform operation 330 without performing operations 310 and 320.
[0100] In operation 340, the electronic device (1000) may obtain the user's biometric information. According to one embodiment, the electronic device (1000) may obtain the user's biometric information sensed by a sensor of the electronic device (1000). Additionally, the electronic device (1000) may receive the user's biometric information sensed by a sensor of another electronic device from another electronic device. For example, the user's biometric data may be data regarding at least one of heart rate, sleep pattern, or exercise amount, but is not limited thereto. According to one embodiment, the user's biometric information may also be obtained as the user's logging data.
[0101] Additionally, the electronic device (1000) may acquire sensing information regarding the surrounding environment of the electronic device (1000) carried by the user. The sensing information regarding the surrounding environment may include, but is not limited to, sensing data regarding at least one of ambient temperature, noise level, or lighting conditions. According to one embodiment, the sensing information regarding the surrounding environment of the electronic device (1000) may also be acquired as user logging data.
[0102] At operation 350, the electronic device (1000) may generate a first input prompt for generating first summary data.
[0103] In one embodiment, the first input prompt may include, for example, information regarding the properties of the first summary data to be generated by the first artificial intelligence model. The properties of the first summary data may include preset information related to the type and number of media data to be included in the first summary data, the period corresponding to the first summary data, and the method of generating the first summary data. In addition, the first input prompt may include, for example, elements for determining the style, subject matter, and / or length of the first summary data generated by the first artificial intelligence model.
[0104] According to one embodiment, the first input prompt may be generated according to criteria preset in the electronic device (1000). The electronic device (1000) may generate the first input prompt in real time or select at least one of a plurality of pre-generated first input prompts.
[0105] According to one embodiment, the electronic device (1000) may generate or select a first input prompt to generate first summary data according to a preset cycle. For example, to generate first summary data every hour, the first input prompt "Summarize the user's logging data by category for the last hour" may be generated or selected.
[0106] According to one embodiment, the electronic device (1000) may generate or select a first input prompt to generate first summary data when a specific event occurs. The specific event may include, but is not limited to, an event in which the location of the electronic device (1000) changes by a predetermined threshold or more, an event in which a predetermined application or function is executed within the electronic device (1000), and / or an event in which sensed data detected through a sensor of the electronic device (1000) changes by a threshold or more. In this case, according to one embodiment, the electronic device (1000) may generate a first input prompt according to the specific event that has occurred. For example, when an event in which the location of the electronic device (1000) changes by a predetermined threshold or more occurs, the electronic device (1000) may generate a first input prompt saying, “Generate first summary data based on recent logging data related to location movement.” For example, when an event occurs in which a predetermined application or function is executed within the electronic device (1000), the electronic device (1000) may generate a first input prompt saying, “Generate first summary data based on logging data related to the executed application.” For example, when an event occurs in which sensing data detected through a sensor of the electronic device (1000) changes by a threshold or more, the electronic device (1000) may generate a first input prompt saying, “Generate first summary data related to a change in recent sensing data.”
[0107] In action 360, the electronic device (1000) can input a first input prompt to the first artificial intelligence model.
[0108] According to one embodiment, the electronic device (1000) may input a generated or selected first input prompt into the first artificial intelligence model. In addition, the electronic device (1000) may input the analysis result of logging data related to the first input prompt into the first artificial intelligence model together with the first input prompt. For example, the electronic device (1000) may input at least one of the object recognition result obtained in operation 310, the natural language interpretation result obtained in operation 320, the object recognition result and the natural language interpretation result obtained in operation 330, or the sensing information obtained in operation 340 into the first artificial intelligence model.
[0109] Although the above description describes inputting the analysis results of logging data into the first artificial intelligence model, the present invention is not limited thereto. If the first artificial intelligence model is an artificial intelligence model trained to generate first summary data based on various types of media data, the electronic device (1000) may not input the analysis results of the logging data obtained in operations 310, 320, and 330 into the first artificial intelligence model. In this case, the electronic device (1000) may input the logging data into the first artificial intelligence model along with the first input prompt.
[0110] FIG. 4 is a diagram illustrating an example of a first artificial intelligence model generating first summary data based on input values according to one embodiment.
[0111] Referring to FIG. 4, a first input prompt and logging data can be input to a first artificial intelligence model (40), and first summary data can be output from the first artificial intelligence model (40).
[0112] In one embodiment, logging data may be input to the first artificial intelligence model (40) along with a first input prompt. In this case, the logging data input to the first artificial intelligence model (40) may be data converted to have the format of input data used during training of the first artificial intelligence model (40).
[0113] In one embodiment, the analysis results of the logging data may be input to the first artificial intelligence model (40) along with the first input prompt. In this case, the analysis results of the logging data input to the first artificial intelligence model (40) may be data converted to have the format of the input data used during training of the first artificial intelligence model (40).
[0114] According to one embodiment, the calculation of the first artificial intelligence model (40) may be performed by at least one of a central processing unit (CPU), a graphical processing unit (GPU), or a neutral processing unit (NPU) of the electronic device (1000), but is not limited thereto. For example, the calculation of the first artificial intelligence model (40) may also be performed by another electronic device and / or a server. In this case, data for the calculation of the first artificial intelligence model (40) may be provided from the electronic device (1000) to the other electronic device and / or the server.
[0115] The first summary data output from the first artificial intelligence model (40) according to one embodiment may be classified by category. The first summary data classified by category will be described later in FIG. 5.
[0116] FIG. 5 is a diagram illustrating an example of first summary data classified according to a plurality of categories according to one embodiment.
[0117] Referring to FIG. 5, the first summary data generated according to one embodiment may be classified, for example, according to date / time, location, and person.
[0118] For example, from the new logging data 1 generated at the first time, the first summary data A1 of the place A category and the first summary data B1 of the person B category can be generated. For example, A1 may be a place category called Central Park, and B1 may be a person category called Jiwoo, who is the user's child. If the logging data indicates what Jiwoo did at Central Park, the first summary data A1 of the place A category and the first summary data B1 of the person B category can be generated.
[0119] For example, the first summary data A2 of the Place A category and the first summary data C1 of the Person C category can be generated from the new logging data 2 generated at the second time. For example, when the new logging data 2 is generated, the new logging data 2 can indicate the content of what happened with the user's friend Younghee in Central Park. In this case, if the first summary data A1 of the Place A category already exists, the first summary data A2 can be generated by updating the first summary data A1. Additionally, if Younghee is not included in the existing Person category, the first summary data C1 of the Person C category can be generated.
[0120] For example, from the new logging data 3 generated at the third time, the first summary data A3 of the place A category and the first summary data B2 of the person B category can be generated. For example, if the new logging data 3 is related to place A and person B, the first summary data A3 of the place A category can be generated by updating the first summary data A2, and the first summary data B2 of the person B category can be generated by updating the first summary data B1.
[0121] In FIG. 5, examples of first summary data classified by date / time, location, and person are described, but are not limited thereto.
[0122] FIG. 6 is a flowchart of a method for an electronic device to select first summary data related to a user input for generating second summary data according to one embodiment.
[0123] Actions 610 to 630 of FIG. 6 may correspond to action 240 of FIG. 2.
[0124] In operation 610, the electronic device (1000) may input text according to a user input into an artificial intelligence model for natural language interpretation. According to one embodiment, the electronic device (1000) may input text regarding a user input requesting the generation of second summary data into an artificial intelligence model for natural language interpretation. For example, the electronic device (1000) may input text input by the user and / or text converted from a voice input by the user into an NLU (natural language understanding) model for the generation of second summary data. For example, the electronic device (1000) may input text input by the user and / or text converted from a voice input by the user into an LLM (large language model) model for the generation of second summary data. The LLM model may be, for example, an artificial intelligence model stored within the electronic device (1000) or an artificial intelligence model stored in a server (not shown).
[0125] For example, if the user input requesting the generation of the second summary data is a voice input by the user into the electronic device (1000), the electronic device (1000) may obtain text from the voice data using automatic speech recognition (ASR) technology and / or speech to text (STT) technology, and input the obtained text into an artificial intelligence model trained for natural language interpretation. For example, if the user input requesting the generation of the second summary data is text, the electronic device (1000) may input the text input by the user into an artificial intelligence model trained for natural language interpretation.
[0126] In operation 620, the electronic device (1000) may obtain information regarding the intent of a user input output from an artificial intelligence model for natural language interpretation. According to one embodiment, the electronic device (1000) may obtain an intent and parameters output from the NLU model. The electronic device (1000) may obtain an intent and parameters indicating a request for generating second summary data. According to one embodiment, the intent is information determined by interpreting text using the NLU model and may indicate a user's input intent. The intent may include, for example, information indicating an intent for generating second summary data. The intent may include not only information indicating the user's input intent but also a numerical value corresponding to the information indicating the user's input intent. The numerical value may indicate a probability that the text is related to information indicating a specific intent. When multiple pieces of information indicating the user's intent are obtained as a result of interpreting the text using the NLU model, the intent information having the largest numerical value corresponding to each piece of intent information may be determined as the intent. In addition, for example, parameters may be variable information for determining detailed actions intended by a user in relation to an intent. The parameters may include words related to the detailed actions intended by the user. Parameters are information related to intents, and multiple parameters may correspond to a single intent. Parameters may include not only variable information for determining the detailed actions intended by the user, but also a numerical value indicating the probability that the text is related to the variable information. As a result of interpreting text using a natural language understanding model, multiple pieces of variable information representing parameters may be acquired. In this case, the variable information with the largest numerical value corresponding to each piece of variable information may be determined as the parameter.
[0127] For example, when the electronic device (1000) utilizes a large language model (LLM), the electronic device (1000) can obtain output data related to the meaning of text output from the LLM model. The output data output from the LLM model may include, for example, natural language representing the meaning of the text and / or category values corresponding to the meaning of the text, but is not limited thereto.
[0128] In operation 630, the electronic device (1000) may select at least one first summary data from among a plurality of first summary data based on information about the intent of the user input. According to one embodiment, the electronic device (1000) may select a category corresponding to the meaning of the user input requesting the generation of second summary data from among the plurality of categories of the first summary data. For example, the electronic device (1000) may select a category corresponding to at least one of an intent or parameter output from the NLU model. In this case, the category corresponding to at least one of the intent or parameter may be preset for the first summary data when the first summary data is generated. For example, the electronic device (1000) may select a category corresponding to an intent output from the NLU model. For example, the electronic device (1000) may select a category corresponding to an intent and a parameter output from the NLU model.
[0129] For example, when the electronic device (1000) uses a large language model (LLM), the electronic device (1000) can select a category based on output data output from the LLM model. For example, when natural language representing the meaning of text is output by the LLM model, the electronic device (1000) can select a category corresponding to the natural language representing the meaning of the text. For example, when a category value corresponding to the meaning of text is output from the LLM model, the electronic device (1000) can select a category corresponding to the category value output from the LLM model.
[0130] For example, if the user input for generating the second summary data is “What did you do with your child during the week?”, the electronic device (1000) can select the time category and the “child” category related to “during the week” related to the meaning of the user input.
[0131] In one embodiment, the electronic device (1000) may select first summary data corresponding to the selected category. For example, if the user input for generating second summary data is "What did you do with your child this week?", the electronic device (1000) may select first summary data related to the child from among the first summary data stored for the week, based on the time category related to "this week" and the "child" category.
[0132] FIG. 7 is a diagram illustrating an example in which first summary data related to user input is selected from first summary data classified according to a plurality of categories according to one embodiment.
[0133] Referring to FIG. 7, for example, first summary data for generating second summary data may be selected from among the first summary data classified by category in FIG. 6. For example, when a user input such as “Summarize what happened with B and C” is received, the electronic device (1000) may use an artificial intelligence model for natural language interpretation to understand the meaning of the user input and select the person B category and the person C category to generate the second summary data. In addition, for example, with respect to the person B category, the electronic device (1000) may select the first summary data B1, which is the updated summary data, from among the first summary data B1 and the first summary data B2. In addition, for example, with respect to the person C category, the electronic device (1000) may select the first summary data C1.
[0134] For example, the electronic device (1000) may obtain an intent called “content summary between people” and parameters called “person B” and “person C” from an artificial intelligence model, and select first summary data B1 of the person B category and first summary data C1 of the person C category to generate second summary data.
[0135] FIG. 8 is a diagram illustrating an example of a second artificial intelligence model generating second summary data based on input values according to one embodiment.
[0136] Referring to FIG. 8, a second input prompt and selected first summary data can be input into a second artificial intelligence model (80), and second summary data can be output from the second artificial intelligence model (80).
[0137] In one embodiment, the selected first summary data may be input into the second artificial intelligence model (80) along with a second input prompt. In this case, the first summary data input into the second artificial intelligence model (80) may be data having the format of the input data used during training of the second artificial intelligence model (80).
[0138] According to one embodiment, the calculation of the second artificial intelligence model (80) may be performed by at least one of a central processing unit (CPU), a graphical processing unit (GPU), or a neutral processing unit (NPU) of the electronic device (1000), but is not limited thereto. For example, the calculation of the second artificial intelligence model (80) may also be performed by another electronic device and / or a server. In this case, data for the calculation of the second artificial intelligence model (80) may be provided from the electronic device (1000) to the other electronic device and / or the server.
[0139] The second summary data generated by the second artificial intelligence model (80) according to one embodiment will be described in more detail in FIGS. 9 to 12 described below.
[0140] FIG. 9A is a diagram illustrating an example of first summary data according to one embodiment.
[0141] Referring to FIG. 9A, for example, the electronic device (1000) may generate first summary data (94) related to Younghee on the morning of February 4th. The electronic device (1000) may input logging data acquired on the morning of February 4th into a first artificial intelligence model along with a first input prompt, and may acquire first summary data output from the first artificial intelligence model. For example, the first summary data may include summarized voice data and text describing the summarized voice data. For example, the text corresponding to the summarized voice data may include the text, “At 9:10 a.m. on February 4th, Younghee and I exchanged “hello” at work and asked her if she had decided on attendees for the meeting next week. Younghee responded that she would decide by tomorrow.” In addition, for example, the summarized voice data may include voice data (e.g., audio data) of the logging data corresponding to the text. For example, when a thumbnail (94) representing summarized voice data is displayed on the screen of the electronic device (1000) and the thumbnail (94) is selected by the user, the summarized voice data can be output from the electronic device (1000). In this case, the summarized voice data may not be stored in advance in the electronic device (1000). For example, the electronic device (1000) may store only the text of the summarized voice data, and when the thumbnail (94) is selected by the user, the electronic device (1000) may convert the text of the summarized voice data into voice using TTS technology, thereby generating and outputting summarized voice data in real time. In this case, for example, the thumbnail (94) representing the summarized voice data may include information representing the subject of the summarized voice data (e.g., representative person, representative place, etc.). For example, the thumbnail representing the summarized voice data may be an image representing the subject of the summarized voice data (e.g., representative person, representative place, etc.), but is not limited thereto.For example, a topic related to summarized speech data may be identified from, but is not limited to, metadata of logging data corresponding to the summarized speech data and / or analysis results of the logging data.
[0142] Additionally, for example, the first summary data may correspond to the category “February 4th, AM” and the category “Younghee.”
[0143] FIG. 9b is a diagram illustrating an example of first summary data according to one embodiment.
[0144] Referring to FIG. 9B, for example, the electronic device (1000) may generate first summary data related to the morning of February 4th. The electronic device (1000) may input the logging data acquired on the morning of February 4th into the first artificial intelligence model along with the first input prompt, and may acquire the first summary data output from the first artificial intelligence model. For example, the first summary data may include summarized image data (104) and text describing the summarized image data (104). For example, the text describing the summarized image data (104) may include the text, “On February 4th at 11:40 AM, I ate salad at Subway with Younghee, Minhee, and Cheolsu. We talked about what we would eat on vacation next month, and Younghee said she wanted to eat the famous barbecue in the area, and Cheolsu said he wanted to drink a beverage made with local specialties. Minhee agreed with Cheolsu. After the meal, I took pictures of Younghee, Minhee, and Cheolsu.”
[0145] For example, summary data may be generated based on logging data regarding a function executed on the electronic device (1000) and / or the usage history of devices within the electronic device (1000). For example, based on information indicating that a photographing function was executed on the electronic device (1000), the summary data may include the content, “I took pictures of Younghee, Minhee, and Chulsoo.”
[0146] Additionally, for example, the summarized image data (104) may be images of Younghee, Minhee, and Cheolsu, but is not limited thereto. Furthermore, for example, the first summary data may correspond to the category of "February 4th, morning."
[0147] In addition, for example, the image (104) shown in FIG. 9B may be a representative image among the images that are the source of the summary data. For example, the electronic device (1000) may set at least one or more of the image data that are the source of the summary data as a representative image and display them on the screen, and when the representative image is selected by the user, the electronic device may display summarized image data including multiple images. For example, an image that includes the face of a person related to the summary data may be the representative image, but is not limited thereto. In addition, when there are multiple representative images, the multiple representative images may be displayed alternately. In addition, for example, the multiple representative images may be displayed together in multiple divided image display areas.
[0148] According to one embodiment, the electronic device (1000) can input logging data acquired on the morning of February 4th into the first artificial intelligence model along with the first input prompt, and can acquire the first summary data of FIG. 9A and the first summary data of FIG. 9B from the first artificial intelligence model, respectively.
[0149] FIG. 9c is a diagram illustrating an example of first summary data selected based on user input and second summary data generated based on the selected first summary data according to one embodiment.
[0150] Referring to FIG. 9C, for example, the electronic device (1000) may generate first summary data related to a trip. The electronic device (1000) may input logging data related to the trip into a first artificial intelligence model along with a first input prompt, and may obtain first summary data output from the first artificial intelligence model. For example, the first summary data may include summarized image data (114) and text describing the summarized image data (114). For example, the text describing the summarized image data may include the text, “We drove to Gyeongpodae at 9:50 AM on February 4th. Younghee said the sunset was beautiful, and you also talked about the beautiful scenery, saying that the red color resembled a rose.” For example, the summarized image data (114) may be an image of a sunset taken at Gyeongpodae. Additionally, for example, the first summary data may correspond to the category of “travel.”
[0151] For example, summary data may be generated based on logging data regarding a function executed in the electronic device (1000) and / or a usage history of a device within the electronic device (1000). In addition, for example, the image (114) shown in FIG. 9C may be a representative image among the images that are the source of the summary data. For example, the electronic device (1000) may set at least one or more of the image data that are the source of the summary data as a representative image and display them on the screen, and when the representative image is selected by the user, summarized image data including multiple images may be displayed. For example, an image including a face of a person related to the summary data may be the representative image, but is not limited thereto. In addition, when there are multiple representative images, the multiple representative images may be displayed alternately. In addition, for example, the multiple representative images may be displayed together in multiple divided image display areas.
[0152] According to one embodiment, the electronic device (1000) can input logging data acquired on the morning of February 4th into the first artificial intelligence model along with the first input prompt, and can acquire the first summary data of FIG. 9A, the first summary data of FIG. 9B, and the first summary data of FIG. 9C from the first artificial intelligence model, respectively.
[0153] FIG. 10 is a diagram illustrating an example of second summary data generated based on first summary data from a plurality of electronic devices according to one embodiment.
[0154] Referring to FIG. 10 , the electronic device (1000) may generate second summary data using a second artificial intelligence model based on a user input (120). For example, the electronic device (1000) may receive a user input (120) such as, “Make a video of the activities the children enjoyed during this picnic.” and interpret the meaning of the user input (120) using an artificial intelligence model for natural language interpretation. Furthermore, the electronic device (1000) may generate a second input prompt based on the user input (120) and preset criteria. For example, the electronic device may generate a second input prompt such as, “Location: Central Park. Collect possible summary data from connected devices. If no devices are connected, collect data after connecting. Data containing images that appear fun, identifiable human laughter, or mentions of fun. Briefly summarize the contents of the above data.”
[0155] Thereafter, for example, the electronic device (1000) may select first summary data related to the category of “theme park” and the category of “today” from among the first summary data of the electronic device (1000) and the first summary data of at least one other electronic device (e.g., the user’s wearable electronic device and / or the other user’s electronic device). In this case, for example, the electronic device (1000) may request first summary data related to the category of a place unit and first summary data related to the category of a date unit from at least one other electronic device, and receive the requested first summary data from the at least one other electronic device. In addition, the electronic device (1000) may select first summary data related to both the category of “theme park” and the category of “today” from among the received first summary data.
[0156] Alternatively, for example, the electronic device (1000) may request first summary data of the “Theme Park” category and first summary data of the “Today” category from at least one other electronic device, and receive first summary data of the “Theme Park” category and first summary data of the “Today” category from at least one other electronic device. In addition, the electronic device (1000) may select first summary data related to both “Theme Park” and “Today” from among the first summary data of the “Theme Park” category and the first summary data of the “Today” category.
[0157] Alternatively, for example, the electronic device (1000) may request first summary data related to both the “Theme Park” category and the “Today” category from at least one other electronic device, and receive first summary data related to both the “Theme Park” category and the “Today” category from at least one other electronic device.
[0158] Additionally, the electronic device (1000) can input the selected first summary data and the generated second input prompt into the second artificial intelligence model and obtain the second summary data output from the second artificial intelligence model.
[0159] In one embodiment, for example, the other electronic device may be a wearable electronic device and / or a mobile device worn by another student during a school field trip, and the first summary data of the other electronic device may be first summary data generated by acquiring audio data and / or image data from the other electronic device and summarizing the same. In this case, the user's electronic device (1000) may be a teacher's device, and the electronic device (1000) may collect the summary data from the student's electronic device. In addition, the electronic device (1000) may further summarize the first summary data collected from the other electronic device and use the summarized data to generate second summary data. In this case, the electronic device (1000) may generate the second summary data by considering the characteristics (e.g., audio / video / image) of the data collected from the other electronic device.
[0160] In one embodiment, for example, the second summary data may include summary data (122) based on first summary data generated in the electronic device (1000), summary data (123) based on first summary data generated in another electronic device of the first student, and summary data (124) based on first summary data generated in another electronic device of the second student.
[0161] In one embodiment, for example, a wearable device worn by a student may transmit only summary data that has been previously set as shareable to the teacher's electronic device (1000). For example, the wearable device worn by a student may set restrictions on sharing summary data, such as sharing only content that occurred at a specific time or location, or sharing only content with a specific person.
[0162] FIG. 11 is a diagram illustrating an example of second summary data generated according to user input related to a specific situation according to one embodiment.
[0163] Referring to FIG. 11, the electronic device (1000) may generate second summary data using a second artificial intelligence model based on a user input (130) related to a specific situation. For example, the electronic device (1000) may receive a user input (130) including a request related to the emotions of a specific person, such as “What did Jiwoo like to do in Central Park last time?”, and may interpret the meaning of the user input (130) using an artificial intelligence model for natural language interpretation. In addition, the electronic device (1000) may generate a second input prompt based on the user input (130) and a preset criterion. For example, the second input prompt may be generated as follows: “The location is Central Park, and among the summary data, data classified as Jiwoo. Briefly summarize the content of the data that includes a happy-looking video and the sound of Jiwoo laughing. If there are less than three pieces of data, request data from another wearable device and then summarize it.”
[0164] Thereafter, for example, the electronic device (1000) may select first summary data related to the category of “Central Park” and the category of “Eraser” from among the first summary data of the electronic device (1000) and the first summary data of at least one other electronic device (e.g., the user’s wearable electronic device and / or another user’s electronic device). In addition, the electronic device (1000) may input the selected first summary data and the generated second input prompt into a second artificial intelligence model, and obtain second summary data output from the second artificial intelligence model.
[0165] For example, the second summary data may include summary data (132) based on a photo taken by the electronic device (1000) and summary data (134) based on a photo taken by another electronic device. For example, the summary data (132) may include a photo taken by the user's mobile phone and text describing the content of the photo. Additionally, for example, the summary data (134) may include summary data based on logging data collected from a wearable device worn by Jiwoo.
[0166] In this case, the electronic device (1000) can sequentially output the summary data (132) and the summary data (134) within the second summary data. For example, the summary data (132) based on a photo taken by the electronic device (1000) may be displayed on the screen of the electronic device (1000), and then the summary data (132) based on the photo taken by the electronic device (1000) may be removed from the screen or the summary data (134) based on a photo taken by another electronic device may be displayed on the screen following the summary data (132). Accordingly, the image displayed on the electronic device (1000) may be changed to an image related to the summary data and displayed whenever the content of the summary data is changed. In addition, the text describing the photo taken by the electronic device (1000) may be changed to text describing the photo taken by another electronic device and displayed. For example, text within the second summary data may be gradually displayed, and images related to the displayed text may be changed and displayed. In this case, the method of summarizing the data may differ, for example, depending on the type of summary data and the properties of the device that generated the summary data. For example, the first summary data based on the user's mobile phone may include an image of Jiwoo and may also include Jiwoo's voice. Accordingly, in order to determine whether the first summary data based on the user's mobile phone is related to the user input (130), the electronic device (1000) may select the first summary data to be used for generating the second summary data based on whether the first summary data includes an image of Jiwoo smiling in a photo and whether the first summary data includes voice data in which Jiwoo's laughter is recorded. In addition, the electronic device (1000) may generate data (132) based on the selected first summary data among the first summary data based on the user's mobile phone.
[0167] Additionally, the electronic device (1000) may be used to determine whether the first summary data is related to Jiwoo's pleasant emotions, for example, by using biometric information collected from the electronic device (1000) when generating the first summary data.
[0168] Additionally, for example, the first summary data based on Jiwoo's wearable device may not include an image of Jiwoo. In this case, in order to determine whether the first summary data based on Jiwoo's wearable device is related to the user input (130), the electronic device (1000) may select the first summary data to be used for generating the second summary data based on whether the first summary data includes voice data recording Jiwoo's laughter. Additionally, the electronic device (1000) may generate data (134) based on the selected first summary data among the first summary data based on Jiwoo's wearable device.
[0169] FIG. 12 is a diagram illustrating an example of second summary data generated according to user input according to one embodiment.
[0170] Referring to FIG. 12, the electronic device (1000) may generate second summary data using a second artificial intelligence model based on a user input (140). For example, the electronic device (1000) may receive a user input (140) such as “What did Younghee do yesterday?” and interpret the meaning of the user input (140) using an artificial intelligence model for natural language interpretation. In addition, the electronic device (1000) may generate a second input prompt based on the user input (140) and a preset criterion. For example, the electronic device may generate a second input prompt such as “Time: February 4, 2024, data classified as Younghee among summary data, briefly summarize what happened by time zone. If there are less than three pieces of data, request data from another wearable device and then summarize.”
[0171] Thereafter, for example, the electronic device (1000) may select, from among the first summary data, first summary data related to the category of “February 4, 2024” and the category of “Younghee.” In addition, the electronic device (1000) may input the selected summary data and the generated second input prompt into the second artificial intelligence model, and obtain second summary data output from the second artificial intelligence model.
[0172] For example, the second summary data may include data related to Younghee on February 4, 2024. For example, the second summary data may include data (142) representing a conversation with Younghee at work at 9:10 a.m. on February 4. The data (142) may include a thumbnail representing the summarized voice data and text describing the summarized voice data. In this case, for example, the thumbnail representing the summarized voice data may include information related to the topic of the summarized voice data (e.g., representative person, representative place, etc.). For example, the thumbnail representing the summarized voice data may be, but is not limited to, an image representing the topic of the summarized voice data (e.g., representative person, representative place, etc.). In addition, for example, the topic of the summarized voice data may be identified from, but is not limited to, metadata of logging data corresponding to the summarized voice data and / or an analysis result of the logging data. Also, for example, the text corresponding to the summarized voice data may include the text, “I greeted Younghee at work yesterday morning at 9:10 and briefly talked about the meeting to be held next week.” For example, a thumbnail (94) representing the summarized voice data may be displayed on the screen of the electronic device (1000), and when the thumbnail (94) is selected by the user, the summarized voice data may be output from the electronic device (1000). In this case, the summarized voice data may not be stored in advance in the electronic device (1000). For example, the electronic device (1000) may store only the text of the summarized voice data, and when the thumbnail (94) is selected by the user, the text of the summarized voice data may be converted into voice using TTS technology, thereby generating and outputting the summarized voice data in real time.
[0173] For example, the second summary data may include data (143) describing what happened with Younghee, Minhee, and Cheolsu on the morning of February 4th. For example, data (143) may include images of Younghee, Minhee, and Cheolsu and text describing the images. For example, the text describing the images may include the text, "We briefly talked about the meeting next week. From 11:40 AM, we had a meal with Younghee, Minhee, and Cheolsu, talked about our vacation next month, and took pictures after the meal."
[0174] For example, the second summary data may include data (144) describing the activities related to the trip on February 4th. For example, data (144) may include a summarized video of a sunset shot at Gyeongpodae and text describing the summarized video. For example, the text describing the summarized video may include the text, "We met at Gyeongpodae at 9:50 PM and had a great time talking about the beautiful scenery together."
[0175] According to one embodiment, data (142), data (143), and data (144) included in the second summary data may be sequentially output from the electronic device (1000). For example, data (142), data (143), and data (144) may be output in the order of corresponding times, but is not limited thereto. For example, the output order of data (142, 143, 144) may be determined based on the number of logging data related to data (142, 143, 144) and / or the degree to which logging data related to data (142, 143, 144) is related to the user of the electronic device (1000) (e.g., the time the user was together with another user). In addition, the output order of data (142, 143, 144) may be determined based on, for example, a user request or a user input.
[0176] For example, data (143) and data (144) included in the second summary data may be displayed continuously. For example, text may be displayed by continuously scrolling in response to a scroll input. Furthermore, for example, an image may change to a different image depending on the degree of scrolling. In this case, the timing of the image change may be determined based on the extent to which the text corresponding to the image is displayed on the screen.
[0177] FIG. 13 is a diagram showing an example of first summary data displayed on a screen of an electronic device when the electronic device is a watch-type electronic device according to one embodiment.
[0178] Referring to FIG. 13, an electronic device (1000) according to one embodiment may be a watch-type electronic device. In this case, the electronic device (1000) may provide first summary data related to voice data. For example, the first summary data of FIG. 13 may be summary data generated by the electronic device (1000), which is a watch-type electronic device, but is not limited thereto.
[0179] For example, when thumbnails (13-1, 13-2, 13-3, 13-4) representing first summary data classified by person are displayed on the screen of the electronic device (1000), and one of the thumbnails (13-1, 13-2, 13-3, 13-4) is selected by the user, voice data may be output from the electronic device (1000) based on the first summary data corresponding to the selected thumbnail. In this case, the voice data to be output may not be stored in advance in the electronic device (1000). For example, the electronic device (1000) may store the first summary data as text, and when a thumbnail is selected by the user, the electronic device (1000) may convert the text of the first summary data corresponding to the selected thumbnail into voice using TTS technology, thereby generating and outputting voice data in real time.
[0180] For example, if the electronic device (1000) is a watch-type electronic device, the electronic device (1000) may generate first summary data based solely on audio data and display a thumbnail of the first summary data related to the audio data, but is not limited thereto. For example, if the electronic device (1000) is a watch-type electronic device, the electronic device (1000) may generate first summary data in the form of text based solely on audio data and display a thumbnail indicating that the first summary data in the form of text is based on audio data.
[0181] For example, even if the electronic device (1000) is a watch-type electronic device, if the resolution of the screen of the electronic device (1000) is higher than a preset value, the electronic device (1000) may provide first summary data including image data.
[0182] FIG. 14 is a diagram illustrating an example of displaying various types of summary data on a screen of an electronic device according to one embodiment.
[0183] Referring to FIG. 14, an electronic device (1000) according to one embodiment may be an electronic device that provides a display having a predetermined resolution or higher. In this case, the electronic device (1000) may provide multiple types of first summary data by categorizing them. For example, the first summary data of FIG. 14 may be summary data generated by the electronic device (1000), but is not limited thereto.
[0184] According to one embodiment, the electronic device (1000) may display thumbnails (14-1, 14-2, 14-3, 14-4, 14-5) representing first summary data classified by person on a list of summary data. In addition, when one of the thumbnails (14-1, 14-2, 14-3, 14-4, 14-5) in the list is selected by the user, the electronic device (1000) may output first summary data corresponding to the selected thumbnail. For example, the thumbnails (14-1, 14-2, 14-3, 14-4, 14-5) of the first summary data may include thumbnails (14-1, 14-2) of the first summary data including image data and thumbnails (14-3, 14-4, 14-5) of the first summary data related only to audio data. For example, the thumbnails (14-1, 14-2) of the first summary data including image data may be representative images including the face of a specific person, and the user of the electronic device (1000) may intuitively identify who the specific person is through the representative image. In addition, for example, the image data included in the first summary data may include, but is not limited to, still images and / or video images. For example, the first summary data may include text related to the image data (e.g., text including a description related to the image).
[0185] According to one embodiment, when one of the thumbnails (14-3, 14-4, 14-5) of the first summary data related only to audio data is selected from among the thumbnails (14-1, 14-2, 14-3, 14-4, 14-5) representing the first summary data classified by person, the electronic device (1000) may generate and output voice data in real time by converting the text corresponding to the selected thumbnail into voice using TTS technology. In this case, the voice data may not be stored in advance in the electronic device (1000). For example, the electronic device (1000) may store only the text of the voice data as the first summary data.
[0186] According to one embodiment, the electronic device (1000) may display thumbnails (14-8, 14-9) representing first summary data classified by location on a list of summary data, and when one of the thumbnails (14-8, 14-9) is selected by a user, the electronic device may output first summary data corresponding to the selected thumbnail. For example, the thumbnails (14-8, 14-9) of the first summary data may include thumbnails of the first summary data including image data. In this case, the thumbnails (14-8, 14-9) of the first summary data including image data may include, for example, an image representing a location and / or an image generated by superimposing other photos of the location by time zone, but is not limited thereto. In addition, for example, the image data included in the first summary data may include, but is not limited to, a still image and / or a video image.
[0187] FIG. 15 is a diagram illustrating an example of first summary data stored in an electronic device and another electronic device according to one embodiment.
[0188] Referring to FIG. 15, the first summary data used to generate the second summary data may be stored in the electronic device (1000) and other electronic devices (e.g., 2001, 2002, 2003).
[0189] In one embodiment, the electronic device (1000) and the other electronic device (2001) may be the user's devices. For example, the electronic device (1000) may be the user's mobile phone, and the other electronic device (2001) may be the user's wearable electronic device. For example, the other electronic device (2001) may include, but is not limited to, a watch-type electronic device, an HMD (head mounted device) electronic device, a ring-type electronic device, and an earphone device.
[0190] In one embodiment, the other electronic device (2002) and the other electronic device (2003) may be another user's electronic device (1000). For example, the other electronic device (2002) may be another user's mobile phone, and the other electronic device (2003) may be another user's wearable electronic device. For example, the other electronic device (2003) may include, but is not limited to, a watch-type electronic device, a head-mounted device (HMD) electronic device, a ring-type electronic device, and an earphone device.
[0191] According to one embodiment, the other electronic devices (2001, 2002, 2003) can each generate and store first summary data. For example, the other electronic devices (2001, 2002, 2003) can obtain logging data using a camera, a microphone, and / or various sensors of the other electronic devices (2001, 2002, 2003), and can each generate and store first summary data based on the logging data. In addition, the first summary data stored in the other electronic devices (2001, 2002, 2003) can be provided to the electronic device (1000), and the first summary data provided to the electronic device (1000) can be used by the electronic device (1000) to generate second summary data.
[0192] According to one embodiment, the electronic device (1000) can store first summary data generated by the electronic device (1000) in a first summary data DB (1930-1). The electronic device (1000) can receive first summary data generated by other electronic devices (2001, 2002, 2003) from other electronic devices (2001, 2002, 2003) and store the first summary data DB (1930-1).
[0193] According to one embodiment, the electronic device (1000) can store second summary data generated based on first summary data stored in the first summary data DB (1930-1) in the second summary data DB (1930-2).
[0194] Although the above description describes a user and other users' electronic devices (2001, 2002, 2003) providing the first summary data to the electronic device (1000), this is not limited. For example, the first summary data may be generated by a server (not shown) and provided to the electronic device (1000).
[0195] FIG. 16 is a flowchart of a method for generating second summary data by an electronic device in conjunction with another electronic device according to one embodiment.
[0196] Referring to FIG. 16, the first summary data and the second summary data may be generated by another electronic device (2000). The other electronic device (2000) may be another electronic device of the user, another electronic device of another user, or a server.
[0197] In operation 1600, the electronic device (1000) may obtain a plurality of logging data including a plurality of images and a plurality of voice data. Operation 1600 may correspond to operation 210 of FIG. 2.
[0198] In operation 1605, the electronic device (1000) may provide logging data to another electronic device (2000). The electronic device (1000) may transmit the logging data to the other electronic device (2000) while requesting generation of first summary data. The electronic device (1000) may request generation of the first summary data from the other electronic device (2000) when a specific event occurs. In this case, the specific event may include, but is not limited to, for example, an event in which the location of the electronic device (1000) changes by a predetermined threshold or more, an event in which the electronic device (1000) remains at a predetermined location by a predetermined threshold or more, an event in which logging data exceeding a predetermined threshold is acquired for a predetermined period of time, an event in which a conversation partner changes, an event in which a predetermined application or function is executed within the electronic device (1000), and / or an event in which sensing data detected by a sensor of the electronic device (1000) changes by a threshold or more.
[0199] According to one embodiment, the electronic device (1000) can identify whether another electronic device (2000) has permission to use the logging data, and if the other electronic device (2000) has permission to use the logging data, can request generation of first summary data while providing the logging data to the other electronic device (2000).
[0200] According to one embodiment, the electronic device (1000) may determine the time of requesting the first summary data based on the conversation between the user and the other party. For example, the electronic device (1000) may determine that a situation in which the user's conversation partner has significantly changed is the time to request the first summary data. For example, if the user and A finish their conversation and are now talking to B, the time when the conversation partner changes from A to B may be determined as the time of request. Alternatively, for example, the electronic device (1000) may detect a conversation indicating a farewell during a conversation between the user and A and determine the time of request. For example, if B joins a conversation between the user and A, the electronic device (1000) may determine that the entire conversation has ended and determine the time when B joins as the time of request.
[0201] In operation 1610, another electronic device (2000) may obtain a plurality of first summary data classified according to a plurality of categories by analyzing a plurality of logging data. Operation 1610 may correspond to operation 220 of FIG. 2.
[0202] In operation 1615, another electronic device (2000) may provide the generated plurality of first summary data to the electronic device (1000). The provided plurality of first summary data may be stored in the electronic device (1000).
[0203] In operation 1620, the electronic device (1000) may receive a user input requesting generation of second summary data, in operation 1625, the electronic device (1000) may select at least one first summary data related to the user input from among a plurality of first summary data, and in operation 1630, the electronic device (1000) may generate a second input prompt for generation of second summary data so that the second summary data may be generated based on the user input and the selected first summary data. Operation 1620 may correspond to operation 230 of FIG. 2 , operation 1625 may correspond to operation 240 of FIG. 2 , and operation 1630 may correspond to operation 250 of FIG. 2 .
[0204] In operation 1635, the electronic device (1000) may provide a second input prompt to another electronic device (2000). The electronic device (1000) may provide first summary data related to the second input prompt to the other electronic device (2000) along with the second input prompt. If the first summary data is already stored in the other electronic device (2000), the electronic device (1000) may provide an identification value of the first summary data to the other electronic device (2000).
[0205] In one embodiment, the electronic device (1000) may identify whether another electronic device (2000) has permission to use the first summary data, and if the other electronic device (2000) has permission to use the first summary data, may request generation of the second summary data while providing the first summary data and a second input prompt to the other electronic device (2000).
[0206] At operation 1640, another electronic device (2000) can obtain second summary data by applying a second input prompt to a second artificial intelligence model, and at operation 1645, another electronic device (2000) can provide the second summary data to the electronic device (1000).
[0207] In operation 1650, the electronic device (1000) can display second summary data. The electronic device (1000) can output the second summary data.
[0208] FIG. 17 is a flowchart of a method for generating second summary data by an electronic device in conjunction with another electronic device according to one embodiment.
[0209] Referring to FIG. 17, the first summary data and the second summary data may be generated by another electronic device (2000). The other electronic device (2000) may be another electronic device of the user, another electronic device of another user, or a server. Furthermore, in FIG. 17, unlike FIG. 16, the generation of the second input prompt may be performed by another electronic device (2000).
[0210] Actions 1700 to 1725 correspond to actions 1600 to 1625 of Fig. 16, so their description will be omitted for convenience.
[0211] In operation 1730, the electronic device (1000) may provide user input and at least one selected first summary data to another electronic device (2000).
[0212] In operation 1735, another electronic device (2000) may generate a second input prompt for generating second summary data based on the user input and the selected first summary data. Operation 1735 may correspond to operation 250 of FIG. 2.
[0213] At operation 1740, another electronic device (2000) can obtain second summary data by applying a second input prompt to a second artificial intelligence model, and at operation 1745, another electronic device (2000) can provide the second summary data to the electronic device (1000).
[0214] In operation 1750, the electronic device (1000) can display second summary data. The electronic device (1000) can output the second summary data.
[0215] Although FIGS. 16 and 17 describe another electronic device (2000) as generating the first summary data and the second summary data, the present invention is not limited thereto. For example, the electronic device (1000) may generate a portion of the first summary data and the second summary data, and the other electronic device (2000) may generate the remaining portion of the first summary data and the second summary data. For example, the electronic device (1000) may generate the first summary data, and the other electronic device (2000) may generate the second summary data. Alternatively, for example, the electronic device (1000) may generate the second summary data, and the other electronic device (2000) may generate the first summary data. In this case, the electronic device generating the second summary data may have higher-spec resources (e.g., CPU, GPU, memory capacity, battery capacity) than the electronic device generating the first summary data.
[0216] According to one embodiment, a method for generating summary data by an electronic device may include: obtaining a plurality of logging data including a plurality of images and a plurality of voice data by photographing a surrounding of the electronic device and recording a voice in the surrounding of the electronic device; obtaining a plurality of first summary data classified according to a plurality of categories by analyzing the plurality of logging data; receiving a user input requesting generation of second summary data; selecting at least one first summary data related to the user input from among the plurality of first summary data; generating a second input prompt for generation of the second summary data based on the user input and the selected at least one first summary data; and obtaining the second summary data by applying the second input prompt to a second artificial intelligence model.
[0217] According to one embodiment, the method further includes an operation of generating a first input prompt for summarizing the plurality of logging data including the plurality of images and the plurality of voice data; wherein the operation of obtaining the plurality of first summary data may generate a plurality of first summary data classified according to the plurality of categories by applying the first input prompt to a first artificial intelligence model trained for generating the first summary data.
[0218] According to one embodiment, the method further includes an operation of recognizing objects within the plurality of images; and an operation of interpreting the meaning of the plurality of voice data; wherein the operation of generating the first input prompt may generate the first input prompt to be input to the first artificial intelligence model based on the object recognition result and the meaning interpretation result.
[0219] According to one embodiment, the method further includes an operation of determining the plurality of categories by analyzing logging data including the plurality of images and the plurality of voice data; wherein information regarding at least one of the determined plurality of categories can be included in the first input prompt.
[0220] According to one embodiment, the method may further include the operation of updating at least some of the plurality of first summary data as new logging data including new image and new voice data is acquired.
[0221] In one embodiment, the plurality of categories may be classified based on at least one of person, place, time, emotion, stress, health, or conversation topic.
[0222] According to one embodiment, the plurality of logging data further includes biometric information of a user of the electronic device, and the plurality of categories can be classified based on the biometric information of the user.
[0223] According to one embodiment, the operation of selecting at least one first summary data related to the user input may include: an operation of interpreting an intent of the user input; and an operation of selecting the at least one first summary data related to the intent of the user input.
[0224] According to one embodiment, at least a portion of the plurality of first summary data may be received from another electronic device.
[0225] In one embodiment, the second artificial intelligence model may be a generative artificial intelligence model trained to generate the second summary data including text and images.
[0226] FIG. 18 is a block diagram of an electronic device (1801) within a network environment (1800) according to various embodiments. Referring to FIG. 18, in the network environment (1800), the electronic device (1801) may communicate with the electronic device (1802) via a first network (1898) (e.g., a short-range wireless communication network), or may communicate with the electronic device (1804) or a server (1808) via a second network (1899) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (1801) may communicate with the electronic device (1804) via the server (1808). According to one embodiment, the electronic device (1801) may include a processor (1820), a memory (1830), an input module (1850), an audio output module (1855), a display module (1860), an audio module (1870), a sensor module (1876), an interface (1877), a connection terminal (1878), a haptic module (1879), a camera module (1880), a power management module (1888), a battery (1889), a communication module (1890), a subscriber identification module (1896), or an antenna module (1897). In some embodiments, the electronic device (1801) may omit at least one of these components (e.g., the connection terminal (1878)), or may have one or more other components added. In some embodiments, some of these components (e.g., sensor module (1876), camera module (1880), or antenna module (1897)) may be integrated into a single component (e.g., display module (1860)).
[0227] The processor (1820) may control at least one other component (e.g., a hardware or software component) of the electronic device (1801) connected to the processor (1820) by executing, for example, software (e.g., a program (1840)), and may perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (1820) may store commands or data received from other components (e.g., a sensor module (1876) or a communication module (1890)) in a volatile memory (1832), process the commands or data stored in the volatile memory (1832), and store result data in a non-volatile memory (1834). According to one embodiment, the processor (1820) may include a main processor (1821) (e.g., a central processing unit or an application processor) or an auxiliary processor (1823) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1821). For example, when the electronic device (1801) includes the main processor (1821) and the auxiliary processor (1823), the auxiliary processor (1823) may be configured to use less power than the main processor (1821) or to be specialized for a given function. The auxiliary processor (1823) may be implemented separately from the main processor (1821) or as a part thereof.
[0228] The auxiliary processor (1823) may control at least a portion of functions or states associated with at least one component (e.g., the display module (1860), the sensor module (1876), or the communication module (1890)) of the electronic device (1801), for example, on behalf of the main processor (1821) while the main processor (1821) is in an inactive (e.g., sleep) state, or together with the main processor (1821) while the main processor (1821) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (1823) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (1880) or a communication module (1890)). In one embodiment, the auxiliary processor (1823) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (1801) where the artificial intelligence is performed, or can be performed through a separate server (e.g., server (1808)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0229] The memory (1830) can store various data used by at least one component (e.g., the processor (1820) or the sensor module (1876)) of the electronic device (1801). The data can include, for example, software (e.g., the program (1840)) and input data or output data for commands related thereto. The memory (1830) can include volatile memory (1832) or non-volatile memory (1834).
[0230] The program (1840) may be stored as software in memory (1830) and may include, for example, an operating system (1842), middleware (1844), or an application (1846).
[0231] The input module (1850) can receive commands or data to be used in a component of the electronic device (1801) (e.g., a processor (1820)) from an external source (e.g., a user) of the electronic device (1801). The input module (1850) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0232] The audio output module (1855) can output audio signals to the outside of the electronic device (1801). The audio output module (1855) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0233] The display module (1860) can visually provide information to an external party (e.g., a user) of the electronic device (1801). The display module (1860) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (1860) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0234] The audio module (1870) can convert sound into an electrical signal, or vice versa. According to one embodiment, the audio module (1870) can acquire sound through the input module (1850), output sound through the sound output module (1855), or an external electronic device (e.g., electronic device (1802)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1801).
[0235] The sensor module (1876) can detect the operating status (e.g., power or temperature) of the electronic device (1801) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (1876) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0236] The interface (1877) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1801) with an external electronic device (e.g., the electronic device (1802)). In one embodiment, the interface (1877) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0237] The connection terminal (1878) may include a connector through which the electronic device (1801) may be physically connected to an external electronic device (e.g., the electronic device (1802)). In one embodiment, the connection terminal (1878) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0238] The haptic module (1879) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (1879) may include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0239] The camera module (1880) can capture still images and moving images. In one embodiment, the camera module (1880) may include one or more lenses, image sensors, image signal processors, or flashes.
[0240] The power management module (1888) can manage the power supplied to the electronic device (1801). According to one embodiment, the power management module (1888) can be implemented as at least a part of, for example, a power management integrated circuit (PMIC).
[0241] A battery (1889) may power at least one component of the electronic device (1801). In one embodiment, the battery (1889) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0242] The communication module (1890) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1801) and an external electronic device (e.g., electronic device (1802), electronic device (1804), or server (1808)), and the performance of communication through the established communication channel. The communication module (1890) may operate independently from the processor (1820) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1890) may include a wireless communication module (1892) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1894) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (1804) via a first network (1898) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1899) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a local area network or a wide area network)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1892) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1896) to identify or authenticate the electronic device (1801) within a communication network such as the first network (1898) or the second network (1899).
[0243] The wireless communication module (1892) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimizing terminal power and connecting multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (1892) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (1892) may support various technologies for securing performance in high-frequency bands, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (1892) may support various requirements specified in the electronic device (1801), an external electronic device (e.g., the electronic device (1804)), or a network system (e.g., the second network (1899)). According to one embodiment, the wireless communication module (1892) can support a peak data rate (e.g., 20 Gbps or more) for eMBB implementation, a loss coverage (e.g., 164 dB or less) for mMTC implementation, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC implementation.
[0244] The antenna module (1897) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (1897) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (1897) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1898) or the second network (1899), may be selected from the plurality of antennas by, for example, the communication module (1890). A signal or power may be transmitted or received between the communication module (1890) and the external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (1897).
[0245] According to various embodiments, the antenna module (1897) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
[0246] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0247] According to one embodiment, commands or data may be transmitted or received between the electronic device (1801) and an external electronic device (1804) via a server (1808) connected to a second network (1899). Each of the external electronic devices (1802 or 1804) may be the same or a different type of device as the electronic device (1801). According to one embodiment, all or part of the operations executed in the electronic device (1801) may be executed in one or more of the external electronic devices (1802, 1804, or 1808). For example, when the electronic device (1801) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1801) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1801). The electronic device (1801) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (1801) may provide an ultra-low latency service using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (1804) may include an Internet of Things (IoT) device. The server (1808) may be an intelligent server utilizing machine learning and / or a neural network.According to one embodiment, an external electronic device (1804) or server (1808) may be included within the second network (1899). The electronic device (1801) may be applied to intelligent services (e.g., smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology and IoT-related technology.
[0248] An electronic device (1801) according to one embodiment may correspond to the electronic device (1000) of FIGS. 1 to 17, and the electronic device (1801) may perform the operations of the electronic device (1000) of FIGS. 1 to 17.
[0249] Figure 19 is a block diagram of an electronic device according to one embodiment.
[0250] The electronic device (1000) of FIG. 19 may correspond to the electronic device (1801) of FIG. 18.
[0251] Referring to FIG. 19, the electronic device (1000) may include an input module (1950), a camera module (1980), a display module (1960), a memory (1930), and a processor (1920).
[0252] The input module (1950) can receive commands or data to be used in a component of the electronic device (1000) (e.g., a processor (1920)) from an external source (e.g., a user) of the electronic device (1000). The input module (1950) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0253] The camera module (1980) can capture still images and moving images. In one embodiment, the camera module (1980) may include one or more lenses, image sensors, image signal processors, or flashes.
[0254] The display module (1960) can visually provide information to an external party (e.g., a user) of the electronic device (1000). The display module (1960) may include, for example, a display, a holographic device, or a projector, and a control circuit for controlling the device. In one embodiment, the display module (1960) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0255] The memory (1930) can store various data used by at least one component of the electronic device (1000) (e.g., the processor (1920) or a sensor module (not shown)). The data can include, for example, input data or output data for software and commands related thereto. The memory (1930) can include volatile memory or non-volatile memory.
[0256] According to one embodiment, the memory (1930) can store at least one first summary data (1930-1) and at least one second summary data (1930-2).
[0257] According to one embodiment, the memory (1930) may store at least a first artificial intelligence model (1930-3) and at least one second artificial intelligence model (1930-4), but is not limited thereto. The first artificial intelligence model (1930-3) and the second artificial intelligence model (1940) may be the same artificial intelligence model, in which case one artificial intelligence model may be stored in the memory (1930). In addition, although FIG. 19 illustrates that at least the first artificial intelligence model (1930-3) and at least one second artificial intelligence model (1930-4) are stored in the memory (1930), the present invention is not limited thereto. For example, at least the first artificial intelligence model (1930-3) and at least one second artificial intelligence model (1930-4) may be stored in an internal memory within a separate processor (e.g., an NPU) for computing the artificial intelligence models.
[0258] The processor (1920) may, for example, execute software to control at least one other component (e.g., hardware or software component) of the electronic device (1000) connected to the processor (1820) and perform various data processing or calculations. According to one embodiment, the processor (1920) may include a main processor (e.g., a central processing unit or an application processor) or an auxiliary processor (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that may operate independently or together therewith.
[0259] According to one embodiment, the processor (1920) can control the operations of the electronic device (1000) in FIGS. 1 to 17.
[0260] According to one embodiment, the processor (1920) can obtain a plurality of logging data including a plurality of images and a plurality of voice data.
[0261] According to one embodiment, the processor (1920) can obtain a plurality of first summary data classified according to a plurality of categories by analyzing a plurality of logging data.
[0262] According to one embodiment, the processor (1920) may receive user input requesting generation of second summary data.
[0263] According to one embodiment, the processor (1920) may select at least one first summary data related to the user input from among a plurality of first summary data.
[0264] According to one embodiment, the processor (1920) may generate a second input prompt for generating second summary data based on user input and the selected first summary data.
[0265] According to one embodiment, the processor (1920) can obtain second summary data by applying a second input prompt to a second artificial intelligence model.
[0266] FIG. 20 is a diagram illustrating a system including a generative artificial intelligence model according to one embodiment.
[0267] Referring to FIG. 20, the User Query / Response Interface (2310) can receive a user's input. The user's input may be in the form of natural language, images, and / or videos. Furthermore, context information may also be transmitted when the user's input is transmitted. Context information may include various additional information at the time of user input. For example, information on the application currently being used by the user or information on the user's location. Furthermore, the user's input may be in a mixed form of the aforementioned natural language, images, sounds, and context information. Furthermore, the user's input may also be in a non-natural language form, such as selecting a menu. The User Query / Response Interface (2310) can output the results of a generative artificial intelligence system to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of an action requested by the user. The User Query Interface can output the results of a generative artificial intelligence system to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of an action requested by the user.
[0268] The AI framework (2320) can receive user input and coordinate and control each component necessary to perform the user's intention based on the user's query.
[0269] User input received from the User Query / Response Interface (2310) can be transmitted to the Prompt design component (2321). The Prompt design component (2321) can be used to generate a prompt suitable for inputting the user input into a Large Language Model (LLM) or a large multimodal model (LMM). The Prompt design component (2321) can be an AI component that uses a machine learning algorithm or a neural network to develop better prompts over time. The Prompt design component (2321) can access a knowledge component (e.g., knowledge repositories (2340)) containing user preference data, a prompt library, and prompt examples based on the user input to generate a prompt, and transmit the generated prompt to the LLM or LMM.
[0270] The API / Plug-in management component (2323) can communicate with external information when there is a request for additional information when passing user input as input to a generative model. The API / Plug-in management component (2323) can establish a channel for communicating with the outside of the AI Interface through the API, and can enable access to various data sources (e.g., knowledge repositories (2340)) through the established channel. In addition, if the API / Plug-in management component (2323) needs to perform an action that performs the user input as a final result rather than an intermediate result in an application or service, it can request the action to the application / service component (2330) through the API. Information obtained from an external source can be used to generate a prompt in the prompt design component (2321) together with the user input, or can be passed as input to the generative model.
[0271] The Refiner component (e.g., the output modification component (2325)) can fine-tune the output from a generative model. For example, the Refiner component can verify that the content generated by the LLM and / or LMM is not irrelevant, biased, or harmful. Furthermore, the Refiner component can determine the degree to which the output matches the user's desired result and, if necessary, perform additional processing. The Refiner component can also configure and provide hints to the user to avoid undesirable output.
[0272] Generative AI Model (2350) can generally refer to an artificial intelligence neural network that creates new types of data based on user input information. Generative AI Model (2350) can include an image-generating model and / or a language-generating model. Representative models for generating images include a generative adversarial network (GAN) and a variational autoencoder (VAE), and examples include a diffusion-based generative model that uses a VAE and a transformer structure. A language-generating model is a model trained to statistically output the most appropriate output value based on input values, and representative examples include models such as CHAT-GPT 3 and CHAT-GPT 4. In addition, there is also an LMM that can recognize various types of data input, such as text, images, and voice, and generate new data corresponding to them.
[0273] According to one embodiment, the electronic device (1000) of FIGS. 1 to 19 may be configured to include at least a portion of the User Query / Response Interface (2310), the AI framework (2320), the application / service component (2330), the knowledge repositories (2340), or the Generative AI Model (2350) of FIG. 20. According to one embodiment, at least a portion of the User Query / Response Interface (2310), the AI framework (2320), the application / service component (2330), the knowledge repositories (2340), or the Generative AI Model (2350) of FIG. 20 may be included in another electronic device (e.g., another user's electronic device and / or a server).
[0274] According to one embodiment, an electronic device for generating summary data comprises: a camera; a microphone;
[0275] A memory storing commands; and one or more processors; wherein the commands, when executed by the one or more processors, cause the electronic device to: obtain a plurality of logging data including a plurality of images and a plurality of voice data by photographing the surroundings of the electronic device through the camera and recording a voice of the surroundings of the electronic device through the microphone, obtain a plurality of first summary data classified according to a plurality of categories by analyzing the plurality of logging data, receive a user input requesting generation of second summary data, select at least one first summary data related to the user input from among the plurality of first summary data, generate a second input prompt for generation of the second summary data based on the user input and the selected at least one first summary data, and obtain the second summary data by applying the second input prompt to a second artificial intelligence model.
[0276] According to one embodiment, the instructions, when executed by the one or more processors, may cause the electronic device to: generate a first input prompt for summarizing the plurality of logging data including the plurality of images and the plurality of voice data, and apply the first input prompt to a first artificial intelligence model trained for generation of the first summary data, thereby generating a plurality of first summary data classified according to the plurality of categories.
[0277] According to one embodiment, the instructions, when executed by the one or more processors, may cause the electronic device to: recognize objects within the plurality of images, interpret meaning of the plurality of voice data, and generate the first input prompt to be input to the first artificial intelligence model based on the object recognition result and the meaning interpretation result.
[0278] According to one embodiment, the instructions, when executed by the one or more processors, cause the electronic device to: determine the plurality of categories by analyzing logging data including the plurality of images and the plurality of voice data, wherein information regarding at least one of the determined plurality of categories can be included in the first input prompt.
[0279] According to one embodiment, the instructions, when executed by the one or more processors, may cause the electronic device to: update at least some of the plurality of first summary data as new logging data, including new image and new audio data, is acquired.
[0280] In one embodiment, the plurality of categories may be classified based on at least one of person, place, time, emotion, stress, health, or conversation topic.
[0281] According to one embodiment, the plurality of logging data further includes biometric information of a user of the electronic device, and the plurality of categories can be classified based on the biometric information of the user.
[0282] According to one embodiment, the instructions, when executed by the one or more processors, may cause the electronic device to: interpret the intent of the user input, and select the at least one first summary data related to the intent of the user input.
[0283] According to one embodiment, at least a portion of the plurality of first summary data may be received from another electronic device.
[0284] According to one embodiment, a computer-readable recording medium having recorded thereon a program for executing a method, the method comprising: obtaining a plurality of logging data including a plurality of images and a plurality of voice data by photographing the surroundings of an electronic device and recording voices in the surroundings of the electronic device; obtaining a plurality of first summary data classified according to a plurality of categories by analyzing the plurality of logging data; receiving a user input requesting generation of second summary data; selecting at least one first summary data related to the user input from among the plurality of first summary data; generating a second input prompt for generation of the second summary data based on the user input and the selected at least one first summary data; and obtaining the second summary data by applying the second input prompt to a second artificial intelligence model.
[0285] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0286] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0287] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0288] Various embodiments of the present document may be implemented as software (e.g., a program (1840)) including one or more instructions stored in a storage medium (e.g., an internal memory (1836) or an external memory (1838)) readable by a machine (e.g., an electronic device (1801)). For example, a processor (e.g., a processor (1820)) of the machine (e.g., an electronic device (1801)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0289] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0290] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In a method for an electronic device to generate summary data, An operation of obtaining a plurality of logging data including a plurality of audio data by recording sounds surrounding the electronic device; An operation of obtaining a plurality of first summary data classified according to a plurality of categories by analyzing the plurality of logging data; An action that receives user input requesting the generation of second summary data; An operation of selecting at least one first summary data related to the user input from among the plurality of first summary data; An operation of generating a second input prompt for generating the second summary data, such that the second summary data is generated based on the user input and the at least one selected first summary data; and An operation of obtaining the second summary data by applying the second input prompt to the artificial intelligence model; A method comprising:
2. In paragraph 1, An operation of generating a first input prompt for summarizing the plurality of logging data including the plurality of audio data; Including more, A method wherein the operation of obtaining the plurality of first summary data generates a plurality of first summary data classified according to the plurality of categories by applying the first input prompt to the artificial intelligence model trained for generating the first summary data.
3. In paragraph 2, The above plurality of logging data further includes a plurality of images obtained by photographing the surroundings of the electronic device, The above method, An operation of recognizing objects within the plurality of images; and An operation of interpreting the meaning of the plurality of audio data; Including more, A method wherein the operation of generating the first input prompt generates the first input prompt to be input to the artificial intelligence model based on the object recognition result and the semantic interpretation result.
4. In paragraph 2, An operation of determining the plurality of categories by analyzing logging data including the plurality of audio data; Including more, A method wherein information regarding at least one of the determined plurality of categories is included in the first input prompt.
5. In paragraph 1, An operation of updating at least some of the plurality of first summary data as new logging data including image data and / or audio data is acquired; A method further comprising:
6. In paragraph 1, A method wherein the artificial intelligence model is a generative artificial intelligence model trained to generate the second summary data including text and images.
7. In an electronic device that generates summary data, camera; mike; memory that stores commands; and One or more processors; Includes, The above instructions, when executed by the one or more processors, cause the electronic device to: By recording the surrounding sound of the electronic device through the microphone, a plurality of logging data including a plurality of audio data are obtained, By analyzing the above plurality of logging data, a plurality of first summary data classified according to a plurality of categories are obtained, Receive user input requesting the generation of second summary data, Among the plurality of first summary data, at least one first summary data related to the user input is selected, Generating a second input prompt for generating the second summary data based on the user input and the at least one selected first summary data, An electronic device that obtains the second summary data by applying the second input prompt to an artificial intelligence model.
8. In paragraph 7, The above instructions, when executed by the one or more processors, cause the electronic device to: Generate a first input prompt for summarizing the plurality of logging data including the plurality of images and the plurality of audio data, An electronic device that generates a plurality of first summary data classified according to the plurality of categories by applying the first input prompt to the artificial intelligence model trained for generating the first summary data.
9. In paragraph 8, The above plurality of logging data further includes a plurality of images obtained by photographing the surroundings of the electronic device, The above instructions, when executed by the one or more processors, cause the electronic device to: Recognize objects within the above multiple images, Interpret the meaning of the above multiple audio data, An electronic device that generates the first input prompt to be input into the artificial intelligence model based on the object recognition result and the semantic interpretation result.
10. In paragraph 8, The above instructions, when executed by the one or more processors, cause the electronic device to: By analyzing the logging data including the plurality of audio data, the plurality of categories are determined, An electronic device, wherein information regarding at least one of the determined plurality of categories is included in the first input prompt.
11. In paragraph 7, The above instructions, when executed by the one or more processors, cause the electronic device to: An electronic device configured to update at least some of said plurality of first summary data as new logging data including image and audio data is acquired.
12. In paragraph 7, An electronic device wherein the above categories are classified according to at least one of person, place, time, emotion, stress, health, or conversation topic.
13. In paragraph 7, The plurality of logging data further includes biometric information of the user of the electronic device, An electronic device wherein the above categories are classified based on the biometric information of the user.
14. In paragraph 7, The above instructions, when executed by the one or more processors, cause the electronic device to: Interpret the intent of the above user input, An electronic device that selects at least one first summary data related to the intent of the user input.
15. In paragraph 7, An electronic device, wherein at least a portion of the plurality of first summary data is received from another electronic device.
Citation Information
Patent Citations
Systems and methods for automatically recognizing and classifying daily behavior patterns using wearable sensors
KR102465318B1
Voice based realtime event logging
US20190043500A1