Method for processing utterance of user and electronic device supporting same

A multimodal generation model trained on personalized data addresses the lack of effective user speech processing in electronic devices, enabling devices to generate personalized multimedia content based on user images and captions, enhancing user interaction and response accuracy.

WO2026116979A1PCT designated stage Publication Date: 2026-06-04SAMSUNG ELECTRONICS CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2025-11-26
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing electronic devices lack effective methods for processing user speech and generating personalized multimedia content based on user images and captions, limiting their ability to provide relevant and engaging responses.

Method used

A multimodal generation model is trained using personalized data to map image frames and text captions, allowing the generation of related images or text based on user queries or captured images, utilizing a processor to select highlight frames and generate captions, and integrating this model into electronic devices for interactive responses.

Benefits of technology

Enables devices to provide personalized and relevant multimedia outputs, enhancing user interaction by leveraging user-specific data to generate images or text related to user experiences, improving engagement and response accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025019774_04062026_PF_FP_ABST
    Figure KR2025019774_04062026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device disclosed in the present document may comprise a memory, and at least one processor. The memory may store instructions that, when executed individually or collectively by the at least one processor, instruct the electronic device to: acquire a first video related to a user; select a highlight frame on the basis of a rank for each frame constituting the first video; generate a caption related to the highlight frame and at least one object included in the highlight frame; generate personalized data for the user by linking the highlight frame and the caption; and train a multi-modal generation model by using the personalized data. Other embodiments identified through the specification are possible.
Need to check novelty before this filing date? Find Prior Art

Description

Method for processing user speech and electronic device supporting the same

[0001] The embodiments disclosed in this document relate to a method for processing user speech using a multimodal generation model and an electronic device supporting the same.

[0002] Electronic devices (e.g., smartphones, tablets, desktops, or VST (virtual studio technology) devices) can store photos or videos (or video). Photos or videos can be stored with captions. Captions can be simple text containing a description of the photo or video.

[0003] When a user searches for a photo or video stored on an electronic device, the electronic device may display a photo or video corresponding to the user's input (e.g., touch input, voice input). For example, the electronic device may find a photo or video similar to a sentence based on the user's input, or generate and display a random image related to a word entered by the user.

[0004] An electronic device according to one embodiment may include a memory and at least one processor. The memory may store instructions that, when executed individually or collectively by the at least one processor, cause the electronic device to acquire a first image related to a user, select a highlight frame based on a frame-by-frame rank constituting the first image, generate a caption related to the highlight frame and at least one object included in the highlight frame, generate personalized data for the user by linking the highlight frame and the caption, and train a multimodal generation model using the personalized data.

[0005] A method for processing a user's speech according to one embodiment may be performed in an electronic device. The method may include the operation of acquiring a first image related to a user, the operation of selecting a highlight frame based on a rank per frame constituting the first image, the operation of generating a caption related to the highlight frame and at least one object included in the highlight frame, the operation of generating personalized data for the user by linking the highlight frame and the caption, and the operation of training a multimodal generation model using the personalized data.

[0006] A computer-readable storage medium according to one embodiment may store instructions executable by a processor. When the instructions are executed, the processor of an electronic device may perform the following operations: acquiring a first image related to a user; selecting a highlight frame based on a rank per frame constituting the first image; generating a caption related to the highlight frame and at least one object included in the highlight frame; generating personalized data for the user by linking the highlight frame and the caption; and training a multimodal generation model using the personalized data.

[0007] FIG. 1 is a block diagram of an electronic device in a network environment according to various embodiments.

[0008] FIG. 2 is a configuration diagram of a voice recognition system according to one embodiment.

[0009] FIG. 3a is a flowchart illustrating a learning method for a multimodal generation model according to one embodiment.

[0010] FIG. 3b is a flowchart illustrating a method of using a multimodal generation model according to one embodiment.

[0011] FIG. 4a illustrates the learning and use of a multimodal generation model according to one embodiment.

[0012] FIG. 4b shows the output of a multimodal generation model according to one embodiment.

[0013] FIG. 5a illustrates a learning method of a multimodal generation model according to one embodiment.

[0014] FIG. 5b illustrates the operation of a multimodal generation model according to one embodiment.

[0015] FIGS. 6a and 6b show a web search unit according to one embodiment.

[0016] FIG. 6c shows a data storage unit according to one embodiment.

[0017] FIG. 7 shows a video summary according to one embodiment.

[0018] FIGS. 8A, FIGS. 8B, and FIGS. 8C illustrate the operation of a streaming caption model according to one embodiment.

[0019] FIG. 9 shows a photo recognition model according to one embodiment.

[0020] FIG. 10 shows a user interface regarding the use of a multimodal generation model according to one embodiment.

[0021] FIG. 11 shows the output of a multimodal generation model corresponding to a user's utterance according to one embodiment.

[0022] FIG. 12 shows a user interface for using an image-based multimodal generation model according to one embodiment.

[0023] FIG. 13 shows the output of a multimodal generation model corresponding to a user's utterance related to a place according to one embodiment.

[0024] FIG. 14 shows the output of a multimodal generation model that provides sequential responses in response to a user's utterance according to one embodiment.

[0025] FIG. 15 shows the output of a multimodal generation model based on web crawling according to one embodiment.

[0026] FIG. 16 shows the output of a multimodal generation model in a foldable device according to one embodiment.

[0027] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.

[0028] Hereinafter, various embodiments of this document are described with reference to the accompanying drawings. However, this is not intended to limit the technology described in this document to specific embodiments and should be understood to include various modifications, equivalents, and / or alternatives to the embodiments of this document. In relation to the description of the drawings, similar reference numerals may be used for similar components.

[0029] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (or display) (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).

[0030] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), for example, and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.

[0031] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence is performed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0032] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).

[0033] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0034] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0035] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0036] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0037] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).

[0038] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0039] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0040] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0041] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0042] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0043] The power management module (188) can manage the power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0044] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0045] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).

[0046] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for realizing eMBB, loss coverage (e.g., 164 dB or less) for realizing mMTC, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for realizing URLLC.

[0047] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).

[0048] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0049] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.

[0050] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0051]

[0052] FIG. 2 is a configuration diagram of a voice recognition system according to one embodiment.

[0053] Referring to FIG. 2, the voice recognition system (201) may include a wearable device (e.g., smart glasses) (210) and an electronic device (220) (e.g., the electronic device (101) of FIG. 1). FIG. 2 illustrates a case where the wearable device (210) and the electronic device (220) are configured separately, but is not limited thereto. For example, the wearable device (210) and the electronic device (220) may be implemented as a single device.

[0054] A wearable device (e.g., smart glasses) (210) may be worn on a user's body. The wearable device (210) may include a microphone, a camera, or a button for user input. The wearable device may receive and process the user's voice through the microphone. The wearable device (210) may acquire images (photos or videos) using the camera. The wearable device (210) may start taking photos and videos in response to user input occurring on the button.

[0055] According to one embodiment, the wearable device (210) can capture images of an object or scene of interest to the user in accordance with user input or voice commands. For example, the wearable device (210) can start taking photos or videos through voice commands "take photo" or "take video". As another example, the wearable device (210) can proceed with taking photos when the user presses a designated button "once" and start taking videos when the user presses a designated button "twice". The wearable device (210) can transmit the captured video or photos to an electronic device (220).

[0056] According to one embodiment, the electronic device (220) may be a device connected to the wearable device (210) via a designated wireless communication. For example, the electronic device (220) may be a device such as a smartphone or a tablet PC. Some operations of the electronic device (220) may be performed through an external server (e.g., the server (108) of FIG. 1).

[0057] An electronic device (220) (e.g., the electronic device (101) of FIG. 1) can receive video or photos from a wearable device (210). The electronic device (220) can receive video or photos and obtain personalized data specific to the user.

[0058] According to one embodiment, the electronic device (220) may include a video summarizing unit (230), a caption generating unit (240), a photo processing unit (250), a web search unit (255), a personalized data storage unit (260), and a multimodal generation model (270). Each component of the electronic device (220) in FIG. 2 is classified according to the function it performs, but is not limited thereto. In the following, the operation of the video summarizing unit (230), the caption generating unit (240), the photo processing unit (250), and the multimodal generation model (270) may be an operation of the processor of the electronic device (220) (e.g., the processor (120) in FIG. 1) or an operation performed by the execution of an algorithm or program in the processor. The data storage unit (350) may be part of a memory (e.g., the memory (130) in FIG. 1). Each component of the electronic device (220) in FIG. 2 may be integrated or separated into sub-components.

[0059] According to one embodiment, the video summary unit (230) and the caption generation unit (240) can extract personalized data from a video (video, streaming video) received from a wearable device (210). The video summary unit (230) and the caption generation unit (240) may be referred to as a video highlight model.

[0060] The video highlight model can select video frames according to rank from a video (video, streaming video) received from a wearable device (210). The video highlight model can store caption information for the selected frames along with the selected frames. For example, if a man walking down the street and eating a sandwich is filmed, the video highlight model can generate a caption consisting of two sentences: "A man walks into a shop. The man is eating a sandwich."

[0061] According to one embodiment, the video summarizing unit (230) can find highlight sections by selecting input video frames according to importance rank. The video summarizing unit (230) can perform video summarization after selecting frames according to rank.

[0062] According to one embodiment, the video summarizing unit (230) can select frames such that the duration (t) of the selected frames is maintained at less than a first reference value (e.g., 15%) for the total video time (T). If the duration (t) of the selected frames is less than a second reference value (e.g., 5%) for the total video time (T), the video summarizing unit (230) can determine that the frames are not important video frames and delete the selected frames. Additional information regarding the operation of the video summarizing unit (230) may be provided through FIG. 7.

[0063] According to one embodiment, the caption generation unit (240) can generate a caption through a frame that is input via streaming during video recording. The caption may be simple text describing the video.

[0064] According to one embodiment, the caption generation unit (240) can generate detailed captions in stages 1 through 3. In the first stage, the caption generation unit (240) can generate captions in a first time unit (hereinafter, clip unit). A clip unit may have a frame length of approximately 10 seconds or less. In a clip unit, a caption including words or a brief description may be generated. In the second stage, the caption generation unit (240) can generate captions in a second time unit (hereinafter, segment unit). A segment unit may have a frame length between approximately 10 seconds and approximately 10 minutes. In the third stage, the caption generation unit (240) can generate captions in a third time unit (hereinafter, whole video unit). A whole video unit may have a frame length between approximately 10 minutes and several hours.

[0065] According to one embodiment, the caption generation unit (240) can recognize a user's speech that occurs during the process of acquiring an image and convert it into text. The caption generation unit (240) can convert the user's speech or the conversation of people around it into text and reflect it in generating a caption. For example, the caption generation unit (240) can generate a caption by inputting it into a language model (LM) using the caption information generated by the first to third steps and a script transcribing the user's speech. Additional information regarding the operation of the caption generation unit (240) may be provided through FIGS. 8a to 8c.

[0066] The photo processing unit (250) can process a photo taken from a wearable device (210). The photo processing unit (250) can extract multiple objects from a single image. The photo processing unit (250) can separate parts corresponding to each object from a single image through an image element split model and create new images for each separated object. For example, when "a cat and a dog are overlapping each other in a park" within a single image, the photo processing unit (250) can create three images for "park", "cat", and "dog".

[0067] According to one embodiment, the web search unit (255) can perform web crawling related to photos and selected frames. The web search unit (255) can read text or similar images related to photos and selected frames. Through this, the reliability of the personalized data can be increased.

[0068] According to one embodiment, the web search unit (255) can label more accurate information through user feedback on the stored data to increase the reliability of the data. Additional information regarding the web search unit (255) may be provided through FIGS. 6a and FIGS. 6b.

[0069] According to one embodiment, the data storage unit (260) can store photos and selected frames together with captions. The photos and selected frames and the captions describing them can be stored as paired data.

[0070] According to one embodiment, the multimodal generation model (270) can be trained with personalized data. If the personalized data stored in the data storage unit (260) is of a certain data size or larger (e.g., user data is about 1 TB or larger), the multimodal generation model (270) can be trained using the personalized data. Training can be performed by fine-tuning the pre-train model to adapt to the user's personalized data. The multimodal generation model (270) can map image frames (selected frames) and text (captions) to each other in a vector space dimension and perform alignment to form a correlation between the two features.

[0071] According to one embodiment, the multimodal generation model (270) can perform learning using personalized data when the electronic device (220) is charging or when the user does not use the electronic device (220) for more than a specified amount of time.

[0072] According to one embodiment, the multimodal generation model (270) can generate related images or text by taking previously experienced information regarding a user's question (query). The multimodal generation model (270) can provide the generated images or text to the user visually or provide them as speech through text-to-speech (TTS). Additional information regarding the learning and operation of the multimodal generation model (270) may be provided through FIGS. 5A and FIGS. 5B.

[0073]

[0074] FIG. 3a is a flowchart illustrating a learning method for a multimodal generation model according to one embodiment.

[0075] Referring to FIGS. 1 to 3a, in operation 311, the processor (120) can acquire a first image related to the user. The first image may be an image captured through a wearable device (210).

[0076] According to one embodiment, the first image may be a photo or video (or video) taken by a user wearing a wearable device (210) in daily life. For example, the first image may be an image taken while performing a life logging function. Alternatively, the first image may be a photo or video taken by a user of an area of ​​interest (e.g., scene, object) while moving or in daily life.

[0077] In operation 313, according to one embodiment, the processor (120) may select highlight frames based on the frame-by-frame rank constituting the first image. The frame-by-frame rank may refer to an index that indicates a high probability of being selected as a highlight frame. For example, for a video of a person playing basketball, the rank of an image frame containing a person shooting may be determined to be higher than that of an image frame containing a person in a normal dribbling state. The processor (120) may summarize the first image by combining the image frames with high ranks.

[0078] According to one embodiment, the processor (120) can summarize the first image using various algorithms that compare and analyze feature points of frames constituting the first image. For example, the video summarizing unit (230) can extract features of image frames (feature extraction), learn frame importance (multi-head attention), learn frame location information (position-wise feed forward), and convert into probabilities (softmax) to label ranks. Additional information regarding frame selection may be provided through FIG. 7.

[0079] In operation 315, according to one embodiment, the processor (120) can generate a caption related to a highlight frame and at least one object included in the highlight frame. The caption may be simple text describing the video. The processor (120) can generate captions for the video by dividing them into clip units, segment units, and whole video units. If the voice of a user or people nearby is generated during the process of acquiring the first video, the processor (120) can generate a caption for the first video by reflecting the voice.

[0080] In operation 317, according to one embodiment, the processor (120) may generate personalized data for a user by linking a highlight frame and a caption. The personalized data may be paired data in which the highlight frame and the caption corresponding to the highlight frame are paired with each other.

[0081] According to one embodiment, the processor (120) can generate personalized data by generating stepwise captions through user feedback and obtaining highlight frames according to rank.

[0082] In operation 319, according to one embodiment, the processor (120) can train a multimodal generation model using personalized data. If the stored personalized data is larger than a certain data size (e.g., user data is about 1 TB or more), the processor (120) can train a multimodal generation model (270) using the personalized data.

[0083] According to one embodiment, the processor (120) can remove personalized data after training the multimodal generation model (270). By doing so, the processor (120) can secure storage space by deleting unnecessary information regarding logging.

[0084]

[0085] FIG. 3b is a flowchart illustrating a method of using a multimodal generation model according to one embodiment.

[0086] Referring to FIGS. 1 to 3b, in operation 321, according to one embodiment, the processor (120) can train a multimodal generation model using personalized data. The processor (120) can enable the multimodal generation model (270) to map highlight frames (selected frames) and captions to each other in a vector space dimension and perform alignment by forming a correlation between the two features.

[0087] In operation 323, according to one embodiment, the processor (120) may receive a second image captured through a camera of a wearable device worn by the user or a speech input from the user. The second image may be an image captured through a wearable device (210) worn by the user. The speech input from the user may be a question (query) from the user that occurs together with or separately from the second image.

[0088] In operation 325, according to one embodiment, the processor (120) may generate a third image or a first text in a multimodal generation model (270) based on a second image or a user's speech input. The third image or the first text may be generated based on information previously experienced by the user (e.g., captured images, speech content).

[0089] In operation 327, according to one embodiment, the processor (120) may output at least one of a third image or a first text. The processor (120) may provide the third image or the first text to the user visually or provide it as speech through text-to-speech (TTS).

[0090]

[0091] FIG. 4a illustrates the learning and use of a multimodal generation model according to one embodiment.

[0092] Referring to FIG. 2 and FIG. 4a, according to one embodiment, a user (405) may wear a wearable device (210). The wearable device (210) may capture an image (photo or video) (215) through a camera. The wearable device (210) may acquire the user's (405) speech (406) through a microphone. The wearable device (210) may transmit the image (215) and speech (406) to an electronic device (220).

[0093] According to one embodiment, the electronic device (220) can generate video summaries and captions and store personalized data in a personal data storage unit (260). The electronic device (220) can incorporate the utterance (406) of the user (405) into the caption generation. The electronic device (220) can generate and provide information related to the user's previous experiences using a multimodal generation model (270).

[0094] For example, a wearable device (210) or an electronic device (220) may provide a response (450) to a user's query (430) through a voice recognition application. Additionally, the wearable device (210) or the electronic device (220) may provide an image (455) generated in relation to the user's previous experience.

[0095]

[0096] FIG. 4b shows the output of a multimodal generation model according to one embodiment.

[0097] Referring to FIG. 2 and FIG. 4b, the multimodal generative model (270) can generate and provide information related to the user's previous memories or experiences.

[0098] According to one embodiment, the multimodal generation model (270) can generate one output (image or text) for one input (image or text). For example, in response to a user's query (voice input) (451), the multimodal generation model (270) may provide an image (461) generated by the multimodal generation model (270) in relation to the user's previous memory or experience. In another example, in response to an image (452) being captured through a camera, the multimodal generation model (270) may provide text (462) generated by the multimodal generation model (270) in relation to the user's previous memory or experience. The text (462) may include a description of a background and object that the user has previously seen when the user sees the background and object while walking along a path.

[0099] According to one embodiment, the multimodal generation model (270) can generate multiple outputs (images or text) for a single input (image or text). For example, in response to a user's query (voice input) (453), the multimodal generation model (270) can provide an image (463a) and text (463b) generated by the multimodal generation model (270) in relation to the user's previous memory or experience. In another example, in response to an image (454) being captured through a camera, the multimodal generation model (270) can provide text (464a) and an image (464b) generated by the multimodal generation model (270) in relation to the user's previous memory or experience.

[0100] Although not shown in FIG. 4b, the multimodal generation model (270) can generate multiple outputs (images or text) for multiple inputs (images or text).

[0101]

[0102] FIG. 5a illustrates a learning method of a multimodal generation model according to one embodiment.

[0103] Referring to Figures 2 and 5a, the Discriminative model (501) learns various data and can classify data with similar information through learning.

[0104] A multimodal generative model (270) can categorize data (selected frames and captions) obtained through the user's experience within the model. The multimodal generative model (270) can perform clustering within the model while learning pair data by utilizing the user's previous personal data as multi-input and multi-output. The multimodal generative model (270) can take clustered elements (520) as output and output an image (521) or text (522) mapped to the clustered elements (520). The multimodal generative model (270) can provide the search information the user wants more quickly by utilizing the user's previous experience.

[0105]

[0106] FIG. 5b illustrates the operation of a multimodal generation model according to one embodiment.

[0107] Referring to FIG. 5b, the multimodal generation model (270) can provide multi-input and multi-output. When user personalized data (paired data - image / text) (530) is stored, the multimodal generation model (270) can perform cross alignment by matching the correct answer to image to text and text to image (image to text, text to image). The multimodal generation model (270) may include encoders (541, 542) and decoders (561, 562).

[0108] The encoder (541, 542) may have an attention layer (542) of a CNN model (541) for images and an RNN model for text. The feature representation that passes through the encoder (541, 542) can be converted into an image (532a) or text (531a) by passing through the decoder (561, 562). The multimodal generation model (270) can perform a process of making it similar to the correct answer by checking the loss of the input and output values.

[0109] The core unit (550) can perform alignment by fusing text frames and image frames. The multimodal generation model (270) can map important relationships between image frames and text frames through learning.

[0110] According to one embodiment, when the training of the multimodal generation model (270) using personalized data is completed, the personalized data may be removed.

[0111] According to one embodiment, the output process of the response in the multimodal generation model (270) may be similar to the learning process. When text (user query) (581) or an image (currently captured image frame) (582) is input, text (581a) or an image (582a) may be generated.

[0112]

[0113] FIGS. 6a and 6b show a web search unit according to one embodiment.

[0114] Referring to FIG. 2 and FIG. 6a, the web search unit (255) can perform web crawling related to photos and selected frames. The web search unit (255) may include an encoder (611) and a decoder (612). The web search unit (255) can generate similar images (631, 632) that more clearly represent objects included in the photos and selected frames (620). The web search unit (255) can add information (641) related to the similar images (631, 632) through web crawling (640). Through this, the reliability of the personalized data can be increased.

[0115] Referring to FIG. 2 and FIG. 6b, the web search unit (255) can label more accurate information through user feedback on stored data to increase the reliability of the data. For example, the web search unit (255) may include a first interface (681) that allows the user to select an image or a second interface (682) that allows the user to select text.

[0116]

[0117] FIG. 6c shows the storage of pair data in a data storage unit according to one embodiment.

[0118] Referring to FIG. 2 and FIG. 6c, the data storage unit (260) can store highlight frames (652, 662, 672) together with captions (651, 661, 671). The captions (651, 661, 671) may be text describing the highlight frames (652, 662, 672) and the objects contained in the highlight frames (652, 662, 672). The highlight frames (652, 662, 672) and the captions (651, 661, 671) may be stored as paired data. When the multimodal generation model (270) is trained, it may be trained together in units of paired data.

[0119]

[0120] FIG. 7 shows a video summary according to one embodiment. FIG. 7 is exemplary and is not limited thereto.

[0121] Referring to FIGS. 2 and FIGS. 7, the video summarizing unit (230) may receive a video frame (or image frame) (710) from a wearable device (210). The video summarizing unit (230) may select the video frame (710) according to rank. The rank per frame may refer to an index that indicates a high probability of being selected as a highlight frame.

[0122] The video summary unit (230) can predict the rank of each video frame (710) using the video rank summary model (720). The video rank summary model (720) can generate a frame (730) with the rank labeled for each video frame (710).

[0123] According to one embodiment, the video summary unit (230) can extract features of the video frame (710), learn frame importance (multi-head attention), learn frame position information (position-wise feed forward), and convert it into a probability (softmax) to label the rank.

[0124] According to one embodiment, the video summary unit (230) may label important moments of visually consistent segments with ranks. For example, frames (731, 735) containing general scenery may be labeled as a relatively low rank, such as rank 3 or 4. Alternatively, frames (732) featuring houses or commercial buildings, frames (733) featuring famous tourist attractions, and frames (734) featuring people may be labeled as a relatively high rank, such as rank 1 or 2.

[0125] According to one embodiment, the video summarizing unit (230) can summarize video frames (710) by performing a Rank sorting strategy. In the first step, the video summarizing unit (230) can merge frames of the same rank (740). In the second step, the video summarizing unit (230) can select a specified rank (e.g., Rank 1, Rank 2) among the frames merged with the same rank (750). In the third step, the video summarizing unit (230) can ensure that frames selected by segment according to rank are less than or equal to a specified time. For example, the video summarizing unit (230) can select two adjacent segments, select only one segment, or ensure that adjacent segments are not selected.

[0126] According to one embodiment, the video summarizing unit (230) can select frames such that the duration (t) of the selected frames is maintained at less than a first reference value (e.g., 15%) for the total video time (T). If the duration (t) of the selected frames is less than the first reference value (e.g., 15%), the video summarizing unit (230) can add frames of a lower rank so that the duration (t) exceeds the first reference value (e.g., 15%).

[0127] According to one embodiment, the video summary unit (230) may determine that the selected frame is not an important video frame and delete the selected frame if the duration (t) of the selected frame is less than a second reference value (e.g., 5%) with respect to the total video time (T). The video summary unit (230) may save the selected frame if the duration (t) of the selected frame is greater than or equal to the second reference value (e.g., 5%) with respect to the total video time (T).

[0128]

[0129] FIGS. 8a, 8b, and 8c illustrate the operation of a streaming caption model according to one embodiment. FIGS. 8a, 8b, and 8c are exemplary and are not limited thereto.

[0130] FIG. 8a shows a streaming caption model according to one embodiment, and FIG. 8b may be a form that embodies the streaming caption model of FIG. 8a. FIG. 8c shows an example of a caption generated by a streaming caption model according to one embodiment.

[0131] Referring to FIG. 2 and FIG. 8a, the caption generation unit (240) can learn a streaming captioning model (870) in a recurrent manner. The input to the streaming captioning model (870) may be a clip, a segment, or a whole frame. A clip, a segment, or a whole frame may be a window size capable of holding image frames. For example, a clip may be a window size for holding frames between 0 and 10 seconds, a segment may be a window size for holding frames between 10 seconds and 1 minute, and a whole frame may be a window size that combines two or more segments.

[0132] For example, a streaming caption model (870) can read image frames in clip units and fuse the image and text (vision-text) using a vision-text attention alignment model. At this time, alignment between words in the vision frame and text can be performed. A text decoder generates text and can perform learning by computing a loss with key words or brief sentences for the ground truth clip.

[0133] Subsequently, when a segment-length image frame is input, the streaming caption model (870) can generate a vector by performing vision-text attention alignment on the segment-length image frame and the sentence summarized in clip units through the previous text decoder. The streaming caption model (870) can generate a new sentence by inputting the generated vector into the text decoder. For the segment-unit sentence, the streaming caption model (870) can perform a loss operation using a more detailed and longer sentence than a keyword or a brief sentence.

[0134] Finally, the streaming caption model (870) can input sentences for every segment into the LM (851) to generate caption sentences for the entire video.

[0135] Referring to FIG. 2 and FIG. 8b, the caption generation unit (230) can receive an image frame (812) that is streamed from a wearable device (210). The streaming frame may be a frame that is streamed during video recording. The image frame (812) may be received together with the user's speech (811).

[0136] The caption generation unit (230) can generate a caption for the streaming frame (812) according to the first to third steps.

[0137] In the first step (820), the caption generation unit (240) can generate a caption in a first time unit (clip unit). A clip unit may be a frame length of approximately 10 seconds or less. The caption generation unit (240) may utilize a learned vision-text attention alignment model. The vision-text attention alignment model may be a model that performs alignment between images and text by fusing them. The vision-text attention alignment model can output text that matches the image.

[0138] According to one embodiment, a caption containing words or brief descriptions may be generated at the clip unit. For example, in the first step (820), captions such as "road, sign, mountain," "many trees are placed at the foot of the mountain," and "there is a walking path around a beautiful lake" may be generated.

[0139] In the second step (830), based on the result of the first step (820), the caption generation unit (240) can generate a caption in a second time unit (segment unit). A segment unit may have a frame length between approximately 10 seconds and approximately 10 minutes. The caption generation unit (240) can input the caption text generated in each clip unit (clip_1, clip_2, ... clip_N) into a vision-text attention alignment model to fuse it with the video in the segment unit. The vision-text attention alignment model can output the caption text in the segment unit.

[0140] For example, in the second step (830), a caption such as “Let’s follow the road signs, and a beautiful lake appeared. We are walking along the path around the lake.” can be generated.

[0141] In the third step (840), based on the result of the second step (830), the caption generation unit (240) can generate a caption for a third time unit (entire video unit). The entire video unit may have a frame length between about 10 minutes and several hours. When video recording ends, a caption for the entire video unit can be generated.

[0142] For example, a caption corresponding to the entire video, such as “I followed the road signs, a beautiful lake appeared, and I walked along the path around the lake and met my friends again.”, can be generated. According to one embodiment, the caption generation unit (230) may receive the user’s speech along with image frames from the wearable device (210). The caption generation unit (230) may extract information contained in the user’s speech independently of the caption generation process using the image frames.

[0143] According to one embodiment, the caption generation unit (230) can search for meaningful speech segments that are not background noise using VAD (voice activity dictation). The caption generation unit (230) can distinguish between the user's speech and the speech of another speaker through speaker diarization. The caption generation unit (230) can output meaningful speech content between the user and the speaker as text (automatic speech recognition, ASR) using a speaker-specific stt (speech to text) model (860).

[0144] Referring to FIGS. 8b and 8c, the caption generation unit (240) can generate a single caption (852) by inputting a caption (813) for the entire video frame and text (861) transcribed from the user's speech into a language model (LM) (851). The caption (852) can be stored as pair data together with a summarized image frame (812a).

[0145]

[0146] FIG. 9 shows a photo recognition model according to one embodiment.

[0147] Referring to FIGS. 2 and FIGS. 9, the photo processing unit (250) can process a photo taken from a wearable device (210). The photo processing unit (250) can extract multiple objects from a single image. The photo processing unit (250) can separate parts corresponding to each object from a single image through an image element split model and create a new image for each separated object.

[0148] The image element segmentation model can separate backgrounds and objects using a clustering method. The image element segmentation model can automatically identify potential objects from random tokens in an input image (901) through a Self-Attention layer (910). The image element segmentation model can separate backgrounds and objects through pre-clustering (920) (902). The image element segmentation model can localize objects through cross-attention (930) and filtered masks (940). The image element segmentation model can separate each object through post-clustering (950) (903). Depending on the image, objects may overlap or only partially be visible. If objects overlap or only partially are visible, the image element segmentation model can regenerate objects (904a, 904b, 904c) using a Generator (960).

[0149] The photo processing unit (250) can obtain various information by generating images of multiple objects through a single input image (901).

[0150] According to one embodiment, an image element segmentation model can be trained to separate masked objects and calculate the loss of the actual object and the image of the actual object to display the actual image.

[0151]

[0152] FIG. 10 shows a user interface regarding the use of a multimodal generation model according to one embodiment.

[0153] Referring to FIGS. 2 and FIGS. 10, a multimodal generation model (270) can be learned by personalized data that reflects the experience of a user (405). A wearable device (210) can capture an image (215). The multimodal generation model (270) can provide information related to the user's previous experience corresponding to the current image (215).

[0154] For example, if the image (215) includes first to fourth objects (e.g., picture frame, plant, violin, flowerpot), the multimodal generation model (270) can generate and provide an image (1030) similar to an image previously taken by the user among the first to fourth objects (e.g., picture frame, plant, violin, flowerpot).

[0155] For example, the multimodal generation model (270) can identify pair data related to the user's previous experience corresponding to the first object (e.g., a picture frame). If the multimodal generation model (270) does not have personalized data (pair data) related to the user's previous experience, it can output a first generated image (1031) corresponding to the first object (e.g., a picture frame) and a description (1041) regarding the first generated image (1031) in a TTS manner through web crawling. The description (1041) may include information regarding the object included in the first generated image (1031) (e.g., price, store or shopping mall where it is sold).

[0156] For example, the multimodal generation model (270) can identify personalized data (pair data) related to the user's previous experience corresponding to a second object (e.g., sunflower) or a fourth object (e.g., flowerpot). If the multimodal generation model (270) has personalized data (pair data) related to the user's previous experience, it can use the image of the personalized data to generate a second generated image (1032) related to the second object (e.g., sunflower) or a fourth generated image (1034) related to the fourth object (e.g., flowerpot). There may be multiple second generated images (1032) or fourth generated images (1034). The second generated images (1032) or fourth generated images (1034) may be a video or frames that transition over time. The multimodal generation model (270) can output a description (1042) regarding the second generated image (1032) or a description (1043) regarding the fourth generated image (1034) in a TTS manner using the text of the pair data. The second generated image (1032) or the fourth generated image (1034) may be an image previously viewed by the user (or an image generated based on an image viewed by the user). The description (1042) or the description (1043) may include information regarding the location where an object identical or similar to the second object (e.g., sunflower) or the fourth object (e.g., flowerpot) was viewed, the schedule of the day it was viewed, or the weather.

[0157] For example, if the multimodal generation model (270) does not have personalized data (pair data) related to the user's previous experience or web crawl data, it may simply display a third generated image (1033) containing a third object (e.g., sunflower).

[0158]

[0159] FIG. 11 shows the output of a multimodal generation model corresponding to a user's utterance according to one embodiment.

[0160] Referring to FIGS. 2 and FIGS. 11, the wearable device (210) can receive a user's utterance (216) without a separate image input. The multimodal generation model (270) can provide information related to the user's previous experience corresponding to the user's utterance (216).

[0161] For example, the user's utterance (216) may be a query related to a memory previously experienced by the user. The multimodal generation model (270) may generate and display an image previously experienced by the user (an image viewed by the user while wearing the wearable device (210)) (1110) (hereinafter, experience image) based on the user's utterance (216). The experience image (1110) may be generated using an image of personalized data (pair data) related to the user's previous experience. The experience image (1110) may be an image containing an object (e.g., a figure) mentioned by the user at the place and time mentioned in the user's utterance (216).

[0162] According to one embodiment, individual information about an object included in an experience image (1110) may be provided. For example, if the experience image (1110) includes a first object (e.g., robot figure 1) and a second object (e.g., robot figure 2), the multimodal generation model (270) may output a first additional image (1120) and a first additional description (1121) for the first object (e.g., robot figure 1) and a second additional image (1130) and a second additional description (1131) for the second object (e.g., robot figure 2). The first additional image (1120), the first additional description (1121), the second additional image (1130), and the second additional description (1131) may be generated through personalized data (pair data) that the user has previously experienced and stored, or through web crawling. The first additional explanation (1121) and the second additional explanation (1131) can be output in a TTS manner.

[0163]

[0164] FIG. 12 shows a user interface for using a multimodal generation model based on images and utterances according to one embodiment.

[0165] Referring to FIGS. 2 and FIGS. 12, a wearable device (210) can capture an image (1215). Additionally, the wearable device (210) can receive a user utterance (1230) along with the image (1215). The image (1215) may include a first object (e.g., a monitor), and the user utterance (1230) may be a query requesting a description of the first object (e.g., a monitor).

[0166] The multimodal generation model (270) can provide information related to the user's previous experience corresponding to the current image (1215) by personalized data reflecting the user's experience. For example, the multimodal generation model (270) can generate and provide an image (1235) similar to an image that the user has previously seen (or taken) of among the first objects (e.g., monitors). In addition, the wearable device (210) can output classified text (1240) along with the generated image (1235) in a TTS manner.

[0167]

[0168] FIG. 13 shows the output of a multimodal generation model corresponding to a user's utterance related to a place according to one embodiment.

[0169] Referring to FIGS. 2 and FIGS. 13, a wearable device (210) can receive a user's utterance (216). A multimodal generation model (270) can be learned by personalized data that reflects the user's (405) experience. The multimodal generation model (270) can provide information related to the user's previous experience corresponding to the user's utterance (216).

[0170] For example, if the user's utterance (216) is related to a previously visited place, the multimodal generation model (270) can generate and display a video or photo (1310) related to the user's previous experience. According to one embodiment, when multiple videos are searched, the multimodal generation model (270) can arrange the multiple videos along a horizontal axis regarding importance or relevance. A first video (1310) regarding an object (e.g., a lake) with the highest relevance to the user's utterance (216), and a second video (1315) regarding an object (e.g., a schedule in Switzerland) with the second highest relevance to the user's utterance (216) can be arranged sequentially along the horizontal axis. According to one embodiment, the multimodal generation model (270) can arrange highlight images of each video along a vertical axis in chronological order.

[0171] In addition, text (1320) describing a video or photo (1310) and text (1330) describing the user's (405) past schedule at a visited place can be output in a TTS manner.

[0172]

[0173] FIG. 14 shows the output of a multimodal generation model that provides sequential responses in response to a user's utterance according to one embodiment.

[0174] Referring to FIGS. 2 and FIGS. 14, a multimodal generation model (270) can be learned by personalized data that reflects the experience of a user (405). A wearable device (210) can receive a first utterance (1411) of the user. The multimodal generation model (270) can provide information related to the user's previous experience corresponding to the first utterance (1411) of the user.

[0175] For example, if the user's first utterance (1411) is a query related to a previously experienced memory, the multimodal generation model (270) may generate and display a video (1425) related to the user's previous experience. Additionally, a first response (1411a) in clip units may be provided for the first utterance (1411). A clip unit may be a frame length of approximately 10 seconds or less. The first response (1411a) may be generated based on text containing words or brief descriptions corresponding to the user's first utterance (1411). Subsequently, when the user's second utterance (1412) is received, a second response (1412a) in clip units for the second utterance (1421) may be provided. The second response (1412a) may be generated based on text containing words or brief descriptions corresponding to the user's second utterance (1412).

[0176] When receiving a third utterance (1413) from a user requesting additional information or details, the multimodal generation model (270) may provide a third response (1413a) in segment units for the third utterance (1413). A segment unit may be a frame length between approximately 10 seconds and approximately 10 minutes. The third response (1413a) may consist of multiple sentences corresponding to the first through third utterances (1411, 1412, 1413).

[0177] When a fourth utterance (1414) from a user requesting full details is received, a fourth response (1414a) for the entire video unit regarding the fourth utterance (1414) may be provided. The fourth response (1414a) may be generated based on text summarizing the entire video (1425).

[0178]

[0179] FIG. 15 shows the output of a multimodal generation model based on web crawling according to one embodiment.

[0180] Referring to FIG. 2 and FIG. 15, the multimodal generation model (270) can be trained by personalized data that reflects the experience of the user (405). The wearable device (210) can capture an image (1510) in response to the user's voice command (216). The multimodal generation model (270) can recognize and distinguish objects included in the currently captured image (1510).

[0181] For example, if the image (1510) includes first to third objects (e.g., shoes, bar, pants), the multimodal generation model (270) may generate and provide an image similar to an image previously taken by the user among the first to third objects (e.g., shoes, bar, pants). Alternatively, text (1560) may be output in a TTS manner based on web crawled information (1550) related to the first to third objects (e.g., shoes, bar, pants).

[0182]

[0183] FIG. 16 shows the output of a multimodal generation model in a foldable device according to one embodiment.

[0184] Referring to FIGS. 2 and FIGS. 16, the multimodal generation model (270) can be trained by personalized data that reflects the user's experience. The multimodal generation model (270) can receive the user's utterance through an electronic device (220) and output a response corresponding to the user's utterance.

[0185] Depending on the state of the electronic device (220), a response may be output in a different way. For example, if the electronic device (220) is in a folding state (1601), information related to the user's previous experience may be provided in a text-based manner. If the electronic device (220) is in an unfolded state (1602), information related to the user's previous experience may be provided in an image and text-based manner.

[0186]

[0187] When a user searches for a photo or video stored on an electronic device, the device may display a photo or video corresponding to the user's input (e.g., touch input, voice input). For example, the device may find a photo or video similar to a sentence based on user input, or generate and display a random image related to a word entered by the user. In this case, the device cannot provide specific information regarding a long video. Furthermore, the randomized image has low relevance to the user, resulting in reduced user satisfaction.

[0188]

[0189] An electronic device according to one embodiment disclosed in this document can train a multimodal generation model using personalized data in which an image and text are paired.

[0190] An electronic device according to one embodiment disclosed in this document can retrieve previously experienced information in response to a user's question and generate and provide related images or text.

[0191] An electronic device according to one embodiment disclosed in this document can provide information about objects and scenes in stages as user feedback by using a multimodal generation model.

[0192] An electronic device according to one embodiment may include a memory and at least one processor. The memory may store instructions that, when executed individually or collectively by the at least one processor, cause the electronic device to acquire a first image related to a user, select a highlight frame based on a frame-by-frame rank constituting the first image, generate a caption related to the highlight frame and at least one object included in the highlight frame, generate personalized data for the user by linking the highlight frame and the caption, and train a multimodal generation model using the personalized data.

[0193] According to one embodiment, the first image may be captured through a camera of a wearable device worn by the user.

[0194] According to one embodiment, the first image may be transmitted to the electronic device along with sound occurring while the first image is being captured. When the instructions are executed individually or collectively by the at least one processor, the electronic device may generate the caption by reflecting the sound.

[0195] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may acquire a photograph related to a user, extract at least one object included in the photograph, and generate the caption based on the at least one object.

[0196] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be able to operate the multimodal generation model when the personalized data is greater than a specified capacity.

[0197] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may delete the personalized data used to train the multimodal generation model.

[0198] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may map the highlight frame and the caption to each other in a vector space dimension and perform alignment to form a correlation between the highlight frame and the caption to train the multimodal generation model.

[0199] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may generate a first caption by dividing the first image into first time units, generate a second caption in second time units longer than the first time unit based on the first caption, and generate a third caption for the entire first image based on the second caption.

[0200] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may receive a second image captured through a camera of a wearable device worn by the user or a first speech input of the user, generate a third image or a first text in a multimodal generation model based on the second image or the first speech input, and output at least one of the third image or the first text.

[0201] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may receive a second speech input requesting additional information, generate a fourth image or a second text in a multimodal generation model based on the third image or the second speech input, and output at least one of the third image or the second text.

[0202] A method for processing a user's speech according to one embodiment may be performed in an electronic device. The method may include the operation of acquiring a first image related to a user, the operation of selecting a highlight frame based on a rank per frame constituting the first image, the operation of generating a caption related to the highlight frame and at least one object included in the highlight frame, the operation of generating personalized data for the user by linking the highlight frame and the caption, and the operation of training a multimodal generation model using the personalized data.

[0203] According to one embodiment, the first image may be captured through a camera of a wearable device worn by the user.

[0204] According to one embodiment, the first image may be transmitted to the electronic device along with sound generated while the first image is being captured. The operation of generating the caption may include an operation of generating the caption by reflecting the sound.

[0205] According to one embodiment, the method may further include the operation of acquiring a photograph related to a user, the operation of extracting at least one object included in the photograph, and the operation of generating the caption based on the at least one object.

[0206] According to one embodiment, the operation of training the multimodal generation model may include the operation of training the multimodal generation model when the personalized data is greater than or equal to a specified capacity.

[0207] According to one embodiment, the method may further include the operation of deleting the personalized data used to train the multimodal generation model.

[0208] According to one embodiment, the operation of training the multimodal generation model may include the operation of training the multimodal generation model by mapping the highlight frame and the caption to each other in a vector space dimension and performing alignment to form a correlation between the highlight frame and the caption.

[0209] According to one embodiment, the operation of generating the caption may include the operation of generating a first caption by dividing the first image into first time units, the operation of generating a second caption in second time units longer than the first time unit based on the first caption, and the operation of generating a third caption for the entire first image based on the second caption.

[0210] According to one embodiment, the method may include receiving a second image captured through a camera of a wearable device worn by the user or a first speech input of the user, generating a third image or a first text in a multimodal generation model based on the second image or the first speech input, and outputting at least one of the third image or the first text.

[0211] According to one embodiment, the method may include receiving a second speech input requesting additional information, generating a fourth image or a second text in a multimodal generation model based on the third image or the second speech input, and outputting at least one of the third image or the second text.

[0212]

[0213] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., first) component is referred to as “coupled” or “connected” to another (e.g., second) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., by wire), wirelessly, or through a third component.

[0214] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0215] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0216] According to one embodiment, the method according to the various embodiments disclosed herein may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0217] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device, Memory; and It includes at least one processor comprising processing circuitry, and When the above memory is executed individually or collectively by the at least one processor, the electronic device, Acquire a first image related to the user, and Select highlight frames based on the frame-by-frame rank constituting the first video above, and Generates a caption related to the above highlight frame and at least one object included in the above highlight frame, and Personalized data for the user is generated by linking the above highlight frame and the above caption, and An electronic device that stores instructions for training a multimodal generative model using the above-mentioned personalized data.

2. In paragraph 1, the first image is An electronic device captured through the camera of a wearable device worn by the above user.

3. In paragraph 2, the first image is The above first image is transmitted to the electronic device along with the sound generated while the image is being captured, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that generates the above caption by reflecting the above voice.

4. In paragraph 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, Acquire photos related to the user, Extract at least one object included in the above photograph, and An electronic device that generates the caption based on at least one object.

5. In paragraph 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that enables the multimodal generation model to operate when the above personalized data exceeds a specified capacity.

6. In paragraph 5, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device for deleting the personal data used to train the multimodal generation model.

7. In paragraph 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that maps the highlight frame and the caption to each other in a vector space dimension and performs alignment to form a correlation between the highlight frame and the caption to train the multimodal generation model.

8. In paragraph 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, The first video is divided into first time units to generate a first caption, and Based on the first caption above, a second caption is generated in a second time unit longer than the first time unit, and An electronic device that generates a third caption for the entire first image based on the second caption above.

9. In paragraph 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, Receiving a second image captured through a camera of a wearable device worn by the user or a first speech input from the user, Generate a third image or a first text in a multimodal generation model based on the second image or the first speech input, and An electronic device that outputs at least one of the above-mentioned third image or the above-mentioned first text.

10. In paragraph 9, when the instructions are executed individually or collectively by the at least one processor, the electronic device, Receive a second speech input requesting additional information, Generate a fourth image or second text in a multimodal generation model based on the third image or the second speech input, and An electronic device that outputs at least one of the above-mentioned third image or the above-mentioned second text.

11. A method for processing a user's utterance performed in an electronic device, The operation of acquiring a first image related to the user; An operation of selecting highlight frames based on the frame-by-frame rank constituting the first image; The operation of generating a caption related to the above highlight frame and at least one object included in the above highlight frame; The operation of generating personalized data for the user by linking the above highlight frame and the above caption; and A method comprising the operation of training a multimodal generation model using the above-mentioned personalized data.

12. In Paragraph 11, Action of acquiring a photo related to the user; An operation to extract at least one object included in the above photograph; A method further comprising the operation of generating the caption based on at least one object.

13. In Clause 11, the operation of training the multimodal generation model is, A method comprising the operation of mapping the highlight frame and the caption to each other in a vector space dimension, and performing an alignment that forms a correlation between the highlight frame and the caption to train the multimodal generation model.

14. In paragraph 11, the operation of generating the above caption is, The operation of generating a first caption by dividing the first video into first time units; Based on the first caption above, the operation of generating a second caption in a second time unit longer than the first time unit; and A method comprising the operation of generating a third caption for the entire first image based on the second caption above.

15. In Paragraph 11, An operation of receiving a second image captured through a camera of a wearable device worn by the user or a first speech input from the user; The operation of generating a third image or a first text in a multimodal generation model based on the second image or the first speech input; A method comprising the operation of outputting at least one of the third image or the first text.