Electronic device and content generation method of electronic device
Patent Information
- Application Number
- PCT/KR2026/001286
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-04-15
- Filing Date
- 2026-01-21
- Publication Date
- 2026-10-01
Smart Images

Figure KR2026001286_01102026_PF_FP_ABST
Abstract
Description
Electronic device and method of creating content for the electronic device
[0001] The embodiments disclosed in this document relate to an electronic device and a method for creating content of the electronic device.
[0002] Conventional image editing technologies involve the inconvenience of requiring users to go through multiple manual steps to achieve desired effects. This complicates the editing process and can negatively impact the quality of the output or work efficiency. There is also the issue of repetitive manual editing being performed multiple times to reflect user intent.
[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art in relation to the present disclosure.
[0004] An electronic device according to one embodiment disclosed in this document may include an input module, a memory for storing instructions, and at least one processor. When the instructions are executed individually or collectively by at least one processor, the electronic device may acquire first data including at least one object, receive a first user input including at least one request information for generating at least one content associated with at least one of at least one object included in the first data through the input module, receive a second user input for specifying at least one of at least one object in the first data through the input module, identify one or more request information associated with at least one of at least one object among the at least one request information based on at least one of the segments of each of the first user input and the second user input received at a time when each of the first user input and the second user input is received or at a specified time, generate at least one prompt associated with at least one of at least one object based on the identified one or more request information, and generate content associated with at least one of at least one object based on at least one prompt through a generative artificial intelligence (AI) model.
[0005] A method for generating content of an electronic device according to an embodiment disclosed in this document may include: acquiring first data including at least one object; receiving a first user input including at least one request information for generating at least one content associated with each of at least one object through an input module of the electronic device; receiving a second user input for specifying each of at least one object in original data through a display of the electronic device; identifying one or more request information associated with each of at least one object among at least one request information based on at least one interval of each of the first user input and the second user input received at a time when each of the first user input and the second user input is received or at a predetermined time; generating at least one prompt associated with each of at least one object based on the identified one or more request information; and generating content associated with each of at least one object based on at least one prompt through a generative artificial intelligence (AI) model.
[0006] A non-transitory computer-readable medium according to an embodiment disclosed in this document may store at least one instruction that, when executed by at least one processor, enables an electronic device to perform the following operations: acquiring first data including at least one object; receiving a first user input including at least one request information for generating at least one content associated with at least one of at least one object included in the first data through an input module of the electronic device; receiving a second user input for specifying at least one of at least one object in the first data through an input module; identifying one or more request information associated with at least one of at least one object among at least one request information based on at least one interval of each of the first user input and the second user input received at a time when each of the first user input and the second user input is received or at a specified time; generating at least one prompt associated with at least one of at least one object based on the identified one or more request information; and generating content associated with at least one of at least one object based on at least one prompt through a generative artificial intelligence (AI) model.
[0007] FIG. 1 is a block diagram of an electronic device according to one embodiment of the present disclosure.
[0008] FIG. 2 is a diagram illustrating the operation of an electronic device generating content according to one embodiment of the present disclosure.
[0009] FIG. 3 is a flowchart of a content creation method according to one embodiment of the present disclosure.
[0010] FIG. 4 is a flowchart of a content creation method according to one embodiment of the present disclosure.
[0011] FIG. 5 is a drawing for explaining a method for creating content of an electronic device according to one embodiment of the present disclosure.
[0012] FIG. 6 is a diagram showing a graph in which a first user input and a second user input are synchronized in time according to one embodiment of the present disclosure.
[0013] FIG. 7 is a diagram illustrating an operation to identify request information associated with each object according to one embodiment of the present disclosure.
[0014] FIG. 8 is a drawing for explaining a method for generating a prompt associated with each object according to one embodiment of the present disclosure.
[0015] FIG. 9 is a drawing for explaining a method for creating content of an electronic device according to one embodiment of the present disclosure.
[0016] FIG. 10 is a drawing illustrating a method for identifying request information associated with each of at least one object according to one embodiment of the present disclosure.
[0017] FIG. 11 is a drawing showing a graph in which a first user input and a second user input are time-synchronized according to one embodiment of the present disclosure.
[0018] FIG. 12 is a diagram illustrating a method for identifying request information associated with each of at least one object according to one embodiment of the present disclosure.
[0019] FIG. 13 is a drawing showing a graph in which a first user input and a second user input are time-synchronized according to one embodiment of the present disclosure.
[0020] FIG. 14 is a drawing illustrating a method for identifying request information associated with each of at least one object according to one embodiment of the present disclosure.
[0021] FIG. 15 is a drawing for explaining a method for creating content according to one embodiment of the present disclosure.
[0022] FIG. 16 is a drawing for explaining a method for creating content according to one embodiment of the present disclosure.
[0023] FIG. 17 is a drawing illustrating a method for creating content according to one embodiment of the present disclosure.
[0024] FIG. 18 is a diagram illustrating a graph in which a first user input and a second user input are time-synchronized according to one embodiment of the present disclosure.
[0025] FIG. 19 is a drawing illustrating a method for identifying request information associated with each of at least one object according to one embodiment of the present disclosure.
[0026] FIG. 20 is a drawing for explaining how to display content generated according to one embodiment of the present disclosure through a display.
[0027] FIG. 21 is a drawing illustrating a method for creating content according to one embodiment of the present disclosure.
[0028] FIG. 22 is a drawing showing a graph in which a first user input and a second user input are time-synchronized according to one embodiment of the present disclosure.
[0029] FIGS. 23a and FIGS. 23b are drawings illustrating a method for identifying request information associated with each of at least one object by a content creation method according to one embodiment of the present disclosure.
[0030] FIG. 24 is a drawing for explaining a method for creating content according to one embodiment of the present disclosure.
[0031] FIG. 25 is a drawing illustrating a method for creating content of an electronic device according to one embodiment of the present disclosure.
[0032] FIG. 26 is a block diagram of an electronic device in a network environment according to one embodiment.
[0033] FIG. 27 is a block diagram of a generative artificial intelligence (AI) system according to one embodiment.
[0034] FIG. 28 is a block diagram of an artificial intelligence framework (AI framework) according to one embodiment.
[0035] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0036] FIG. 1 is a block diagram of an electronic device according to one embodiment of the present disclosure.
[0037] Referring to FIG. 1, an electronic device (100) according to one embodiment of the present disclosure (e.g., the electronic device (2510) of FIG. 25 or the electronic device (2601) of FIG. 26) may include a display (110) (e.g., 2520 of FIG. 25, 2660 of FIG. 26), an input module (120) (e.g., 2650 of FIG. 26), a memory (130) (e.g., 2630 of FIG. 26), and a processor (140) (e.g., 2620 of FIG. 26).
[0038] According to one embodiment, the display (110) can visually display information. For example, the display (110) can provide various screens individually, in parallel, or sequentially depending on the user's input and / or the processing status of the electronic device. For example, the display (110) can display a screen that provides data (e.g., images, text) to the user. For example, the display (110) can display a screen for selecting or specifying an object. For example, the display (110) can display a screen that confirms user input, including content creation request information. For example, the display (110) can display a screen that confirms the created prompt and / or content. For example, the display (110) can display a screen that confirms the user's request to change the created content.
[0039] According to one embodiment, the display (110) may be integrally formed with an input module (120) (e.g., a touch panel). For example, the display (110) may include a touch screen display (110). For example, the display (110) may receive inputs requesting the selection of an object or the creation of content by recognizing inputs such as a user's touch, gesture, drag, or gaze. For example, in an electronic device (100) that provides functions such as VR (virtual reality), AR (augmented reality), MR (mixed reality), and XR (extended reality), the display (110) may select an object or perform specific inputs by detecting the user's gaze movements.
[0040] According to one embodiment, the input module (120) can receive various user inputs and transmit them to the processor (140). For example, the input module (120) may include hardware such as a microphone, speaker, pressure sensor, mouse, keyboard, camera, physical button, etc. The input module (120) may be configured to detect and process various user inputs, including at least one of user voice input, touch input, keyboard input, mouse input, gesture input, drag input, and eye tracking input.
[0041] According to one embodiment, the memory (130) may store instructions that control the operation of the electronic device (100) when executed individually or collectively by at least one processor (140). For example, the instructions may be stored in one memory (130) or multiple memories (130). The memory (130) may store information and / or data related to the operations of the electronic device (100) at least temporarily. The memory (130) may include a memory (130) independent of at least one processor (140) and may include a processor internal memory (130) mounted on at least one processor (140).
[0042] According to one embodiment, the memory (130) can store data and computation results related to an artificial intelligence (AI) model. For example, the memory (130) can be utilized in the process of AI-based content creation and processing by storing training data, model weights and parameters, prompts and intermediate processing results, and finally generated content of a generative AI model. For example, the electronic device (100) can use the memory (130) to generate a prompt based on a natural language command entered by a user, and to store the generated content after the generative AI model processes the prompt. Additionally, if the AI model performs continuous learning, the memory (130) can support gradual performance improvement by storing additional data sets reflecting user preferences, feedback information, and model update files.
[0043] According to one embodiment, the memory (130) may be configured to interact with other components such as a display (110), an input module (120), and a processor (140). For example, the memory (130) may at least temporarily store data such as images, text, object information, content associated with objects, and user change requests displayed on the display (110). The display (110) may display data stored in the memory (130). The memory (130) may at least temporarily store user input data received from the input module (120). The processor (140) may analyze and process instructions and data stored in the memory (130) to process user input and output processing information through the display (110).
[0044] According to one embodiment, the processor (140) can control the operations of the electronic device (100) by executing instructions stored in memory (130) individually or collectively. For example, the operations described in this disclosure as being performed by the electronic device (100) may be understood as being performed by the processor (140). For example, the processor (140) may include at least one processor (140). For example, the operations described in this disclosure as being performed by the processor (140) may be understood as being performed individually or collectively by at least one processor (140). For example, at least one processor (140) may control each of the operations of the electronic device (100) described below, either independently or collectively. According to various embodiments, at least one processor (140) may include circuits such as a central processing unit (CPU), a micro-processor unit (MPU), an application processor (AP), a communication processor (CP), a system on chip (SoC), and / or an integrated circuit (IC).
[0045] According to one embodiment, the processor (140) may acquire input data comprising at least one object. For example, the input data may be provided by a user through an input module (120) or received from a database, an external device, a server, etc., through a communication circuit. For example, the input data may include images, text, video, audio data, etc. However, the format of the input data is not limited thereto. For example, when including at least one object that is the target of content creation, the input data may be configured in various forms.
[0046] According to one embodiment, the processor (140) can analyze acquired input data to identify a specific object or extract information associated with the object. For example, the processor (140) can visually display at least one of the acquired input data, the identified specific object, or the extracted information through a display (110) or output it in various forms through an output module.
[0047] According to one embodiment, the processor (140) may receive a first user input through an input module (120) (e.g., a microphone, a keyboard). For example, the first user input may include request information associated with each of at least one object. For example, the first user input may include at least one request information for creating at least one content associated with at least one of the at least one objects. For example, the first user input may be received in the form of voice, including the user's speech, through a microphone. However, the form or method of receiving the first user input is not limited thereto. For example, the first user input may be composed of text input.
[0048] According to one embodiment, the processor (140) may receive a second user input through an input module (120) (e.g., a touch screen included in the display (110). For example, the second user input may include at least one of a touch input, a gesture input, a drag action, or an eye-tracking input for specifying at least one of at least one object in the first data displayed through the display (110).
[0049] According to one embodiment, the processor (140) can identify one or more request information associated with at least one of at least one object among at least one request information included in the first user input. For example, the processor (140) can identify one or more request information associated with at least one of at least one object among at least one request information included in the first user input based on the time at which the first user input and the second user input are each received. For example, the processor (140) can identify one or more request information associated with at least one of at least one object among at least one request information included in the first user input based on the interval of the first user input and the second user input received at a specified time.
[0050] According to one embodiment, the processor (140) can identify one or more request information associated with a first object included in at least one object among the first user inputs. For example, the processor (140) can identify a first segment of the first user input received while receiving a second user input for specifying a first object included in at least one object. For example, the processor (140) can identify the request information associated with the identified first segment as one or more request information associated with the first object. For example, the processor (140) can identify at least a portion of the first user input, where the similarity with the first segment is greater than or equal to a specified threshold value, as a segment associated with the first segment or request information associated with the first segment. For example, the processor (140) can identify a third segment of the first user input received before receiving a second user input for specifying a second object included in at least one object. For example, the processor (140) can identify the third section as at least one of the section associated with the first section, the request information associated with the first section, or one or more request information associated with the first object.
[0051] According to one embodiment, the processor (140) can identify one or more request information associated with the first object included in at least one object among the first user inputs. For example, the processor (140) can identify a second segment of the first user input received while receiving a second user input for specifying a second object. For example, the processor (140) can identify the request information associated with the second segment as one or more request information associated with the second object. For example, the processor (140) can identify at least a portion of the first user input, for which the similarity with the second segment is greater than or equal to a specified threshold, as a segment associated with the second segment or a part associated with the second segment. For example, the processor (140) can identify a fourth segment of the first user input received during at least one period among (1) while receiving a second user input for specifying a second object, (2) during a specified time after receiving the second user input for specifying a second object, or (3) before receiving the second user input for specifying a third object included in the first data. For example, the processor (140) can identify the fourth segment as one or more request information associated with the second object.
[0052] According to one embodiment, the processor (140) may identify each of at least one first interval received while receiving a second user input for specifying each of at least one object among the first user inputs as an interval that corresponds temporally to each of at least one object. For example, the processor (140) may calculate a first similarity between each of at least one first interval that corresponds temporally to each of at least one object and each of at least one object.
[0053] According to one embodiment, if the calculated first similarity is greater than or equal to a specified first threshold value, the processor (140) can identify request information associated with each of at least one first interval as one or more request information associated with each of at least one object.
[0054] According to one embodiment, if the calculated first similarity is less than a specified first threshold, the processor (140) may identify one or more first intervals that are not temporally corresponding to each of at least one object among at least one first interval. For example, if the second similarity between each of one or more first intervals and each of at least one object is greater than or equal to a specified second threshold, the processor (140) may identify at least one request information associated with each of one or more first intervals as one or more request information associated with each of at least one object.
[0055] According to one embodiment, if the calculated first similarity is less than a specified first threshold, the processor (140) can identify at least one second segment other than at least one first segment in the first user input. For example, if the second similarity between each of the at least one second segment and each of the at least one object is greater than or equal to a specified second threshold, the processor (140) can identify at least one request information associated with each of the at least one second segment as one or more request information associated with each of the at least one object.
[0056] According to one embodiment, the processor (140) may generate at least one prompt associated with each of at least one object based on each of one or more request information. For example, the processor (140) may determine a content generation method including the number of contents to be generated associated with each of at least one object based on one or more request information associated with each of at least one object. For example, the processor (140) may generate at least one prompt associated with each of at least one object based on the determined content generation method and one or more request information.
[0057] According to one embodiment, the processor (140) can generate content associated with each of at least one object based on each of at least one prompt. For example, the processor (140) can display the generated content through an output module (e.g., a display (110), a speaker).
[0058] According to one embodiment, a communication circuit (not shown) can transmit and receive information and / or data to and from an external device. For example, the communication circuit may include at least one modem. For example, the communication circuit may transmit a transmission signal corresponding to a user's utterance to an external device during a call, and receive a reception signal corresponding to a utterance of the user of the external device (i.e., the call partner) from the external device.
[0059] FIG. 2 is a diagram illustrating the operation of an electronic device generating content according to one embodiment of the present disclosure.
[0060] The operation described in this disclosure as being performed by the electronic device (100) can be understood as the electronic device (100) performing the operation by having at least one processor (140) execute instructions stored in memory (e.g., 130 in FIG. 1, 2630 in FIG. 26) individually or collectively. In the following description of FIG. 2, content that overlaps with FIG. 1 may be omitted or described briefly.
[0061] According to one embodiment, an electronic device (100) may acquire first data (210) comprising at least one object (212, 214). The electronic device (100) may receive the first data (210) from a user through an input module (120) or from a database, an external device and / or an external server through a communication circuit. The electronic device (100) may display or output the acquired first data through a display (110) and / or an output module.
[0062] According to one embodiment, the first data (210) may include an image containing at least one object (212, 214). For example, the at least one object may include a background, an object and / or a person. For example, the background may include a wall, a floor, the sky, and / or a chair. For example, the object may include clouds, a top, a bottom, and / or shoes. For example, the person may include one or more people. According to various embodiments, the background, object, and person are not limited to those listed above. For example, the first object (212) (e.g., a wall) may refer to a wall or a part of a wall located at the top of the person on the far left, and the second object (214) (e.g., pants) may refer to the pants or bottom of the person on the far right.
[0063] According to one embodiment, the electronic device (100) may receive a first user input (220). The first user input (220) may include at least one request information for creating at least one content associated with at least one of at least one object. For example, the first user input (220) may include request information for creating content associated with the first object (212) (e.g., a fireworks display) (e.g., "Draw a fireworks display on this wall"). Additionally, the first user input (220) may include request information for creating content associated with the second object (214) (e.g., jeans) (e.g., "Draw these pants as jeans").
[0064] According to one embodiment, an electronic device (100) may receive a first user input (220) through an input module (e.g., the input module (120) of FIG. 1). For example, the electronic device (100) may receive the first user input (220) (e.g., "Draw a fireworks symbol on this wall, and draw these pants as jeans") as voice through an input module including a microphone. However, the method by which the electronic device (100) receives the first user input (220) is not limited thereto. For example, the electronic device (100) may receive the first user input (220) entered via a keyboard in the form of text.
[0065] According to one embodiment, the electronic device (100) may receive a second user input (230, 232). For example, the second user input (230, 232) may include at least one input for specifying or designating at least one of at least one object (212, 214) in the first data (210). For example, the second user input (230, 232) may include a second user input (230) for designating the first object (212) in the first data (210) and a second user input (232) for designating the second object (214) in the first data (210).
[0066] According to one embodiment, the electronic device (100) may receive a second user input (230, 232) through an input module (e.g., a touch screen included in the display (110). For example, the electronic device (100) may receive a second user input (230) for designating the first object (212) in such a way that the first object (212) is selected through the display (110). For example, the electronic device (100) may receive a second user input (230) for designating the first object (212) in such a way that an area associated with the first object (212) (e.g., a part of a wall) is selected through the display (110). For example, when the second user input (230) is received to designate the first object (212) with a circular gesture through the display (110), the electronic device (100) may specify or designate the first object (212). For example, when receiving user input through the display (110) such that the first object (212) is touched or dragged for a specified period of time, the electronic device (100) may identify or specify the first object (212). For example, when receiving user input through the display (110) such that the user gazes at the first object (212) or an area associated with the object for a specified period of time, the electronic device (100) may identify the first object (212) by tracking the user's gaze. However, the method of implementing the second user input (230) for identifying or specifying the first object (212) within the first data (210) may include various variations according to other embodiments of the present invention and is not limited thereto.
[0067] According to one embodiment, the electronic device (100) may receive a second user input (232) for specifying a second object (214) through an input module. For example, the electronic device (100) may specify or designate the second object (214) within the first data (210) based on the second user input (232). For example, the electronic device (100) may receive a second user input (232) for drawing a circular gesture around the area of the second object (214) through a display (110). For example, the electronic device (100) may receive a second user input (232) for touching or dragging the second object (214) for a specified time through a display (110). For example, the electronic device (100) may receive a second user input (232) containing gaze information for gazing at a second object (214) and / or an area associated with the second object (214) for a specified period of time through a camera. However, the method of implementing the second user input (232) for identifying or designating the second object (214) within the first data (210) may include various variations according to other embodiments of the present invention and is not limited thereto.
[0068] According to one embodiment, the electronic device (100) may receive a second user input (e.g., 230, 232) for specifying each of at least one object (e.g., 212, 214) independently or individually. Depending on the characteristics of each object (e.g., 212, 214) included in the first data (210), the user environment for the electronic device (100), etc., the user may input the second user input (230, 232) in an appropriate manner. For example, the electronic device (100) may specify the first object (212) based on the user's gesture (e.g., circular gesture) on the display (110), and specify the second object (214) by tracking the user's gaze while gazing at the specific object for a specified period of time. As another example, the electronic device (100) may specify a first object (212) by a user's tap action on the display (110) and specify a second object (214) by an input of pressing the specific object for a specified time. However, the above examples are merely illustrative of some embodiments of the present invention, and the combination and implementation method of the second user input (e.g., 230, 232) for specifying or designating at least one object (e.g., 212, 214) may include various variations according to other embodiments of the present invention.
[0069] According to one embodiment, the electronic device (100) may sequentially receive a second user input (e.g., 230, 232) for specifying each of at least one object (e.g., 212, 214). For example, the electronic device (100) may recognize the order in which at least one second user input (e.g., 230, 232) is received sequentially. For example, the electronic device (100) may first receive a user input for specifying a first object (212) through a display (110), and then receive a user input for specifying a second object (214). For example, the electronic device (100) may receive a second user input (230) (e.g., a circular gesture) for specifying the first object (212), and then receive a second user input (232) (e.g., a circular gesture) for specifying the second object (214).
[0070] According to one embodiment, the sequential input method of the second user input (230, 232) for specifying each of at least one object (212, 214) may vary depending on the user environment or the state of the input device. For example, an electronic device (100) can identify the number of objects (212, 214) to be created by analyzing a first user input (220) containing at least one request information for creating content. The electronic device (100) can receive the second user input (230, 232) for specifying at least one object (212, 214). Subsequently, the electronic device (100) can compare the number of received second user inputs (e.g., 230, 232) with the number of objects to be created identified through the analysis of the first user input (220). If it is determined that the user has specified all the objects required for content creation, the electronic device (100) can perform subsequent operations (e.g., creating a prompt or creating content) according to a specified process.
[0071] According to one embodiment, if the number of second user inputs (230, 232) is less than the number of objects to be created for content generation, the electronic device (100) may induce sequential input of the second user inputs (230, 232). For example, the electronic device (100) may provide a UI or guidance message that induces additional object designation. For example, the electronic device (100) may receive a first user input (220) (e.g., "Draw a fireworks symbol on this wall and draw these pants as jeans") and a second user input (230) for designating a first object (212) (e.g., wall), but may not receive a second user input (232) for designating a second object (214) (e.g., pants). In this case, the electronic device (100) may display a guidance (e.g., "Would you like to designate other objects as well?") that induces additional object designation through a display (110) and / or an output module. As another example, the electronic device (100) may recommend an additional object to be designated based on the first user input (220) and the received second user input (230, 232) when there is no user input to additionally designate an object even after a specified time has elapsed.
[0072] According to one embodiment, if the number of objects specified through the second user input (230, 232) is greater than the number of objects to be created for content determined through the analysis of the first user input (220), the electronic device (100) may perform an operation to verify the request information and / or adjust the object designation. For example, the electronic device (100) may identify unnecessary objects among the objects specified through the second user input (230, 232) and exclude the identified unnecessary objects from the content creation target. As another example, the electronic device (100) may induce the user to confirm the object designation through a confirmation message for the second user input (230, 232) (e.g., "It is determined that more objects have been specified than intended by the user. Do you want to create content for all specified objects?").
[0073] According to one embodiment, the electronic device (100) can accurately determine the user's intention to create content by determining unnecessary objects or inducing the user to specify additional objects based on the result of comparing the first user input (220) and the second user input (230, 232).
[0074] According to one embodiment, the electronic device (100) can generate content associated with at least one of at least one object (212, 214) based on a first user input (220) and a second user input (230, 232) for specifying at least one object (212, 214).
[0075] According to one embodiment, the electronic device (100) can identify one or more request information associated with at least one of at least one object (212, 214) among the first user input (220) based on the first user input (220) and the second user input (230, 232). For example, the electronic device (100) can identify one or more request information associated with at least one of at least one object (212, 214) among at least one request information included in the first user input (220) based on the order in which the first user input (220) and the second user input (230, 232) are each received. As another example, the electronic device (100) can identify one or more request information associated with at least one of at least one object (212, 214) among at least one request information based on at least a portion of each of the first user input (220) and the second user input (230, 232) received at a specified time.
[0076] According to one embodiment, the electronic device (100) may receive a first user input (220) through a microphone. For example, the first user input (220) may include request information associated with a first object (212) (e.g., a wall) (e.g., "draw a fireworks mark on the wall") and request information associated with a second object (214) (e.g., pants) (e.g., "draw pants as jeans"), or may be a combination thereof. For example, the first user input (220) may be in the form of "draw a fireworks mark on this wall, and draw these pants as jeans." For example, the electronic device (100) may receive a second user input (230) for identifying or specifying the first object (212) in the first data (210). For example, if a user specifies a part of a wall with a circular gesture while saying "on this wall," the electronic device (100) may receive this as an input to specify a first object (212). For example, the electronic device (100) may receive a second user input (232) to specify or designate a second object (214) from the first data (210). For example, if a user specifies pants with a circular gesture while saying "these pants," the electronic device (100) may receive this as an input to designate a second object (214).
[0077] According to one embodiment, the electronic device (100) can analyze the first user input (220) and the second user input (230, 232) to match each request information with an appropriate object (212, 214). The electronic device can process the voice data of the first user input (220) in real time and synchronize it with the object designation action of the second user input (230, 232). For example, if the user says "on this wall" and designates "wall" with a circular gesture, the electronic device (100) can detect this and connect the request information "draw a firework" with the corresponding first object (212) (e.g., wall). Additionally, if the user says "these pants" and designates "pants" with a circular gesture, the electronic device (100) can match the request information "draw jeans" with the second object (214) (e.g., pants).
[0078] According to one embodiment, an electronic device (100) may generate at least one prompt associated with at least one of at least one object (212, 214) based on one or more identified request information. For example, the request information may include natural language utterance (e.g., voice command), text input, etc., as source data of the first user input. For example, the prompt may include data converted into a specific input format that a generative AI model can process based on the request information. For example, the electronic device (100) may interpret the user's request information (e.g., "Draw a fireworks symbol on this wall") and organize it according to a format that an AI model can understand. For example, the electronic device (100) may derive necessary metadata (e.g., style, color, background information, masking information, etc.) in addition to the user's request information. For example, the electronic device (100) can generate a prompt (e.g., prompt 1, create an image, to match the description: Add a fireworks display effect on the wall section of the provided image. Do not modify any other part of the image). based on the results and / or metadata of analyzing the user's request information.
[0079] According to one embodiment, an electronic device (100) can generate at least one content associated with each of at least one object based on at least one generated prompt through a generative AI model. For example, the electronic device (100) can generate second data including at least one content associated with each of at least one object (212, 214) based on at least one content and first data (210). The second data may correspond to data in which the changed object is reflected in the existing first data (210). The second data may be generated through processes such as editing the first data (210), changing the color, or inserting another object into the object and / or area of the first data (210).
[0080] According to one embodiment, the electronic device (100) may display at least one generated content and / or a second data through a display (110). For example, the electronic device (100) may provide a preview screen before applying the generated content to the first data. For example, the electronic device (100) may provide a user interface (UI) that can adjust the generated content.
[0081] FIG. 3 is a flowchart of a content creation method according to one embodiment of the present disclosure.
[0082] An operation described in the present disclosure as being performed by an electronic device (e.g., electronic device (100) of FIG. 1, electronic device (2510) of FIG. 25, or electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 140 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively.
[0083] According to one embodiment of the present disclosure, an electronic device (100) may perform a method (300) for generating content. The order of operations described below in relation to FIG. 3 is an example, and the embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be executed differently from the order of FIG. 3 or may be executed substantially simultaneously with other operations of FIG. 3. In the following description of FIG. 3, content that overlaps with FIG. 1 and FIG. 2 may be omitted or briefly described.
[0084] According to one embodiment, in operation 310, the electronic device (100) may acquire first data including at least one object. For example, the electronic device (100) may display the acquired first data through a display (e.g., 110 in FIG. 1, 2520 in FIG. 25, 2660 in FIG. 26). For example, a user may visually recognize the first data displayed through the display (110) and identify at least one object included in the first data (e.g., an object that is the target of the content to be created (e.g., a wall, pants)).
[0085] According to one embodiment, in operation 320, the electronic device (100) may receive a first user input including at least one request information for creating at least one content associated with each of at least one object through an input module (e.g., 120 in FIG. 1, 2650 in FIG. 26). For example, the first user input may include at least one request information requesting the creation of content associated with each of at least one object. For example, the at least one request information may include a sentence, word, or phrase requesting the creation of content. Additionally, the first user input may be input in the form of voice input, text input, natural language commands, etc.
[0086] According to one embodiment, in operation 330, the electronic device (100) may receive a second user input for the input module (120) to specify each of at least one object in the first data. The second user input may include an input action performed by the user to select or specify a specific object on the display (110). For example, the user may select a specific object by directly touching it on the display (110). As another example, the user may specify it by a gesture of drawing a circle around the specific object or inputting a specific pattern. As yet another example, the user may specify a specific object by dragging it on the display using a finger or a pen. As yet another example, if the user gazes at a specific object for a specified period of time, the electronic device (100) may specify the specific object based on the user's gaze information obtained through a camera.
[0087] According to one embodiment, in operation 340, the electronic device (100) can identify one or more request information associated with each of at least one object among at least one request information included in the first user input. For example, the electronic device (100) can identify request information associated with each object by analyzing the temporal relationship in which the first user input and the second user input are received. For example, the electronic device (100) can identify request information associated with each object based on the time when the first user input and the second user input are each received. For example, the electronic device (100) can identify one or more request information based on the interval of each of the first user input and the second user input received at a specified time.
[0088] According to one embodiment, in operation 340, the electronic device (100) may receive a first user input and a second user input sequentially or in parallel. For example, the electronic device (100) may receive the first user input first, and then receive the second user input sequentially. For example, the electronic device (100) may determine that the user first specifies an object and requests the creation of content associated with the specified object. For example, if the user first specifies a specific object and then inputs request information, the electronic device (100) may match the object with the subsequent request information. As another example, the electronic device (100) may receive the second user input first, and then receive the first user input sequentially. For example, the electronic device (100) may determine that the user first requests the creation of content associated with a specific object and specifies an object corresponding to the target for content creation. For example, if a user first inputs request information and then specifies a particular object, the electronic device (100) can identify that the most recently received request information is associated with that object.
[0089] According to one embodiment, in operation 340, the electronic device (100) can identify request information associated with each object by analyzing the temporal relationship in which the first user input and the second user input are received, respectively. For example, if a user touches a wall while uttering specific request information (e.g., "Draw a firework symbol on this wall"), the electronic device (100) can identify that the request information is associated with the wall. For example, if the user utters specific request information (e.g., "Draw a firework symbol on this wall") within a specified time after touching the wall, the electronic device (100) can match that request information with the request information for the object.
[0090] According to one embodiment, in operation 350, the electronic device (100) can generate at least one prompt associated with at least one of at least one object based on one or more identified request information. The electronic device (100) can generate a prompt in a form that an AI model can process by combining request information extracted from a first user input and object information designated as a second user input.
[0091] According to one embodiment, in operation 350, the electronic device (100) may generate a prompt based on rules or generate a prompt by utilizing at least one of a natural language processing (NLP) model or an AI model. For example, the electronic device (100) may analyze request information (e.g., "Draw a fireworks display on this wall") and generate a prompt based on specified rules (e.g., [Object] Wall, [Action] Draw or add, [Content Type] Fireworks display). For example, the electronic device (100) may generate a prompt by utilizing an NLP model to understand the context of the request information. For example, the electronic device (100) may utilize an AI model to generate or correct a prompt, or add detailed configurations (e.g., color, brightness, style) to the generated prompt.
[0092] According to one embodiment, in operation 360, the electronic device (100) can generate at least one piece of content associated with at least one of at least one object based on at least one prompt through a generative artificial intelligence (AI) model. For example, the electronic device (100) can generate new second data by combining the generated content with existing first data, or display the generated second data through a display (110).
[0093] FIG. 4 is a flowchart of a content creation method according to one embodiment of the present disclosure.
[0094] An operation described in the present disclosure as being performed by an electronic device (e.g., electronic device (100) of FIG. 1, electronic device (2510) of FIG. 25, or electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 140 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively.
[0095] According to one embodiment of the present disclosure, an electronic device (100) can perform a content creation method (400). The order of operations described below in relation to FIG. 4 is an example, and the embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be executed differently from the order of FIG. 4 or may be executed substantially simultaneously with other operations of FIG. 4. In the following description of FIG. 4, content that overlaps with FIG. 1 to 3 may be omitted or briefly described.
[0096] According to one embodiment, the content creation method (400) of FIG. 4 may correspond to operations 340 and 350 of the content creation method (300) of FIG. 3.
[0097] According to one embodiment, the electronic device (100) can analyze the temporal relationship between the first user input (220) and the second user input (240) to more accurately match object-specific request information.
[0098] According to one embodiment, an electronic device (100) may receive, through an input module, at least one first user input including content creation request information associated with at least one object and / or at least one second user input for specifying at least one object in parallel or individually. For example, the electronic device (100) may receive at least one second user input at a specific time while receiving the first user input. For example, the electronic device (100) may receive the first user input after receiving the second user input. For example, the electronic device (100) may receive at least one second user input after receiving the first user input.
[0099] According to one embodiment, in operation 410, the electronic device (100) can identify one or more request information associated with a first object in a first user input. For example, while receiving a second user input for specifying a first object, the electronic device (100) can identify a portion of the first user input received at that time as a first portion. For example, the electronic device (100) can analyze the first portion to extract request information associated with the first portion and match or identify it as one or more request information associated with the first object.
[0100] According to one embodiment, the electronic device (100) can identify a specific section (e.g., a first section) of the first user input received within a given time based on time information in which the second user input received while the first user input was received. For example, the electronic device (100) can synchronize the time when the first user input and the second user input occur. For example, the electronic device (100) can match the request information entered at a specific point in time with a designated object based on the synchronized first user input and the second user input.
[0101] According to one embodiment, the electronic device (100) can synchronize the occurrence times of a first user input and a second user input. For example, the electronic device (100) can record the time at which each user input occurs as a timestamp through an internal clock and / or an external synchronization device (e.g., a global positioning system (GPS), a network time protocol (NTP)). For example, the electronic device (100) can determine the reception time information of each input (e.g., the start time of reception, the time of intersection, and / or the order of reception) based on the temporal synchronization of each input.
[0102] For example, an electronic device (100) can receive a first user input (e.g., voice input via a microphone) and a second user input (e.g., gesture or touch input via a display and / or gaze input via a camera). For example, the electronic device (100) can record the start time of receiving the first user input and the second user input as a timestamp. For example, the electronic device (100) can identify the order of receiving each input based on the timestamp. For example, the electronic device (100) can identify an overlapping section of multiple inputs (e.g., a specific section of the first user input) based on analyzing the timestamp of the second user input that occurred at a specific point in time while the first user input is being received. For example, the electronic device (100) can calculate the time interval between receiving multiple inputs based on the completion time of receiving each input and the overall order of inputs. For example, the electronic device (100) can determine whether to recognize each input as the same event based on whether the calculated reception time interval is within a specified time interval.
[0103] Through this configuration, the electronic device (100) can more precisely determine the start time of reception, the intersection time, the completion time of reception, and / or the entire sequence of the first user input and the second user input based on synchronized time information. The electronic device (100) can designate an object to match the user's intention, effectively integrate and match it with the request information included in the user input, and generate related content.
[0104] According to one embodiment, an electronic device (100) can identify request information that is temporally and / or semantically associated with a first object designated by a second user input within a first user input, based on reception time information of a synchronized first user input and a second user input. For example, the electronic device (100) can identify, among the first user inputs, a portion of the first user input (a first portion) received during the time of receiving the second user input for designating the first object. For example, the electronic device (100) can additionally identify, among the first user inputs, portions (a portion associated with the first portion) where the similarity to the first portion is greater than or equal to a specified threshold value. For example, the electronic device (100) can identify the first portion and / or portion associated with the first portion among the first user inputs as request information associated with the first object.
[0105] According to one embodiment, when a user sequentially designates a plurality of objects, the electronic device (100) can identify request information associated with each object by separating them in time. The electronic device (100) can extract request information associated with each of the plurality of objects based on time information in which a first user input and a second user input for designating each of the plurality of objects are received. For example, the electronic device (100) may receive a second user input for designating a second object after receiving a second user input for designating a first object. The electronic device (100) may receive a second user input for designating a first object before receiving a second user input for designating a second object. The electronic device (100) may identify a portion of the first user input received before receiving the second user input for designating a second object. The electronic device (100) may identify the corresponding portion as request information associated with the first portion or one or more request information associated with the first object.
[0106] According to one embodiment, the electronic device (100) can analyze the meaning of each of the first user input and the first segment and extract request information associated with the first segment. For example, the electronic device (100) can identify one or more request information associated with each object based on the similarity between the first user input and the second user input. At least a portion of the first user input, where the similarity with the first segment is greater than or equal to a specified threshold value, can be identified as a portion associated with the first segment.
[0107] According to one embodiment, the electronic device (100) can receive voice data spoken by a user, such as “Draw a fireworks mark on this wall and draw these pants as jeans,” as a first user input through a microphone. Additionally, the electronic device (100) can receive a circular gesture input for specifying a ‘wall’ as a second user input through a display (e.g., a touch screen) while the user says “on this wall.”
[0108] For example, the electronic device (100) can obtain information on the reception time of a first user input (e.g., "Draw a fireworks symbol on this wall and draw these pants as jeans") and a second user input (e.g., a circular gesture input to specify 'wall'). For example, the electronic device (100) can temporally synchronize the first user input and the second user input based on the time at which the input received first among the first user input and / or the second user input begins. For example, the electronic device (100) can identify, based on the information where each input is temporally synchronized, the time at which the first user input begins to be received, the time at which the first user input is completed to be received, the time at which the second user input begins to be received, the time at which the second user input is completed to be received, the time interval during which the first user input is received, and / or a portion of the first user input received during the time at which the second user input is received.
[0109] For example, the electronic device (100) can identify a first segment of the received first user input (e.g., "on this wall") while a second user input (e.g., a circular gesture input to specify 'wall') is received among the first user input (e.g., "Draw a fireworks mark on this wall and draw these pants as jeans"). For example, in the present disclosure, the first segment of the first user input (e.g., "on this wall") may be referred to as a first segment that corresponds temporally to a first object (e.g., 'wall'). A detailed description thereof is provided later in FIGS. 5 to 7.
[0110] For example, the electronic device (100) can identify request information associated with a first interval (e.g., "on this wall") in the first user input by considering the temporal reception relationship between the first user input and the second user input for specifying a plurality of objects. A detailed explanation of this is provided later in FIGS. 5 to 7.
[0111] As another example, the electronic device (100) can identify request information associated with a first segment of the first user input (e.g., "on this wall") among the first user inputs. For example, the electronic device (100) can semantically analyze the first user input and identify a segment (e.g., "draw a fireworks mark on this wall") that has semantic or contextual similarity to the first segment of the first user input (e.g., "on this wall") above a specified threshold value as request information associated with the first segment (e.g., "on this wall"). A detailed description of this is provided later in FIGS. 9 and 10.
[0112] According to one embodiment, in operation 420, the electronic device (100) can identify one or more request information associated with a second object among the first user inputs. For example, the electronic device (100) can identify request information associated with a second section of the first user input received while receiving a second user input for specifying the second object as one or more request information associated with the second object.
[0113] According to one embodiment, the electronic device (100) can more accurately identify request information based on semantic association as well as temporal association between the first user input and the second user input. For example, the electronic device (100) can identify at least one sentence, at least one phrase, and / or at least one word constituting the first user input. For example, the electronic device (100) can analyze the contextual meaning of the identified at least one sentence, phrase, and / or word. For example, based on the results of this analysis, the electronic device (100) can evaluate the degree of semantic association between the object designated as the second user input and the identified at least one sentence, phrase, and / or word. For example, the electronic device (100) can convert each phrase included in the first user input into an embedding vector using a natural language processing (NLP) model (e.g., BERT (bidirectional encoder representations from transformers), Word2Vec()). For example, the embedding vector quantifies the semantic characteristics of each phrase in a high-dimensional space, and keywords associated with the same object (e.g., ('wall', 'surface', 'building exterior'), ('bottoms', 'pants', 'skirt')) may have similar embedding vector values. For example, the electronic device (100) can calculate the cosine similarity between the calculated vector and the vector corresponding to the object. Through this configuration, the electronic device (100) can calculate the similarity between the first user input and the second user input to extract the request information that is semantically most relevant.
[0114] For example, if the first user input is "Draw a fireworks symbol on this wall and draw these pants as jeans," the electronic device (100) can evaluate the semantic similarity between the phrase "draw a fireworks symbol" and the object "wall." At this time, the phrases "fireworks symbol" and "draw" have a high degree of association with "add visual effect," "surface," "scene," etc., and can have a high degree of semantic similarity with the object "wall." Accordingly, the electronic device (100) can identify the phrase "draw a fireworks symbol" that follows, along with the first section "on this wall," as request information associated with the first object (wall).
[0115] For example, the electronic device (100) can extract keywords such as "jeans" and "draw" from the phrase "draw jeans" and evaluate the semantic similarity with the "pants" object so that the phrase is identified as request information associated with the second object (pants). In this way, the electronic device (100) can analyze the contextual meaning of each phrase within the input text and extract and map the most relevant request information based on the semantic similarity with the object designated through the second user input. A detailed description of identifying request information associated with the object designated as the second user input among the first user inputs based on semantic similarity is provided below in FIGS. 9 to 14.
[0116] According to one embodiment, the electronic device (100) may determine whether to integrate multiple request information into a single request information or separate them into separate request information when multiple request information is associated with the same object. For example, a user may sequentially input request information that is contradictory or conflicting with one another regarding the same object. For example, a user may first say, "Make the jeans look three-dimensional," and then say, "Remove the three-dimensional effect of the jeans." The electronic device (100) may analyze the contradictory request information and estimate the user's intention. For example, the electronic device (100) may interpret that the user has ultimately changed their intention and generate content based on the latter request information. As another example, if the latter request information is estimated to be an unintended request, the electronic device (100) may generate content based on the former request information. As another example, the electronic device (100) may interpret two requests as having the intention of being compared or expressed in parallel, generate content based on each of the former and latter request information, and display it sequentially or in parallel.
[0117] According to one embodiment, in operation 430, the electronic device (100) can generate at least one prompt associated with each of the first object and the second object based on one or more request information associated with each of the first object and the second object. The electronic device (100) can provide the generated prompt to an AI model to automatically generate object-specific content (e.g., fireworks display, jeans). The electronic device (100) can generate new data in which the content is reflected in the original data (e.g., an image of a fireworks display on a wall, an image of pants drawn as jeans).
[0118] FIG. 5 is a drawing (500) for explaining a method for creating content of an electronic device according to one embodiment of the present disclosure.
[0119] An operation described in this disclosure as being performed by an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 140 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 5 that overlaps with FIG. 1 through FIG. 4 may be omitted or briefly described.
[0120] According to one embodiment, an electronic device (100) may acquire first data (210) comprising a plurality of objects (e.g., a first object (212) (e.g., a wall) and a second object (214) (e.g., trousers)). The electronic device (100) may display the first data (210) through a display (e.g., 110 in FIG. 1, 2520 in FIG. 25, 2660 in FIG. 26).
[0121] According to one embodiment, the electronic device (100) may receive a first user input (510) through an input module (e.g., 120 in FIG. 1, 2650 in FIG. 26). The first user input (510) may include request information (512) for creating content associated with a first object (e.g., "Draw a fireworks mark on this wall") and request information (514) for creating content associated with a second object (e.g., "Draw these pants as jeans").
[0122] According to one embodiment, the electronic device (100) may receive a second user input (520, 522) through a display. The second user input (520, 522) may include a second user input (520) for specifying a first object (212) (e.g., a wall) among a plurality of objects included in the first data, and a second user input (522) for specifying a second object (214) (e.g., pants).
[0123] FIG. 6 is a drawing showing a graph (600) in which a first user input and a second user input are time-synchronized according to one embodiment of the present disclosure.
[0124] An operation described in this disclosure as being performed by an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 140 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 6 that overlaps with FIG. 1 through 5 may be omitted or briefly described.
[0125] In the graph (600) of FIG. 6, the horizontal axis may represent the reception times of the first user input (610) and the second user input (620, 622) that are time-synchronized. For example, the electronic device (100) may align each input based on the time when the first user input (610) is first received. Through this, the electronic device (100) can identify relative time relationships, such as the start time of reception, the overlap time, and / or the completion time of reception of each input over the entire time flow.
[0126] According to one embodiment, the first user input (610) may include request information such as a user's voice utterance (e.g., "Draw a fireworks symbol on this wall, and draw these pants as jeans"). At the same time, the electronic device (100) may receive a second user input (620, 622) (e.g., a circular gesture for selecting or designating a specific object (e.g., a wall, pants) on a display). However, the method and / or time at which the electronic device (100) receives the first user input and the second user input, respectively, are not limited thereto.
[0127] According to various embodiments of the present invention, an electronic device (100) may receive a first user input in the form of text input via a keyboard. For example, the electronic device (100) may receive a second user input (620, 622) that designates at least one object and / or area through a gesture of clicking at least one object and / or at least one area on a display for a specified time. For example, the electronic device (100) may receive a second user input (620, 622) that designates at least one object and / or area through gaze information of a user gazing at at least one object and / or at least one area via a camera. For example, the electronic device (100) may receive a second user input (620) that designates a first object among at least one object, and after a certain time (e.g., T2) has elapsed, receive a second user input (622) that designates a second object among at least one object.
[0128] According to one embodiment, the electronic device (100) can identify a first section of the first user input of operation 410 of FIG. 4 based on a first user input and a second user input that are time-synchronized. For example, the electronic device (100) can receive a section of the first user input, "on this wall," during time T1, and simultaneously receive a second user input (620) that designates a first object (e.g., a wall). For example, the electronic device (100) can identify a section of the first user input received during time T1 (e.g., "on this wall") as a first section of the first user input (e.g., a first section of the first user input of operation 410 of FIG. 4). For example, the electronic device (100) receives the segment "draw a fireworks display" among the first user input (610) during time T2, and can identify the segment as request information that is temporally and / or semantically associated with the first segment (e.g., one or more request information associated with the first object of operation 410 of FIG. 4). A detailed description of the method for identifying request information that is temporally associated with the first segment is described below in FIG. 7. A detailed description of the method for identifying request information that is semantically associated with the first segment is described below in FIG. 9 and FIG. 10.
[0129] According to one embodiment, the electronic device (100) can identify a second segment of the first user input of operation 420 of FIG. 4 based on a first user input (610) and a second user input (620, 622) that are time-synchronized. For example, the electronic device (100) can receive a segment of the first user input saying "this pants" during time T3, and simultaneously receive a second user input (622) specifying a second object (e.g., pants). For example, the electronic device (100) can identify a segment of the first user input received during time T3 (e.g., "this pants") as a second segment of the first user input (e.g., the second segment of the first user input of operation 420 of FIG. 4). For example, the electronic device (100) receives a segment of the first user input (610) during time T4, which is "draw jeans," and can identify the segment as request information that is temporally and / or semantically associated with the second segment (e.g., one or more request information associated with the second object of operation 420 of FIG. 4). A detailed description of the method for identifying the request information that is temporally associated with the second segment is described below in FIG. 7. A detailed description of the method for identifying the request information that is semantically associated with the second segment is described below in FIG. 9 and FIG. 10.
[0130] According to one embodiment, the electronic device (100) can more precisely identify user request information corresponding to each object. Subsequently, the electronic device (100) can generate a prompt based on the identified user request information (e.g., generating at least one prompt associated with each of the first object and the second object of operation 430 of FIG. 4). A detailed description regarding the generation of such prompts is provided later in FIG. 8.
[0131] Through this configuration, the electronic device (100) can identify the reception time, reception order, and / or overlapping interval of each of the first user input and the second user input, and accurately distinguish request information for each object to perform content generation that matches the user's intention.
[0132] FIG. 7 is a drawing for illustrating the identification of request information associated with each object according to one embodiment of the present disclosure.
[0133] An operation described in this disclosure as being performed by an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 140 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 7 that overlaps with FIG. 1 through 6 may be omitted or briefly described.
[0134] According to one embodiment, the electronic device (100) can identify a segment of the first user input (e.g., 510 in FIG. 5) that corresponds temporally to each object designated by the second user input (e.g., 520, 522 in FIG. 5). For example, the electronic device (100) can identify a first segment (722) of the first user input (e.g., on this wall) received while receiving the second user input (e.g., 520 in FIG. 5) for designating the first object (710) (e.g., a wall). For example, the electronic device (100) can identify a second segment (732) of the first user input (e.g., these pants) received while receiving the second user input (e.g., 522 in FIG. 5) for designating the second object (712) (e.g., pants) (e.g., 522 in FIG. 5) received while receiving the second user input (e.g., 522 in FIG. 5) for designating the second object (712) (e.g., pants).
[0135] According to one embodiment, the electronic device (100) can identify request information associated with a segment of a first user input (e.g., 510 in FIG. 5) that corresponds temporally to each object. For example, the electronic device (100) can identify request information associated with each of the first segment (722) and the second segment (732) based on the temporal relationship in which each of the first user input (e.g., 510 in FIG. 5) and the second user input (e.g., 520, 522 in FIG. 5) is received.
[0136] According to one embodiment, the electronic device (100) may determine, based on the temporal relationship in which the first user input and the second user input are each received, that request information associated with the first object (710) is likely to have been entered before the second user input (e.g., 522 in FIG. 5) for specifying the second object (712) is received. This is because it can be interpreted that the request information for the object was uttered together with the utterance before or immediately after the time of the user specifying the first object.
[0137] In one embodiment, the electronic device (100) may receive a first user input (720, 730) (e.g., draw a fireworks symbol on this wall, draw these pants as jeans) over a plurality of time intervals. For example, the electronic device (100) may receive the interval (722) of the first user input "on this wall" during time T1, and then receive the interval (724) of "draw a fireworks symbol" during time T2. For example, the electronic device (100) may receive the interval (732) of the first user input "these pants" during time T3, and then receive the interval (734) of "draw jeans" during time T4.
[0138] In one embodiment, the electronic device (100) may receive a second user input for specifying each of a plurality of objects in a time-divided manner. For example, the electronic device (100) may receive a second user input for specifying a second object (712) after receiving a second user input for specifying a first object (710). For example, the electronic device (100) may receive a second user input for specifying the first object (710) (e.g., a wall) at time T1. For example, the electronic device (100) may receive a second user input for specifying the second object (712) (e.g., pants) at time T3.
[0139] According to one embodiment, the electronic device (100) can identify request information associated with each of the first object (710) and the second object (712) from the first user input (720, 730) based on the time at which the first user input and the second user input are each received. Here, the request information may include natural language phrases (e.g., "Draw a fireworks symbol," "Change to jeans") that instruct the user to perform an action (e.g., add visual effects, change styles, etc.) on a specific object.
[0140] For example, the electronic device (100) may identify a third segment (720) of the first user input received before specifying the second object (e.g., times T1 and T2), and may identify the third segment as request information associated with the first object (710). For example, the electronic device (100) may identify a fourth segment (732) of the first user input received during at least one of the following times: (i) the time when the second user input specifying the second object (712) is received (e.g., time T3), (ii) the time specified after receiving the second user input specifying the second object (712) (e.g., time T4), and (iii) before receiving the second user input specifying the third object included in the first data (not shown). For example, the electronic device (100) may identify the fourth segment (732) as request information associated with the second object.
[0141] Through this configuration, the electronic device (100) can utilize temporal synchronization information to identify request information associated with a temporally corresponding interval to each object designated as a second user input among the first user inputs, and by accurately matching the request information to each object, it can perform the content creation task intended by the user.
[0142] FIG. 8 is a drawing (800) for explaining a method of generating a prompt associated with each object according to one embodiment of the present disclosure.
[0143] An operation described in this disclosure as being performed by an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 140 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 8 that overlaps with FIG. 1 through 7 may be omitted or briefly described.
[0144] According to one embodiment, if one or more request information (e.g., 720, 730 in FIG. 7) associated with at least one object (e.g., 212, 214 in FIG. 2; 212, 214 in FIG. 5; 710, 712 in FIG. 7) is identified, the electronic device (100) can generate a prompt (810, 820) associated with the object (710, 712) based on this information. For example, the electronic device (100) can generate a prompt optimized to match the user's intent by considering not only the user's explicit request information but also additional conditions (e.g., masking information, style, color).
[0145] According to one embodiment, the electronic device (100) can identify one or more request information associated with the first object (710) (e.g., a wall) (e.g., draw a fireworks display on this wall). For example, if a user explicitly requests "draw a fireworks display on this wall," the electronic device (100) can generate a prompt (810) (e.g., Prompt 1. Create an image[1] to match the description: Add a firework display effect on the wall section of the provided image. Ensure that the fireworks effect is vibrant and well-integrated into the scene. Use realistic lighting and color blending to maintain a natural appearance.) including detailed conditions as well as the request content.
[0146] According to one embodiment, the electronic device (100) can identify one or more request information (e.g., draw these pants as jeans) associated with a second object (712) (e.g., pants). The electronic device (100) can generate a prompt (820) (e.g., Prompt 2. Create an image[2] to match the description: Modify the pants in the provided image to appear as blue denim jeans. Ensure realistic fabric texture, natural folds, and appropriate shading. Retain the original fit and position of the pants while applying the denim texture. Do not modify any other elements in the image.) that includes additional conditions (e.g., jeans texture, wrinkle shape, fit) as well as what the user explicitly requested.
[0147] Through this configuration, the electronic device (100) can precisely identify content creation request information associated with each of at least one object and location information of the object by analyzing a combination of the first user input and the second user input. The electronic device (100) can generate a prompt that more accurately reflects the user's intention and generate content that specifically and precisely reflects the requirements for a specific object.
[0148] FIG. 9 is a drawing for explaining a method for creating content of an electronic device according to one embodiment of the present disclosure.
[0149] An operation described in the present disclosure as being performed by an electronic device (e.g., electronic device (100) of FIG. 1, electronic device (2510) of FIG. 25, or electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 140 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively.
[0150] According to one embodiment, the electronic device (100) may perform a content creation method (900). The order of operations described below in relation to FIG. 9 is an example, and the embodiments of the present disclosure are not limited thereto. For example, at least some of the operations may be executed differently from the order of FIG. 9 or may be executed substantially simultaneously with other operations of FIG. 9. In the following description of FIG. 9, content that overlaps with FIG. 1 through 8 may be omitted or briefly described.
[0151] According to one embodiment, in operation 910, the electronic device (100) may identify each of at least one first segment received while receiving a second user input for specifying each of at least one object among the first user inputs as a segment that corresponds in time to each of at least one object. Here, "corresponds in time" may mean a case where a portion of the first user input is received simultaneously during the time when the second user input for specifying a specific object is received. This may serve as a criterion for determining that the segment is highly likely to be associated with the object.
[0152] For example, a user may say, as a first user input, "Draw a firework symbol on this wall and draw these pants as jeans," and select 'pants' as a second user input while saying "on this wall," and select 'wall' as a second user input while saying "these pants." In this case, the electronic device (100) can identify the section "on this wall" as a section corresponding temporally to pants. Additionally, it can identify the section "these pants" as a section corresponding temporally to the wall.
[0153] According to one embodiment, in operation 920, the electronic device (100) can calculate a first similarity between each of at least one first interval corresponding to each of at least one object and each of at least one object. The electronic device (100) can determine whether there is a high semantic association by calculating the first similarity between a pair of temporally corresponding intervals and objects. Here, the first similarity may be calculated based on the degree of semantic agreement between the request information included within the temporally corresponding interval and the attributes or characteristics of the object, keyword relevance, or the degree of matching between predefined tags. For example, the first similarity may represent the direct similarity with the interval that was uttered simultaneously when the user points to the object. Through the first similarity, the electronic device (100) can determine whether there is a high semantic association between a pair of temporally corresponding intervals and objects.
[0154] According to one embodiment, the electronic device (100) can convert each phrase of the first user input into an embedding vector using a natural language processing (NLP) model (e.g., BERT, and / or Word2Vec). For example, the electronic device (100) can calculate the cosine similarity between an embedding vector of a specific segment within the first user input (e.g., "on this wall") and a predefined tag vector associated with the object (e.g., an embedding vector of the word "wall").
[0155] According to one embodiment, the electronic device (100) has temporal similarity ( )(or degree of temporal correspondence, temporal association) and semantic similarity( Combine ) to integrate similarity( ) can be calculated. For example, the electronic device (100) can calculate integrated similarity by integrating temporal similarity and semantic similarity in a weighted average manner.
[0156]
[0157] Here, represents a weight indicating the importance of temporal similarity (e.g., between 0 and 1), and can represent a weight (e.g., 0 or greater, 1 or less) indicating the importance of semantic similarity.
[0158] For example, the electronic device (100) can calculate the similarity between at least one first interval of the first user input (e.g., "on this wall") and the object corresponding thereto in time (e.g., pants). Additionally, the electronic device (100) can calculate the similarity between another of at least one first interval (e.g., "these pants") and the object corresponding thereto in time (e.g., wall).
[0159] According to one embodiment, in operation 930, if the calculated first similarity is greater than or equal to a specified threshold value, the electronic device (100) can identify request information associated with each of at least one first interval corresponding temporally to each of at least one object as one or more request information associated with each of at least one object.
[0160] According to one embodiment, the electronic device (100) can determine that there is a high semantic association between the interval and the specified object if the calculated cosine similarity value is 0.88 and the calculated similarity value is greater than or equal to a specified threshold (e.g., 0.75 or 0.8).
[0161] According to one embodiment, if the integrated similarity value between a specific section of a first user input and an object designated as a second user input is greater than or equal to a specified threshold value, the electronic device (100) may determine that a specific section (or request information including the specific section) is semantically strongly associated with the designated object.
[0162] For example, a user may say with a first user input, "Draw a firework symbol on this wall, and draw these pants as jeans," and select 'wall' with a second user input while saying "on this wall," and select 'pants' with a second user input while saying "pants." In this case, the electronic device (100) may calculate that the similarity between a specific interval (e.g., "on this wall") and the object ('wall') that corresponds to it in time is greater than a specified threshold. Additionally, the electronic device (100) may calculate that the similarity between a specific interval (e.g., "pants") and the object ('pants') that corresponds to it in time is greater than a specified threshold.
[0163] According to one embodiment, the similarity between a pair of temporally corresponding segments and objects (e.g., "on this wall" and "pants," "these pants" and "wall") may be less than a specified threshold. The electronic device (100) may correct by fixing the object specified by the second user input and finding another segment within the first user input to compare with the specified object. Additionally, the electronic device (100) may correct by fixing the segment that is temporally corresponding to the specific object in the first user input and finding another object specified by the second user input to compare with each other. Furthermore, if there is a discrepancy between a pair of temporally corresponding segments and objects, the electronic device (100) may request feedback from the user, such as new request information or new object designation input.
[0164] According to one embodiment, in operation 930, if the calculated first similarity is less than a specified first threshold, the electronic device (100) can identify one or more first segments that do not correspond temporally to each of at least one object among at least one first segment. The electronic device (100) can calculate a second similarity between each of at least one object and one or more first segments that do not correspond temporally to each of at least one object. Here, the second similarity may represent the semantic correlation (or second similarity, secondary similarity, indirect similarity) between a segment uttered when a user specifies another object and a specific object when the semantic correlation (or first similarity, primary similarity, direct similarity) between segments uttered simultaneously when a user specifies a specific object is low. If the second similarity is greater than or equal to a specified threshold, the electronic device (100) can identify at least one request information associated with one or more first segments as one or more request information associated with each of at least one object.
[0165] According to one embodiment, a first section of at least one first user input may represent at least one partial section of the first user input received while a second user input for specifying each of at least one object is received.
[0166] For example, a first similarity between a first object (e.g., pants) specified by a second user input for specifying a first object (e.g., pants) and a first interval (e.g., on this wall) that corresponds to it in time may be less than a specified threshold. At least one first interval (e.g., "these pants") that does not correspond in time to the first object (e.g., pants) can be identified. The electronic device (100) can calculate a second similarity between the identified first interval and the first object (e.g., pants).
[0167] According to one embodiment, in operation 930, the electronic device (100) can identify at least one second segment other than at least one first segment in the first user input when the calculated first similarity is less than a specified first threshold value. For example, the electronic device (100) can calculate a third similarity between at least one second segment and each of at least one object. Here, the third similarity may mean a similarity (also, third similarity) that determines the semantic association between the second segment and the specific object by searching for any other arbitrary segment (hereinafter, second segment) within the first user input when request information that is semantically similar or associated with the specific object is not identified in the first segment corresponding temporally to the time when the user specifies the specific object.
[0168] For example, if the calculated third similarity is greater than or equal to a specified third threshold, the electronic device (100) may identify at least one request information associated with at least one second interval as one or more request information associated with each of at least one object. At least one second interval may represent a portion of the first user input that is not at least one first interval. That is, the second interval may mean a portion of the first user input received at a time when the second user input is not received.
[0169] For example, a user may say "Draw a firework symbol on this wall, draw jeans on these pants, and draw a rose on this vase" as a first user input, and perform a circular gesture selecting 'vase' as a second user input while uttering "on this wall," and perform a circular gesture selecting 'pants' as a second user input while uttering "these pants." In this case, at least one first interval may be "on this wall" and "these pants," and the objects specified by the second user input may be 'vase' and 'pants'. Additionally, temporally corresponding interval / object pairs may be (on this wall / vase) and (these pants / pants). The similarity of the latter pair may be greater than or equal to a specified threshold, and the similarity of the former pair may be less than a specified threshold.
[0170] For example, in the case of the above pair of electronic devices, the electronic device (100) may not identify a first segment (e.g., on this wall, these pants) in which the similarity to the second user input with the specified object (e.g., vase) is greater than a specified threshold. In this case, the device may search through at least one second segment (e.g., "draw a fireworks mark," "draw it," "with jeans," "draw it," "in this vase," "a rose," "draw it") among the first user inputs to identify a second segment (e.g., in this vase, a rose) in which the similarity to the second user input with the specified object (e.g., vase) is greater than a specified threshold.
[0171] According to one embodiment, in operation 940, the electronic device (100) can generate at least one prompt associated with each of at least one object based on one or more identified request information.
[0172] The electronic device (100) can identify request information associated with each of at least one object by primarily utilizing the temporal correspondence relationship between a first user input (e.g., voice, text, etc.) and a second user input (e.g., a circular gesture for designating an object), and secondarily utilizing the semantic similarity relationship. Through this configuration, the electronic device (100) can perform immediate analysis of the user's intuitive input method and, in cases where there is no semantic match, correct the request information to generate optimized content that matches the user's intent.
[0173] Through this configuration, the electronic device (100) can associate an object designated as a second user input with a segment identified in the first user input based on a temporal correspondence, and additionally consider semantic similarity to more accurately reflect the user's intention.
[0174] Through this configuration, the electronic device (100) can consider the immediate association between the user's first user input (voice or text) and second user input (object designation input) by prioritizing the use of temporal correspondence. If a specific segment of the first user input is received together with the second user input while the second user input is being received, it can be presumed that the specific segment is likely to be semantically associated with the object.
[0175] According to one embodiment, if the semantic association may be low based solely on the temporal correspondence, the electronic device (100) can perform additional analysis to more accurately correct the user's request information. If the temporally corresponding interval does not semantically match the object, the electronic device (100) can correct it by searching for and comparing another interval in the first user input.
[0176] For example, if there is a discrepancy between the word order of the first user input and the object designation of the second user input (e.g., saying "draw a firework symbol on this wall and draw these pants as jeans," and designating 'wall' while saying "firework symbol"), the similarity between the first user input (1110) and the second user input (1120) in a specified time interval (e.g., T1, T3) may not reach a preset threshold. A detailed explanation of this is provided later in FIGS. 11 and 12.
[0177] For example, if the order in which objects are specified by the second user input (1120, 1122) is changed (e.g., pants are specified at time T1 and a wall is specified at time T3), the similarity between the first user input (1110) and the second user input (1120) in the specified time interval (e.g., T1, T3) may not reach a preset threshold. A detailed explanation of this is provided later in FIGS. 13 and 14.
[0178] For example, if the object to be created identified by the first user input and the object specified by the second user input are different (e.g., specifying a 'vase' and / or a 'person face' while saying, "Draw a firework mark on this wall and draw these pants as jeans"), the similarity between the first user input (1110) and the second user input (1120) in a specified time interval (e.g., T1, T3) may not reach a preset threshold. A detailed explanation of this is provided later in FIGS. 15 and 16.
[0179] Through this configuration, the electronic device (100) can rapidly analyze request information by utilizing temporal correspondence, while also more precisely reflecting the user's intent through a correction function based on semantic similarity. Additionally, by generating an optimized prompt that an AI model can interpret based on the corrected request information, it is possible to generate accurate content that matches the user's intent. The electronic device (100) can reduce the user's repetitive requests for modification and support an intuitive input method, thereby providing an efficient method for generating content.
[0180] FIG. 10 is a drawing illustrating a method (1000) for identifying request information associated with each of at least one object according to one embodiment of the present disclosure.
[0181] An operation described in this disclosure as being performed by an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 140 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 10 that overlaps with FIG. 1 through 9 may be omitted or briefly described.
[0182] According to one embodiment, the higher the semantic association or similarity between a pair of segments and an object, the thicker the arrow displayed between the pair may be displayed, or a straight line may be displayed instead of a dotted line.
[0183] According to one embodiment, a first object (1010) (e.g., 212 in FIG. 1, 212 in FIG. 5, 710 in FIG. 7) (e.g., a wall) designated by a second user input (e.g., 520 in FIG. 5) and a first interval (1022) of the first user input (e.g., on this wall) that corresponds temporally to it may have a similarity greater than a specified threshold value. Accordingly, request information (1020) associated with the first interval (1022), including the first interval (1022) (e.g., on this wall) and the associated intervals (1024, 1026) (e.g., drawing a fireworks display), may be identified as request information associated with the first object (1010) (e.g., a wall). The electronic device (100) may generate a prompt associated with the first object (1010) (e.g., a wall) based on the identified request information.
[0184] According to one embodiment, a second object (1012) (e.g., 214 in FIG. 1, 214 in FIG. 5, 712 in FIG. 7) (e.g., pants) designated by a second user input (e.g., 522 in FIG. 5) and a first interval (1032) of a first user input that corresponds temporally thereto (e.g., these pants) may have a similarity greater than or equal to a specified threshold. Accordingly, request information (1030) associated with the first interval (1032), including the first interval (1032) (e.g., these pants) and the associated intervals (1034, 1036) (e.g., draw it as jeans), may be identified as request information associated with the second object (1012) (e.g., pants). The electronic device (100) may generate a prompt associated with the second object (1012) (e.g., pants) based on the identified request information.
[0185] FIG. 11 is a drawing showing a graph (1100) in which a first user input and a second user input are time-synchronized according to one embodiment of the present disclosure.
[0186] An operation described in this disclosure as being performed by an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 140 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 11 that overlaps with FIG. 1 through FIG. 10 may be omitted or described briefly.
[0187] According to one embodiment, while receiving a first user input (1110) (e.g., draw a fireworks mark on this wall, draw these pants as jeans), the electronic device (100) may receive a second user input (e.g., a circular gesture for selecting a wall) to specify a first object (e.g., 212 in FIG. 1, 212 in FIG. 5, 710 in FIG. 7, 1010 in FIG. 10) (e.g., a wall) and a second user input (e.g., a circular gesture for selecting pants) to specify a second object (e.g., 214 in FIG. 1, 214 in FIG. 5, 712 in FIG. 7, 1010 in FIG. 10) (e.g., pants).
[0188] In the graph (1100) of FIG. 11, the horizontal axis may represent the reception time of each of the synchronized first user input (1110) and the synchronized second user input (1120, 1122). For example, each of the synchronized first user input (1110) and the synchronized second user input (1120, 1122) may be arranged based on the time when the synchronized first user input was first received.
[0189] According to one embodiment, an electronic device (100) may receive at least a portion of a first user input and a second user input simultaneously at a specific time. For example, at time T1, the electronic device (100) may receive "fireworks display" from the first user input (1110) while simultaneously receiving a second user input (1120) specifying a first object (1010 in FIG. 10) (e.g., a wall). For example, at time T3, the electronic device (100) may receive "these pants" from the first user input (1110) while simultaneously receiving a second user input (1122) specifying a second object (1012) (e.g., pants).
[0190] According to one embodiment, the electronic device (100) may receive at least a portion of a first user input or one of a second user input at a specific time. For example, the electronic device (100) may not receive the second user input while receiving a portion of the first user input (1110) (e.g., "draw on this wall") at time T2. For example, the electronic device (100) may not receive the second user input while receiving a portion of the first user input (1110) (e.g., "draw with jeans") at time T4.
[0191] According to one embodiment, the electronic device (100) can temporally synchronize the first user input (1110) and the second user input (1120, 1122) to perform matching between object designation and request information based on the reception time of each input. For example, if the word order of the first user input (1110) is changed, the similarity between the first user input (1110) and the second user input (1120) received in a specified time interval T1 may not reach a specified threshold value. For example, the electronic device (100) can perform additional semantic similarity analysis to correct the association between input intervals. A detailed explanation of this is provided later in FIG. 12.
[0192] FIG. 12 is a drawing illustrating a method (1200) for identifying request information associated with each of at least one object according to one embodiment of the present disclosure.
[0193] An operation described in this disclosure as being performed by an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 140 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 12 that overlaps with FIG. 1 through FIG. 11 may be omitted or briefly described.
[0194] According to one embodiment, a first object (1210) (e.g., a wall) designated by a second user input (1120 in FIG. 11) and a first interval (1222) of the first user input (e.g., a firework display) that corresponds to it in time may have a first similarity less than a specified threshold. In this case, the electronic device (100) may identify at least one first interval (1232) (e.g., these pants) other than the first interval (1222) (e.g., a firework display) that corresponds to the first object (1210) (e.g., a wall) in time among at least one first interval (1222, 1232). The electronic device (100) may calculate a second similarity between the first object (1210) (e.g., a wall) and the one or more first intervals (1232) (e.g., these pants). The similarity may be less than a specified threshold.
[0195] According to one embodiment, the electronic device (100) can identify at least one second section (1224, 1226, 1234, 1236) other than at least one first section (1222, 1232) among the first user inputs. The third similarity of each of the first object (1210) (e.g., a wall) and at least one second section (1224, 1226, 1234, 1236) can be calculated. The electronic device (100) can identify a second section (1224) (e.g., this wall) in which the third similarity with the first object (1210) (e.g., a wall) is greater than or equal to a specified threshold value. Request information (1220) associated with the second section (1224) (e.g., on this wall) may include the second section (1224) (e.g., on this wall) and associated sections (1232, 1226) (e.g., drawing a fireworks mark). Request information (1220) associated with the second section (1224) (e.g., on this wall) may be identified as request information associated with the first object (1210) (e.g., wall). The electronic device (100) may generate a prompt associated with the first object (1210) (e.g., wall) based on the identified request information.
[0196] Through this configuration, if the word order is changed or if a low similarity is calculated within a specified time interval, the electronic device (100) can perform an additional similarity check and, based on the result, perform an appropriate matching between object and request information. Through this, the electronic device (100) can correctly identify a phrase such as "fireworks display" as request information originally associated with the first object (1210) (e.g., wall).
[0197] According to one embodiment, a second object (1212) (e.g., trousers) designated by a second user input (1122 in FIG. 11) and a first interval (1232) of a first user input that corresponds temporally to it (e.g., these trousers) may have a first similarity greater than or equal to a specified threshold. For example, an electronic device (100) may identify request information (1230) associated with the first interval (1232) as request information associated with the second object (1212) (e.g., trousers). The electronic device (100) may generate a prompt associated with the second object (1212) (e.g., trousers) based on the identified request information.
[0198] FIG. 13 is a drawing showing a graph (1300) in which a first user input and a second user input are time-synchronized according to one embodiment of the present disclosure.
[0199] An operation described in this disclosure as being performed by an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 130 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 13 that overlaps with FIG. 1 through FIG. 12 may be omitted or described briefly.
[0200] According to one embodiment, the order of the first user input (1310) (e.g., user utterance) and the second user input (1320, 1322) (e.g., object designation) may not always be the same order. For example, when a user makes the utterance "Draw a firework symbol on this wall and draw these pants as jeans," the first object (wall) and the second object (pants) are generally designated sequentially, but in some cases, the order of utterance and the order of object designation may be reversed.
[0201] According to one embodiment, while receiving a first user input (1310) (e.g., draw a fireworks mark on this wall and draw these pants as jeans), the electronic device (100) may receive a second user input (1320) for specifying a first object (e.g., a wall) (e.g., a circular gesture for selecting a wall) and a second user input (1322) for specifying a second object (e.g., a wall) (e.g., a circular gesture for selecting pants).
[0202] The graph (1300) of FIG. 13 may represent a graph in which the first user input (1310) and the second user input (1320, 1322) are synchronized in time. In the graph (1300) of FIG. 13, the horizontal axis may represent the reception time of each of the synchronized first user input (1310) and the synchronized second user input (1320, 1322). For example, each of the synchronized first user input (1310) and the synchronized second user input (1320, 1322) may be arranged based on the time when the synchronized first user input was first received. For example, the electronic device (100) may identify the start time and completion time of reception of each input based on the synchronized first user input (1310) and the second user input (1320, 1322).
[0203] According to one embodiment, an electronic device (100) may simultaneously receive at least a portion of a first user input and a second user input at a specific time. For example, at time T1, the electronic device (100) may receive "on this wall" from the first user input (1310) while simultaneously receiving a second user input (1320) specifying a first object (e.g., pants). For example, at time T3, the electronic device (100) may receive "these pants" from the first user input (1310) while simultaneously receiving a second user input (1322) specifying a second object (e.g., a wall).
[0204] According to one embodiment, the electronic device (100) may receive at least a portion of the first user input or one of the second user inputs at a specific time. For example, the electronic device (100) may not receive the second user input while receiving a portion of the first user input (1310) (e.g., draw a fireworks symbol) at time T2. For example, the electronic device (100) may not receive the second user input while receiving a portion of the first user input (1310) (e.g., draw jeans) at time T4.
[0205] According to one embodiment, the electronic device (100) can temporally synchronize the first user input (1310) and the second user input (1320, 1322) and perform matching between object designation and request information based on the reception time of each input. For example, if the order of object designation of the second user input (1320, 1322) is changed, the similarity between the first user input (1310) and the second user input (1320, 1322) received in the specified time intervals T1 and T3 may not reach a specified threshold value. For example, the electronic device (100) can perform additional semantic similarity analysis to correct the association between the input intervals. A detailed explanation of this is provided later in FIG. 14.
[0206] FIG. 14 is a drawing illustrating a method (1400) for identifying request information associated with each of at least one object according to one embodiment of the present disclosure.
[0207] An operation described in this disclosure as being performed by an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 130 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 14 that overlaps with FIG. 1 through FIG. 13 may be omitted or described briefly.
[0208] According to one embodiment, a first object (1410) (e.g., pants) designated by a second user input (1320 in FIG. 13) and a first interval (1422) of the first user input that corresponds temporally to it (e.g., on this wall) may have a first similarity less than a specified threshold value. In this case, the electronic device (100) may identify one or more first intervals (1432) (e.g., these pants) other than the first interval (1422) (e.g., on this wall) that corresponds temporally to the first object (1410) (e.g., pants) among at least one first interval (1422, 1432). For example, the electronic device (100) may calculate a second similarity between the first object (1410) and the one or more first intervals (1432). For example, if the calculated second similarity is greater than or equal to a specified threshold, the electronic device (100) may identify request information (1430) associated with one or more first segments (1432) as request information associated with the first object (1410). For example, the electronic device (100) may generate a prompt associated with the first object (1410) based on the identified request information.
[0209] According to one embodiment, a second object (1412) (e.g., a wall) designated by a second user input (1322 in FIG. 13) and a second interval (1432) of the first user input that corresponds temporally to it (e.g., these pants) may have a first similarity less than a specified threshold value. In this case, the electronic device (100) may identify one or more first intervals (1422) (e.g., these wall) among at least one first interval (1422, 1432) other than the first interval (1432) (e.g., these pants) that corresponds temporally to the second object (1412). The electronic device (100) may calculate a second similarity between the second object (1412) and the one or more first intervals (1422). For example, if the second similarity is greater than or equal to a specified threshold, the electronic device (100) may identify request information (1420) associated with one or more first segments (1422) as request information associated with the second object (1412). For example, the electronic device (100) may generate a prompt associated with the second object (1412) based on the identified request information.
[0210] According to one embodiment, the electronic device (100) can convert the semantic representation of each segment (1422, 1424, 1426, 1432, 1434, 1436) within the first user input into an embedding vector through a natural language processing model (e.g., BERT or Sentence-BERT, etc.), and calculate the similarity by calculating the cosine similarity between this vector and the embedding vector of predefined tag or attribute information for the corresponding object. For example, the electronic device (100) can perform an additional similarity calculation method (e.g., second similarity, third similarity) to further compare and correct the similarity between candidate segments based on the comparison result between the calculated similarity and a specified threshold value (e.g., 0.75 or 0.8).
[0211] According to one embodiment, even when the order of speech and the order of object designation are reversed, the electronic device (100) can perform correct object-request information matching through temporal synchronization and semantic similarity correction. For example, even if a user speaks, "Draw a firework symbol on this wall and draw these pants as jeans," but the order of the second user input is reversed, the electronic device can correct the association of "draw a firework symbol" and "draw jeans" with appropriate objects (wall or pants) respectively through the reception time of each input and semantic similarity analysis.
[0212] Through this configuration, the electronic device (100) can integrate temporal synchronization information and semantic similarity analysis to correct and re-verify the immediate association between the user's intuitive input and object designation input. The electronic device (100) can generate an optimized prompt to perform content creation tasks that match the user's intent.
[0213] FIG. 15 is a drawing for explaining a method for creating content according to one embodiment of the present disclosure.
[0214] An operation described in this disclosure as being performed by an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 130 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 15 that overlaps with FIG. 1 through FIG. 14 may be omitted or briefly described.
[0215] According to one embodiment, the electronic device (100) may receive a first user input saying, "Draw a fireworks mark on this wall, and draw these pants as jeans." While receiving the first user input, the electronic device (100) may receive a circular gesture as a second user input specifying a first object (1510) (e.g., pants) while saying "on this wall."
[0216] According to one embodiment, the similarity between a first object (1510) (e.g., pants) designated as a second user input and a first interval of the first user input ("on this wall") that corresponds to it in time may be less than a specified threshold value. A semantic discrepancy may occur between the object intended by the user and the request information.
[0217] Referring to 1500, according to one embodiment, the electronic device (100) can further identify an object (1520) (wall surface) that has a similarity to a first section (this wall surface) of the first user input greater than a specified threshold value based on at least one of the first data, the first user input, or the second user input.
[0218] According to one embodiment, the electronic device (100) may display a guide (1530) (e.g., "Would you like to change the object / area for which you are requesting content creation?") through the display (110) to induce a request for new content creation associated with an additionally identified object (1520) (wall). Through this guide (1530), the electronic device (100) may induce the user to correct an incorrectly designated object or to clearly specify a request for a desired object.
[0219] According to one embodiment, when the electronic device (100) receives user input that changes the target object requesting content creation from a first object (1510) (e.g., pants) to a second object (1520) (e.g., a wall) in response to a guide (1530), it may create content (e.g., a firework display) associated with the second object (1520) (e.g., a wall) based on the first user input. The electronic device (100) may create output data including the created content (e.g., a firework display). The electronic device (100) may display the created content (e.g., a firework display) and / or output data through a display (110).
[0220] Referring to 1502, according to one embodiment, the electronic device (100) may further identify a first segment of the first user input (e.g., "These pants") which has a similarity to the first object (1510) (e.g., pants) greater than a specified threshold value based on at least one of the first data, the first user input, or the second user input.
[0221] According to one embodiment, the electronic device (100) may display a guidance (1532) (e.g., Would you like to change the content creation request for the designated target object / area?) through the display (110) that induces a new content creation request associated with a first section of the additionally identified first user input (e.g., "These pants").
[0222] According to one embodiment, the electronic device (100) may receive a new first user input requesting the creation of new content associated with a designated first object (1510) (e.g., pants) in response to a guide (1532). Based on the new first user input, the electronic device (100) may create content (e.g., jeans) for the first object (1510) (e.g., pants). The electronic device (100) may create output data containing the created content (e.g., jeans). The electronic device (100) may display the created content (e.g., jeans) and / or output data through a display (110).
[0223] Through this configuration, the electronic device (100) can analyze request information based on the temporal correspondence between user inputs (first user input, second user input) and automatically detect if a semantic discrepancy occurs. Additionally, the electronic device (100) can more accurately reflect the user's content creation intention by automatically identifying additional objects or finding a semantically more appropriate section and guiding the user.
[0224] FIG. 16 is a drawing for explaining a method for creating content according to one embodiment of the present disclosure.
[0225] An operation described in this disclosure as being performed by an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 130 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 16 that overlaps with FIG. 1 through FIG. 15 may be omitted or briefly described.
[0226] According to one embodiment, the electronic device (100) may receive a first user input saying, "Draw a fireworks mark on this wall, and draw these pants as jeans." While receiving the first user input, the electronic device (100) may receive a circular gesture as a second user input specifying a first object (1610) (e.g., pants) while saying "on this wall."
[0227] According to one embodiment, the similarity between a first object (1610) (e.g., pants) designated as a second user input and a first interval of the first user input (on this wall) that corresponds to it in time may be less than a specified threshold value. A semantic discrepancy may occur between the object intended by the user and the request information.
[0228] Referring to 1600, according to one embodiment, an electronic device (100) may display a guide (1630) (e.g., "Please re-specify the object / area for which you are requesting content creation") through a display (110) that induces a change in the target object for which content is to be created in connection with a first section of a first user input (e.g., "on this wall"). In response to the guide (1630), the electronic device (100) may additionally receive a new second user input (e.g., a circular gesture for specifying a wall) to specify a new object (1620) (e.g., a wall). Based on the first user input and the additionally received new second user input, the electronic device (100) may create content that corresponds to the user's intention (e.g., a part of a wall with a fireworks symbol drawn on it).
[0229] Referring to 1602, according to one embodiment, an electronic device (100) may display a guide (1632) (e.g., "Please repeat the request to create content for the designated target object / area") through a display (110) that induces a request to create new content associated with a first object (1610) (e.g., pants) designated as a second user input. In response to the guide (1632), the electronic device (100) may additionally receive a new first user input (e.g., "Draw these pants as jeans") requesting the creation of new content associated with the designated first object (1610). Based on the second user input and the additionally received new first user input, the electronic device (100) may create content that matches the user's intention.
[0230] With this configuration, the electronic device (100) can provide an intuitive interface that allows the user to naturally perform correction through guidance (1630, 1632) that induces additional input.
[0231] FIG. 17 is a drawing (1700) illustrating a method for creating content according to one embodiment of the present disclosure.
[0232] An operation described in this disclosure as being performed by an electronic device (100) (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 130 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 17 that overlaps with FIG. 1 through FIG. 17 may be omitted or described briefly.
[0233] According to one embodiment, an electronic device (100) can acquire first data (210_1) including a first object (212_1) (e.g., pants) and a second object (214_1) (e.g., a wall). The electronic device (100) can display the first data (210_1) through a display (110).
[0234] According to one embodiment, an electronic device (100) may receive a first user input (1710) through an input module (e.g., a microphone). The first user input (1710) may include a first request information (1712) for creating content associated with a first object (212_1) (e.g., make the jeans of the person on the right look three-dimensional), a second request information (1714) for creating content associated with a second object (214_1) (e.g., draw a firework effect over the person on the left), and a third request information (1716) for creating content associated with the first object (212_1) (e.g., erase the three-dimensional effect of the person on the right).
[0235] According to one embodiment, the electronic device (100) can receive a second user input (1720, 1722) through a display (110). The electronic device (100) can receive a second user input (1720) for specifying a first object (212_1) (e.g., pants) and a second user input (1722) for specifying a second object (214_1) (e.g., a wall), respectively.
[0236] FIG. 18 is a drawing illustrating a graph (1800) in which a first user input and a second user input are time-synchronized according to one embodiment of the present disclosure.
[0237] An operation described in this disclosure as being performed by an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 130 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 18 that overlaps with FIG. 1 through FIG. 17 may be omitted or briefly described.
[0238] In the graph (1800) of FIG. 18, the horizontal axis may represent the time at which the first user input (1710 of FIG. 17) is received, the first user input (1810) which is time-synchronized with the first user input (1720, 1722 of FIG. 17), and the second user input (1820, 1822, 1824) which is time-synchronized with the second user input (1820, 1822, 1824) which is time-synchronized with the first user input (1810) which is received. For example, the synchronized first user input (1810) and the synchronized second user input (1820, 1822, 1824) may each be arranged based on the time at which the synchronized first user input (1810) is first received.
[0239] According to one embodiment, an electronic device (100) may receive at least a portion of a first user input and a second user input simultaneously at a specific time. For example, at time T1, the electronic device (100) may receive "jeans of the person on the right" among the synchronized first user inputs (1810), while simultaneously receiving a synchronized second user input (1820) specifying a first object (212_1 in FIG. 17) (e.g., pants). For example, at time T3, the electronic device (100) may receive "on the person on the left" among the synchronized first user inputs (1810), while simultaneously receiving a second user input (1822) specifying a second object (214_1 in FIG. 17) (e.g., a wall). For example, at time T5, the electronic device (100) may receive "person on the right" among the synchronized first user inputs (1810), and at the same time receive a second user input (1824) specifying a first object (212_1 in FIG. 17) (e.g., pants).
[0240] According to one embodiment, the electronic device (100) may receive at least a portion of the first user input or one of the second user inputs at a specific time. For example, the electronic device (100) may not receive the second user input while receiving a portion of the synchronized first user input (1810) (e.g., make it look three-dimensional) at time T2. For example, the electronic device (100) may not receive the second user input while receiving a portion of the synchronized first user input (1810) (e.g., draw a firework effect) at time T4. For example, the electronic device (100) may not receive the second user input while receiving a portion of the synchronized first user input (1810) (e.g., erase the three-dimensional effect) at time T6.
[0241] FIG. 19 is a drawing illustrating a method (1900) for identifying request information associated with each of at least one object according to one embodiment of the present disclosure.
[0242] An operation described in this disclosure as being performed by an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 130 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 19 that overlaps with FIG. 1 through FIG. 18 may be omitted or described briefly.
[0243] According to one embodiment, a first object (1910) (e.g., 212_1 in FIG. 17) (e.g., pants) designated by a second user input (e.g., 1720 in FIG. 17, 1820 in FIG. 18) and a first segment (1922, 1924) of the first user input (e.g., jeans of the person on the right) that corresponds temporally to it at a specific time (e.g., T1 in FIG. 18) may have a first similarity greater than or equal to a specified threshold. For example, the electronic device (100) may identify the first request information (1920) associated with the first segment (1922, 1924) as the request information associated with the first object (1910). The electronic device (100) may generate a prompt associated with the first object (1910) based on the identified first request information.
[0244] According to one embodiment, a first object (1910) (e.g., 212_1 in FIG. 17) (e.g., pants) designated by a second user input (e.g., 1720 in FIG. 17, 1824 in FIG. 18) and a first segment (1942) of the first user input (e.g., person on the right) that corresponds temporally at a different specific time (e.g., T5 in FIG. 18) may have a first similarity greater than or equal to a specified threshold. For example, the electronic device (100) may identify a first request information (1940) associated with the first segment (1942) as second request information associated with the first object (1910). The electronic device (100) may generate a prompt associated with the first object (1910) based on the identified second request information.
[0245] According to one embodiment, although not shown in FIG. 19, a second object (214_1 in FIG. 17) (e.g., a wall) designated by a second user input (e.g., 1722 in FIG. 17, 1822 in FIG. 18) and a first segment (1932) of the first user input that corresponds temporally to it (e.g., on a person on the left) may have a similarity greater than a specified threshold. Accordingly, the first segment (1932) and the associated segment (1934) (e.g., drawing a firework effect) can be identified as request information (1930) associated with the first segment (1932). The electronic device (100) can identify the request information (1930) associated with the first segment as request information associated with the second object (214_1 in FIG. 17). The electronic device (100) can generate a prompt associated with a second object (214 in FIG. 17) (e.g., a wall) based on identified request information.
[0246] FIG. 20 is a drawing for explaining how to display content generated according to one embodiment of the present disclosure through a display.
[0247] An operation described in this disclosure as being performed by an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 130 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 20 that overlaps with FIG. 1 through FIG. 19 may be omitted or briefly described.
[0248] According to one embodiment, the electronic device (100) can generate at least one prompt associated with at least one of at least one object based on at least one request information associated with at least one of at least one object. The electronic device (100) can generate content associated with at least one of at least one object based on the prompt generated through a generative AI model. The electronic device (100) can display the generated content through a display.
[0249] According to one embodiment, there may be at least one request information associated with at least one of at least one object. For example, referring to FIG. 19, the request information associated with a first object (212_1 in FIG. 17, 1910 in FIG. 19) (e.g., pants) may include a first request information (1920) (e.g., make the jeans of the person on the right look three-dimensional) and a second request information (1940) (e.g., remove the three-dimensional effect of the person on the right). The electronic device (100) may generate a single prompt or multiple prompts associated with at least one object based on the first request information (1920) and / or the second request information (1940).
[0250] According to one embodiment, the electronic device (100) can determine a content creation method for at least one of each of at least one object based on at least one request information associated with at least one of at least one object. Here, the content creation method may include the number of contents to be created in association with each object.
[0251] According to one embodiment, the electronic device (100) may presume that the user's intention is not to create a single content that integrates multiple request information, but to create multiple content in which each request information is reflected separately. For example, if multiple request information associated with the same object is semantically contradictory, the electronic device (100) may create multiple prompts based on each request information, and may create multiple content associated with the object based on each of the generated prompts.
[0252] According to one embodiment, the electronic device (100) may interpret the user's intention as modifying or replacing previously received request information with the finally received request information. For example, the electronic device (100) may generate content based on the last received request information. For example, if multiple request information associated with the same object and contradictory to each other is received sequentially, the electronic device (100) may generate a single piece of content that reflects the last received request information.
[0253] According to one embodiment, there may be cases where a user inputs multiple request information while explicitly requesting the creation of multiple contents for a single object. For example, the electronic device (100) may create multiple contents with different styles or effects applied based on each request information.
[0254] Through this configuration, the electronic device (100) can determine the optimal content generation method by analyzing the semantic relationship of the request information when multiple request information is associated with the same object. When contradictory request information exists in parallel, individual content based on each request information can be generated to satisfy the user's various intentions. When sequentially entered request information is contradictory, the final request information is reflected first to more accurately reflect the user's intention.
[0255] According to one embodiment, in 2000, the electronic device (100) can generate content (2010) (e.g., firework display) associated with the second object (e.g., 214_1 in FIG. 17) based on request information (1930 in FIG. 19) (e.g., drawing a firework effect over a person on the left) associated with the second object (e.g., 214_1 in FIG. 17).
[0256] The electronic device (100) can generate content (2020) (e.g., jeans with the 3D effect removed) associated with the first object (212_1 in FIG. 17) (e.g., pants), by reflecting all of the first request information (1920 in FIG. 19) (e.g., make the jeans of the person on the right look 3D) and the second request information (1940 in FIG. 19) (e.g., remove the 3D effect of the person on the right). The electronic device (100) can generate second data by reflecting the generated content (2010, 2020) in the first data (210_1 in FIG. 17). The electronic device (100) can display the generated content (2010, 2020) and / or the second data through a display.
[0257] In 2002, the electronic device (100) can generate content (2010) associated with the second object (214_1 in FIG. 17) based on request information (1930 in FIG. 19) associated with the second object (214_1 in FIG. 17) (e.g., adding a fireworks effect over the person on the left).
[0258] According to one embodiment, an electronic device (100) can generate content (2022) (e.g., jeans that appear three-dimensional) associated with a first object (212_1 in FIG. 17) (e.g., pants) by reflecting only the first request information (1920 in FIG. 19) (e.g., making the jeans of the person on the right appear three-dimensional) associated with a first object (212_1 in FIG. 17). For example, the electronic device (100) can generate new second data by reflecting the generated content (2010, 2022) in the first data. The electronic device (100) can display the generated content (2010, 2022) and / or the second data through a display.
[0259] In 2004, the electronic device (100) may display the second data (2030) generated in 2000 and the second data (2032) generated in 2002 together through a display and display a guide (2040) requesting a selection from the user (e.g., "Please select the image you want from the following"). The electronic device (100) may generate or display final content through a display in response to user input selecting one of a plurality of options according to the guide (2040).
[0260] FIG. 21 is a drawing (2100) illustrating a method for creating content according to one embodiment of the present disclosure.
[0261] An operation described in this disclosure as being performed by an electronic device (100) (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 130 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 21 that overlaps with FIG. 1 through FIG. 20 may be omitted or described briefly.
[0262] According to one embodiment, an electronic device (100) can acquire first data (2110) including a first object (2112) (e.g., large sky on the left), a second object (2114) (e.g., small sky on the left), a third object (2116) (e.g., small sky in the middle) and a fourth object (2118) (e.g., small sky on the right). The electronic device (100) can display the first data (2110) through a display (110).
[0263] According to one embodiment, an electronic device (100) may receive a first user input (2120) through an input module (e.g., a microphone). The first user input (2120) may include first request information (2132) for creating content associated with a first object (e.g., draw a red cloud here), and second request information (2134) for creating content associated with second to fourth objects (e.g., draw a purple eye here, here, here).
[0264] According to one embodiment, the electronic device (100) may receive a second user input (2130, 2132, 2134, 2136) through a display (110). The second user input may include a second user input (2130) for specifying a first object (2112) (e.g., large left sky), a second user input (2132) for specifying a second object (2114) (e.g., small left sky), a second user input (2134) for specifying a third object (2116) (e.g., small middle sky), and a second user input (2136) for specifying a fourth object (2118) (e.g., large right sky).
[0265] FIG. 22 is a drawing showing a graph (2200) in which a first user input and a second user input are time-synchronized according to one embodiment of the present disclosure.
[0266] An operation described in this disclosure as being performed by an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 130 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 22 that overlaps with FIG. 1 through 21 may be omitted or briefly described.
[0267] In the graph (2200) of FIG. 22, the horizontal axis may represent the time at which the first user input (2120 of FIG. 21) is received, the first user input (2210) which is time-synchronized with the first user input (2130, 2132, 2134, 2136 of FIG. 21), and the second user input (2220, 2222, 2224, 2226) which is time-synchronized with the second user input (2220, 2222, 2224, 2226) which is time-synchronized with the first user input (2210) which is received. For example, the synchronized first user input (2210) and the synchronized second user input (2220, 2222, 2224, 2226) may each be arranged based on the time at which the synchronized first user input (2210) is first received.
[0268] According to one embodiment, the electronic device (100) may receive at least a portion of a first user input and a second user input simultaneously at a specific time. For example, at time T1, the electronic device (100) may receive "here" among the synchronized first user inputs (2210) at the same time, receive a synchronized second user input (2220) specifying a first object (2112 in FIG. 21) (e.g., the large sky on the left). For example, at time T2, the electronic device (100) may receive "here," among the synchronized first user inputs (2210) at the same time, receive a second user input (2222) specifying a second object (2114 in FIG. 21) (e.g., the small sky on the left). For example, at time T3, the electronic device (100) may receive "here" among the synchronized first user inputs (2210) and at the same time receive a second user input (2224) specifying a third object (2116 in FIG. 21) (e.g., middle small sky). For example, at time T4, the electronic device (100) may receive "here" among the synchronized first user inputs (2210) and at the same time receive a second user input (2226) specifying a fourth object (2118 in FIG. 21) (e.g., right small sky).
[0269] FIGS. 23a and FIGS. 23b are drawings (2300, 2302) illustrating a method for identifying request information associated with each of at least one object according to one embodiment of the present disclosure. Hereinafter, any content overlapping with FIGS. 1 to 22 in the description of FIGS. 23a and FIGS. 23b may be omitted or briefly described.
[0270] Referring to FIG. 23a, a first object (2312) (e.g., the large sky on the left) specified by a second user input (2130 in FIG. 21) and a first interval (2324) of the first user input (e.g., here) that corresponds temporally to it may have a similarity greater than a specified threshold. The electronic device (100) can identify request information (2320) associated with the first interval, including the first interval (2324) (e.g., here) and the associated intervals (2322, 2326) (e.g., red clouds, draw). The electronic device (100) can identify the request information (2320) associated with the first interval as request information associated with the first object (2312).
[0271] Referring to FIG. 23b, a second object (2314) (e.g., small sky on the left) designated by a second user input (2132 in FIG. 21) and a first segment (2333) of a first user input (e.g., here) that is temporally corresponding to the second object (2314) may have a similarity greater than a specified threshold. The electronic device (100) can identify request information (2330) associated with the first segment (e.g., draw purple eyes here), including the first segment (2333) and segments (2331, 2339) associated with the first segment (2333) (e.g., draw purple eyes). The electronic device (100) can identify the request information (2330) associated with the first segment as request information associated with the second object (2314).
[0272] FIG. 24 is a drawing for explaining a method for creating content according to one embodiment of the present disclosure.
[0273] An operation described in this disclosure as being performed by an electronic device (100) (e.g., the electronic device (100) of FIG. 1, the electronic device (2510) of FIG. 25, or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (100) to perform the operation by having at least one processor (e.g., 130 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 24 that overlaps with FIG. 1 through FIG. 23b may be omitted or briefly described.
[0274] According to one embodiment, the electronic device (100) can generate a prompt associated with each object based on request information associated with the identified object. Based on the generated prompt, the electronic device (100) can generate content associated with each object.
[0275] Referring to 2400, according to one embodiment, an electronic device (100) can generate a prompt (2410) associated with a first object (e.g., Prompt 1. Create an image[1] to match the description: Add a red cloud effect to the specified area in the provided image. Ensure that the red cloud appears soft and naturally blended into the scene. Use smooth gradients and subtle transparency to maintain a realistic atmospheric effect. Do not modify any other elements in the image.) based on request information (2320 in FIG. 23a) associated with the first object. The electronic device (100) can generate a prompt (2412) associated with the second to fourth objects (e.g., Prompt 2. Create an image[2] to match the description: Add three instances of purple-colored eyes in the specified locations within the provided image. Ensure that each eye appears distinct yet cohesive in style. Maintain proper shading, reflections, and depth for a realistic or artistic effect, based on the context of the image. Do not alter any other areas of the image.) based on request information (2530) associated with the second object (2114 in FIG. 21, 2314 in FIG. 23b), and the fourth object (2118 in FIG. 21).
[0276] Referring to 2402, according to one embodiment, an electronic device (100) can generate content (2410, 2420, 2422, 2424) associated with each object (2122, 2124, 2126, 2128 of FIG. 21) through a generative AI model based on generated prompts (2410, 2412). The electronic device (100) can generate second data by reflecting content (2410) associated with a first object (2122 in FIG. 21) (e.g., red cloud), content (2420) associated with a second object (2124 in FIG. 21) (e.g., purple eye), content (2422) associated with a third object (2126 in FIG. 21) (e.g., purple eye), and content (2424) associated with a fourth object (2128 in FIG. 21) (e.g., purple eye) into the first data (2110 in FIG. 21). Subsequently, the electronic device (100) can display the second data through a display (110).
[0277] FIG. 25 is a drawing (2500) illustrating a method for creating content of an electronic device according to one embodiment of the present disclosure.
[0278] An operation described in this disclosure as being performed by an electronic device (2510) (e.g., the electronic device (100) of FIG. 1 or the electronic device (2601) of FIG. 26) may be understood as causing the electronic device (2510) to perform the operation by having at least one processor (e.g., 130 of FIG. 1, 2620 of FIG. 26) execute instructions stored in memory (e.g., 130 of FIG. 1, 2630 of FIG. 26) individually or collectively. Hereinafter, any content in the description of FIG. 25 that overlaps with FIG. 1 through 24 may be omitted or briefly described.
[0279] According to one embodiment, a user can designate objects and create content in a virtual and / or augmented reality environment by utilizing an electronic device (2510) that includes VR (Virtual Reality), AR (Augmented Reality), or MR (Mixed Reality) functions. Although the following description focuses on a VR environment for convenience, the various embodiments of the present disclosure can be applied equally to AR or MR environments.
[0280] According to one embodiment, the electronic device (2510) can display the first data (2530) through a VR display (2520). The first data (2530) includes a scene visually represented in a virtual environment and may include at least one object (e.g., a first object (2540) (e.g., a wall), a second object (2542) (e.g., pants)) for which a user can request the creation of specific content.
[0281] According to one embodiment, an electronic device (2510) can receive a request to create content for a specific object through a first user input (2550). For example, if a user speaks via voice input, "Draw a fireworks symbol on this wall and draw these pants as jeans," the electronic device (2510) can analyze the voice input and extract content creation request information associated with each object.
[0282] According to one embodiment, the electronic device (2510) may receive a second user input (2560, 2562) for specifying a specific object in a VR environment. For example, if the user gazes at the first object (2540) on the VR screen for a specified time as a second user input (2560) for specifying the first object (2540), the electronic device (2510) may receive the gaze input and specify the first object. For example, if the user gazes at the second object (2542) for a specified time as a second user input (2562) for specifying the second object (2542), the electronic device (2510) may specify the object.
[0283] According to one embodiment, the electronic device (2510) can analyze the received second user input (2560, 2562) to match the request information identified in the first user input (2550) with a specific object. The electronic device (2510) can generate content based on the identified request information. For example, if the request information associated with the first object (2540) (e.g., a wall) is "add firework display," the electronic device (2510) can generate content (2570) with the firework display applied. For example, if the request information associated with the second object (2542) (e.g., pants) is "change to jeans," the electronic device (2510) can generate content (2572) with the texture of the pants changed. The electronic device (2510) can reflect the generated content in the original data (first data (2530)) to generate second data (2532) representing a new VR environment.
[0284] Through this configuration, the electronic device (2510) can select an object by combining the user's voice input and eye tracking input in a VR environment and generate content associated with that object. The electronic device (2510) can precisely match content by analyzing the correlation between voice commands (first user input) and eye tracking (second user input) without direct touch input from the user. The electronic device (2510) can reflect the generated content in real time and, accordingly, receive immediate feedback from the user. The electronic device (2510) can enhance the user experience in a VR environment and provide a more immersive content creation experience.
[0285]
[0286] According to one embodiment, the electronic device may include an input module, a display, a memory for storing instructions, and at least one processor. When the instructions are executed individually or collectively by at least one processor, the electronic device may be configured to: acquire first data containing at least one object; receive, through an input module, a first user input containing at least one request information for generating at least one content associated with at least one of at least one object included in the first data; receive, through an input module, a second user input for specifying at least one of at least one object in the first data; identify one or more request information associated with at least one of at least one object based on at least one of the request information based on at least one request information based on at least one request information based on the identified request information; generate at least one prompt associated with at least one of at least one object based on at least one request information; and generate content associated with at least one of at least one object based on at least one prompt through a generative artificial intelligence (AI) model.
[0287] According to one embodiment, at least one object includes a first object and a second object, and instructions, when executed individually or collectively by at least one processor, can cause an electronic device to: when executed individually or collectively by at least one processor, the electronic device: identify request information associated with a first segment of a first user input received while receiving a second user input for specifying a first object as one or more request information associated with the first object, and identify request information associated with a second segment of a first user input received while receiving a second user input for specifying a second object as one or more request information associated with the second object, and generate at least one prompt associated with the first object and the second object based on one or more request information associated with each of the first object and the second object.
[0288] According to one embodiment, when instructions are executed individually or collectively by at least one processor, the electronic device may be configured to: receive a second user input for specifying a first object, receive a second user input for specifying a second object, identify a third segment of the first user input received before receiving the second user input for specifying the second object as one or more request information associated with the first object, and identify a fourth segment of the first user input received during at least one period, such as while receiving the second user input for specifying the second object, during a specified time after receiving the second user input for specifying the second object, or before receiving the second user input for specifying a third object included in the first data, as one or more request information associated with the second object.
[0289] According to one embodiment, when instructions are executed individually or collectively by at least one processor, the electronic device may: identify at least a portion of a first user input whose similarity to a first section is greater than or equal to a specified threshold value as a portion associated with the first section, and identify at least a portion of a first user input whose similarity to a second section is greater than or equal to a specified threshold value as a portion associated with the second section.
[0290] According to one embodiment, when instructions are executed individually or collectively by at least one processor, the electronic device may: determine a content creation method including the number of contents to be created in association with at least one of at least one object based on one or more request information, and generate at least one prompt associated with at least one of at least one object based on the content creation method and one or more request information.
[0291] According to one embodiment, when instructions are executed individually or collectively by at least one processor, the electronic device may: identify each of at least one first interval received while receiving a second user input for specifying each of at least one object among a first user input as an interval corresponding in time to each of at least one object, calculate a first similarity between each of at least one first interval corresponding in time to each of at least one object and each of at least one object, and if the calculated first similarity is greater than or equal to a specified first threshold value, identify request information associated with each of at least one first interval as one or more request information associated with each of at least one object, and generate each of at least one prompt associated with each of at least one object based on the identified one or more request information.
[0292] According to one embodiment, when instructions are executed individually or collectively by at least one processor, the electronic device may be configured to: identify one or more first intervals that do not temporally correspond to each of at least one object among at least one first interval when the calculated first similarity is less than a specified first threshold value; identify at least one request information associated with each of at least one first interval and at least one object when the second similarity between each of at least one first interval and at least one object is greater than or equal to a specified second threshold value; and generate at least one prompt associated with each of at least one object based on the identified one or more request information.
[0293] According to one embodiment, when instructions are executed individually or collectively by at least one processor, the electronic device may be configured to: identify at least one second segment that is not at least one first segment in the first user input when the calculated first similarity is less than a specified first threshold, calculate a second similarity between each of the at least one second segment and each of at least one object, and when the calculated second similarity is greater than or equal to a specified second threshold, identify at least one request information associated with each of the at least one second segment as one or more request information associated with each of the at least one object, and generate at least one prompt associated with each of the at least one object based on the identified one or more request information.
[0294] According to one embodiment, at least one of voice input, text input, or natural language command may be included, which includes request information comprising at least one of a sentence, word, or phrase requesting the creation of content associated with each of at least one object.
[0295] According to one embodiment, the second user input may include at least one of a touch input, a gesture input, a drag action, or an eye tracking input for specifying each of at least one object in the first data displayed through the display of the electronic device.
[0296] According to one embodiment, a method for generating content of an electronic device may include: acquiring first data comprising at least one object; receiving a first user input through an input module of the electronic device, the first user input comprising at least one request information for generating at least one content associated with at least one of at least one object included in the first data; receiving a second user input through an input module for specifying at least one of at least one object in the first data; identifying one or more request information associated with at least one of at least one object among the at least one request information based on at least one interval of each of the first user input and the second user input received at a time when each of the first user input and the second user input is received or at a specified time; generating at least one prompt associated with at least one of at least one object based on the identified one or more request information; and generating content associated with at least one of at least one object based on at least one prompt through a generative artificial intelligence (AI) model.
[0297] According to one embodiment, at least one object may include a first object and a second object. For example, an operation of identifying one or more request information associated with at least one of the at least one object may include an operation of identifying request information associated with a first segment of a first user input received while receiving a second user input for specifying the first object as one or more request information associated with the first object, and an operation of identifying request information associated with a second segment of a first user input received while receiving a second user input for specifying the second object as one or more request information associated with the second object. For example, an operation of generating at least one prompt associated with at least one of the at least one object may include an operation of generating at least one prompt associated with each of the first object and the second object based on one or more request information associated with each of the first object and the second object.
[0298] According to one embodiment, the operation of receiving a second user input for specifying at least one of at least one object may include the operation of receiving a second user input for specifying a second object after receiving a second user input for specifying a first object. The operation of identifying one or more request information associated with at least one of at least one object may include the operation of identifying a third segment of a first user input received before receiving a second user input for specifying a second object as one or more request information associated with the first object, and the operation of identifying a fourth segment of a first user input received during at least one period among receiving a second user input for specifying a second object, during a specified time after receiving a second user input for specifying a second object, or before receiving a second user input for specifying a third object included in the first data as one or more request information associated with the second object.
[0299] According to one embodiment, the operation of identifying one or more request information associated with at least one of at least one object may include the operation of identifying at least a portion of a first user input, the similarity to a first segment being greater than or equal to a specified threshold value, as a portion associated with the first segment, and the operation of identifying at least a portion of a first user input, the similarity to a second segment being greater than or equal to a specified threshold value, as a portion associated with the second segment.
[0300] According to one embodiment, the operation of generating at least one prompt associated with at least one of at least one object may include the operation of determining a content generation method including the number of contents to be generated associated with at least one of at least one object based on one or more request information, and the operation of generating at least one prompt associated with at least one of at least one object based on the content generation method and one or more request information.
[0301] According to one embodiment, the operation of identifying one or more request information associated with at least one of at least one object may include: identifying each of at least one first interval received while receiving a second user input for specifying each of at least one object among a first user input as an interval corresponding in time to each of at least one object; calculating a first similarity between each of at least one first interval corresponding in time to each of at least one object and each of at least one object; and if the calculated first similarity is greater than or equal to a specified first threshold value, identifying the request information associated with each of at least one first interval as one or more request information associated with each of at least one object.
[0302] According to one embodiment, the operation of identifying one or more request information associated with at least one of at least one object may include, when the calculated first similarity is less than a specified first threshold value, the operation of identifying one or more first intervals that do not correspond temporally to each of at least one object among at least one first interval, and when the second similarity between each of one or more first intervals and each of at least one object is greater than or equal to a specified second threshold value, the operation of identifying at least one request information associated with each of one or more first intervals as one or more request information associated with each of at least one object.
[0303] According to one embodiment, the operation of identifying one or more request information associated with at least one of at least one object may include, when the calculated first similarity is less than a specified first threshold, the operation of identifying at least one second segment other than at least one first segment in the first user input, and when the second similarity between each of at least one second segment and each of at least one object is greater than or equal to a specified second threshold, the operation of identifying at least one request information associated with each of at least one second segment as one or more request information associated with each of at least one object.
[0304] According to one embodiment, the second user input may include at least one of a touch input, a gesture input, a drag action, or an eye tracking input for specifying each of at least one object in the first data displayed through the display of the electronic device.
[0305] According to one embodiment, a computer-readable non-transient recording medium may store instructions. When the instructions are executed by at least one processor, an electronic device may be configured to perform: an operation of acquiring first data including at least one object; an operation of receiving a first user input including at least one request information for generating at least one content associated with at least one of at least one object included in the first data through an input module of the electronic device; an operation of receiving a second user input for specifying at least one of at least one object in the first data through an input module; an operation of identifying one or more request information associated with at least one of at least one object among at least one request information based on at least one interval of each of the first user input and the second user input received at a time when each of the first user input and the second user input is received or at a specified time; an operation of generating at least one prompt associated with at least one of at least one object based on the identified one or more request information; and an operation of generating content associated with at least one of at least one object based on at least one prompt through a generative artificial intelligence (AI) model.
[0306]
[0307] FIG. 26 is a block diagram of an electronic device (2601) in a network environment (2600) according to various embodiments. Referring to FIG. 26, in the network environment (2600), the electronic device (2601) may communicate with an electronic device (2602) through a first network (2698) (e.g., a short-range wireless communication network) or may communicate with at least one of an electronic device (2604) or a server (2608) through a second network (2699) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (2601) may communicate with the electronic device (2604) through a server (2608). According to one embodiment, the electronic device (2601) may include a processor (2620), memory (2630), input module (2650), sound output module (2655), display module (2660), audio module (2670), sensor module (2676), interface (2677), connection terminal (2678), haptic module (2679), camera module (2680), power management module (2688), battery (2689), communication module (2690), subscriber identification module (2696), or antenna module (2697). In some embodiments, at least one of these components (e.g., connection terminal (2678)) may be omitted from the electronic device (2601), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (2676), camera module (2680), or antenna module (2697)) may be integrated into a single component (e.g., display module (2660)).
[0308] The processor (2620) can, for example, execute software (e.g., program (2640)) to control at least one other component (e.g., hardware or software component) of the electronic device (2601) connected to the processor (2620) and perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (2620) can store commands or data received from other components (e.g., sensor module (2676) or communication module (2690)) in volatile memory (2632), process the commands or data stored in volatile memory (2632), and store the resulting data in non-volatile memory (2634). According to one embodiment, the processor (2620) may include a main processor (2621) (e.g., a central processing unit or an application processor) or an auxiliary processor (2623) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), a secure processing unit (SPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (2601) includes a main processor (2621) and an auxiliary processor (2623), the auxiliary processor (2623) may be configured to use less power than the main processor (2621) or to be specialized for a specified function. The auxiliary processor (2623) may be implemented separately from the main processor (2621) or as part thereof.
[0309] The auxiliary processor (2623) may control at least some of the functions or states associated with at least one component of the electronic device (2601) (e.g., display module (2660), sensor module (2676), or communication module (2690)) on behalf of the main processor (2621) while the main processor (2621) is in an inactive (e.g., sleep) state, or together with the main processor (2621) while the main processor (2621) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (2623) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (2680) or communication module (2690)). According to one embodiment, the auxiliary processor (2623) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (2601) itself where the artificial intelligence is performed, or through a separate server (e.g., server (2608)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network (DQN), or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0310] The memory (2630) can store various data used by at least one component of the electronic device (2601) (e.g., processor (2620) or sensor module (2676)). The data may include, for example, software (e.g., program (2640)) and input or output data for related commands. The memory (2630) may include volatile memory (2632) or non-volatile memory (2634).
[0311] The program (2640) may be stored as software in memory (2630) and may include, for example, an operating system (2642), middleware (2644), or an application (2646). One or more related applications may form a service configured to handle a series of user requests.
[0312] The input module (2650) can receive commands or data to be used for a component of the electronic device (2601) (e.g., processor (2620)) from outside the electronic device (2601) (e.g., user). The input module (2650) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0313] The sound output module (2655) can output a sound signal to the outside of the electronic device (2601). The sound output module (2655) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0314] The display module (2660) can visually provide information to an external (e.g., user) of the electronic device (2601). The display module (2660) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (2660) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0315] The audio module (2670) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (2670) can acquire sound through the input module (2650) or output sound through the sound output module (2655) or an external electronic device (e.g., electronic device (2602)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (2601).
[0316] The sensor module (2676) can detect the operating state of the electronic device (2601) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (2676) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0317] The interface (2677) may support one or more specified protocols that can be used for the electronic device (2601) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (2602)). According to one embodiment, the interface (2677) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0318] The connection terminal (2678) may include a connector through which the electronic device (2601) can be physically connected to an external electronic device (e.g., electronic device (2602)). According to one embodiment, the connection terminal (2678) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0319] The haptic module (2679) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (2679) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0320] The camera module (2680) can capture still images and video. According to one embodiment, the camera module (2680) may include one or more lenses, image sensors, image signal processors, or flashes.
[0321] The power management module (2688) can manage power supplied to the electronic device (2601). According to one embodiment, the power management module (2688) may be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0322] The battery (2689) can supply power to at least one component of the electronic device (2601). According to one embodiment, the battery (2689) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0323] The communication module (2690) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (2601) and an external electronic device (e.g., electronic device (2602), electronic device (2604), or server (2608)), and the performance of communication through the established communication channel. The communication module (2690) may include one or more communication processors that operate independently of the processor (2620) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (2690) may include a wireless communication module (2692) (e.g., cellular communication module, short-range wireless communication module, or global navigation satellite system (GNSS) communication module) or a wired communication module (2694) (e.g., local area network (LAN) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (2604) via a first network (2698) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (2699) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (2692) can identify or authenticate the electronic device (2601) within a communication network such as the first network (2698) or the second network (2699) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (2696).
[0324] The wireless communication module (2692) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. The NR access technology can support enhanced mobile broadband (Embb), massive machine type communications (mMTC), or ultra-reliable and low-latency communications (URLLC). The wireless communication module (2692) can support high-frequency bands (e.g., mmWave bands) to achieve high data transmission rates, for example. The wireless communication module (2692) can support various technologies for securing performance in high-frequency bands, for example, beamforming, multiple-input and multiple-output (massive MIMO), full-dimensional MIMO (FD-MIMO), array antennas, analog beamforming, or large-scale antennas. The wireless communication module (2692) can support various requirements specified in the electronic device (2601), an external electronic device (e.g., electronic device (2604)), or a network system (e.g., a second network (2699)). According to one embodiment, the wireless communication module (2692) can support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.
[0325] An antenna module (2697) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (2697) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (2697) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (2698) or a second network (2699), may be selected from the plurality of antennas, for example, by a communication module (2690). A signal or power may be transmitted or received between the communication module (2690) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (2697).
[0326] According to one embodiment, the antenna module (2697) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0327] At least some of the above components can be connected to each other and exchange signals (e.g., commands or data) through a communication method between peripheral devices (e.g., bus, general purpose input and output (GPIO), serial peripheral interface (SPI), or mobile industry processor interface (MIPI).
[0328] According to one embodiment, commands or data may be transmitted or received between the electronic device (2601) and an external electronic device (2604) through a server (2608) connected to a second network (2699). Each of the external electronic devices (2602, or 2604) may be the same or a different type of device as the electronic device (2601). According to one embodiment, all or part of the operations performed on the electronic device (2601) may be performed on one or more of the external electronic devices (2602, 2604, or 2608). For example, if the electronic device (2601) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (2601) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least a part of the requested function or service, or additional functions or services related to the request, and transmit the result of the execution to the electronic device (2601).
[0329] The electronic device (2601) may process the above results, either as is or additionally, and provide them as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (2601) may provide ultra-low latency services, for example, by using distributed computing or mobile edge computing. In another embodiment, the external electronic device (2604) may include an Internet of Things (IoT) device. The server (2608) may be an intelligent server using machine learning or a neural network. According to one embodiment, the external electronic device (2604) or the server (2608) may be included within a second network (2699). The electronic device (2601) may be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0330] FIG. 27 illustrates a generative artificial intelligence system (2700) according to one embodiment.
[0331] Referring to FIG. 27, the artificial intelligence (AI) system (2700) may include a user interface (2710), an AI framework (2720), a generative AI model (2730), a database (2740), and an application service module (2750). These components may be operated on one or more of an electronic device (e.g., 2601 in FIG. 26), an external electronic device (e.g., 2602, or 2604 in FIG. 26), or a server (e.g., 2608 in FIG. 26). For example, the user interface (2710) and the artificial intelligence framework (2720) can be operated on an electronic device (e.g., 2601 in FIG. 26), and the generative artificial intelligence model (2730) and the database (2740) can be operated on a server (e.g., 2608 in FIG. 26).
[0332] According to one embodiment, a user interface (2710) may receive user input (e.g., user query). User input may be received in the form of text, images, voice (e.g., natural language), video, menu selection, or a combination thereof. The user interface (2710) may include various context information (e.g., running application or user location) related to the generative artificial intelligence system (2700) at the time the user input is received, in addition to or instead of the user input. The user interface (2710) may provide the user input or the context information to the artificial intelligence framework (2720) and provide the processing results therefrom to the user, for example, through the artificial intelligence framework (2720). According to one embodiment, in addition to user input, the electronic device may provide context information obtained using information included on the screen to the artificial intelligence framework (2720). The results may be provided in the form of text, images, voice, video, actions requested by the user (e.g., execution of a specified function or app), or a combination thereof.
[0333] According to one embodiment, the artificial intelligence framework (2720) can identify (e.g., estimate) a user intent based on at least part of user input or context information received from a user interface (2710), control each of the relevant modules (e.g., 2721, 2723, or 2725) to perform a function or action corresponding to the identified user intent, and coordinate collaboration between two or more modules. The artificial intelligence framework (2720) may include a prompt design module (2721), an API / plugin management module (2723), and an output modification module (2725), as illustrated in FIG. 27.
[0334] According to one embodiment, the prompt design module (2721) can generate a prompt to be input to a generative artificial intelligence model (2730) based on at least some of the user input or context information received from the user interface (2710). For example, the prompt design module (2721) can generate a prompt using user preferences, a prompt library, or prompt examples stored in a database (2740) (e.g., a knowledge repository) based on at least some of the user input or context information.
[0335] According to one embodiment, the API / plugin management module (2723) may communicate, for example, via an API, with various resources (e.g., a database (2740)) that provide said additional information when there is a request for said additional information in relation to user input. Additionally or alternatively, when a specified action (e.g., a function, an app, or a service) is performed in response to said user input, the API / plugin management module (2723) may request the application / service module (2750) to perform said specified action via a corresponding API. The API / plugin management module (2723) may provide information obtained from the database (2740), the application / service module (2750), or another external resource to the prompt design module (2721). The obtained information may be used by the prompt design module (2721) in conjunction with the user input to generate a prompt or provided to a generative artificial intelligence model (2730).
[0336] According to one embodiment, the output modification module (2725) can fine-tune the results obtained through the generative artificial intelligence model (2730) as at least part of the response to user input (e.g., user query). For example, the output modification module (2725) can determine whether the content of the response obtained through the generative artificial intelligence model (2730) is appropriate as a response to a request made by the user input. For example, the output modification module (2725) can determine the degree of correlation, the degree of bias (e.g., political or social bias), or the degree of harmfulness (e.g., sexual or profanity) of the difference between the response obtained through the generative artificial intelligence model (2730) and the user input. Additionally or generally, the output modification module (2725) can request that additional AI processing be performed on the obtained response, or provide the user with a hint to avoid unwanted output. For example, additional prompts can be generated through the prompt design module (2721) to obtain a response again through the generative artificial intelligence model (2730).
[0337] According to one embodiment, a generative artificial intelligence model (2730) may form at least part of an artificial intelligence neural network and may include a model that generates images or a model that generates language. An image generation model may include, for example, a generative adversarial network (GAN), a variational autoencoder (VAE), or a diffusion-based model using a VAE and a transformer. A language generation model may include, for example, a large language model (LLM), a large multimodal model (LMM), a large vision model (LVM), or a large action model (LAM). An LAM may automatically generate actions for an environment (e.g., a robot, a car, an electronic device (e.g., 2601 in FIG. 26) or a program (e.g., 2640 in FIG. 26). In addition, for at least some AI models (e.g., LLM), there may be a LoRA (low-rank adaptation) adapter fine-tuned for, for example, specific tasks or specific situations.
[0338] FIG. 28 illustrates an artificial intelligence framework (2720) having on-device AI processing capabilities according to one embodiment.
[0339] In this case, the artificial intelligence framework (2720) may generate and learn a response to the user input by using resources within the device, instead of sending the user input received through a user interface (e.g., 2710 in FIG. 27) operating on the same device (e.g., electronic device (e.g., 2601 in FIG. 26)) to a generative AI model (e.g., 2730 in FIG. 27) operating on an external device (e.g., server (e.g., 2608 in FIG. 26)), or additionally. Referring to FIG. 28, the artificial intelligence framework (2720) may include a cross-application action module (2810), a personal data managing module (2830), an on-device AI model (2850), and an orchestration module (2870).
[0340] According to one embodiment, the cross-application action module (2810) determines one or more additional applications required for the operation of an executed application (e.g., an assistant app) and may connect or suggest operations between the app and at least one additional application, or between a plurality of additional applications. For example, the cross-application action module (2810) may execute one or more additional applications to be used to respond to a user request through the assistant app sequentially or at least partially and simultaneously. Additionally, the cross-application action module (2810) may communicate with the additional applications so that the result of the execution of one additional application (e.g., content) can be shared with other additional applications.
[0341] According to one embodiment, the personal data management module (2830) may provide personal information (e.g., schedule, contact, or message information) about a user of the application (e.g., assistant app) or the additional application running on the device (e.g., electronic device (e.g., 2601 in FIG. 26)) or other related individuals (e.g., family or friends) to another module of the artificial intelligence framework (2720) or a related module (e.g., generative AI model (e.g., 2730 in FIG. 27)) running on another device.
[0342] According to one embodiment, the on-device AI model (2850) may include at least one model among one or more AI models (e.g., GAN, VAE, LLM, LMM, LVM, or LAM) operated on an external device (e.g., a server (e.g., 2608 in FIG. 26)) or a corresponding lightweight AI model. Additionally, for said model or said lightweight model, there may be, for example, a LoRA adapter.
[0343] According to one embodiment, the orchestration module (2870) may select one or more AI models to be used to obtain a response to user input (e.g., user query). For example, the orchestration module (2870) may select one or more AI models from among an on-device AI model (2850), an AI model operating on an external device (e.g., a server (e.g., 2608 in FIG. 26)) (e.g., a generative AI model (e.g., 2730 in FIG. 27)), or a third AI model (not shown) operating on another external device. When multiple AI models are selected, the orchestration module (2870) may communicate with the selected models or devices so that the operation between the selected AI models and the processing of the results thereof can be coordinated between the relevant models or devices.
[0344] According to one embodiment, two or more modules of a generative AI system (e.g., 2700 of FIG. 27) (e.g., cross-application action module (2810) and orchestration module (2870)) can be implemented as a single module to maintain the same functionality. Various variations are possible.
Claims
1. In an electronic device, Input module; Memory for storing instructions; and It includes at least one processor, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Acquire first data including at least one object, and Through the input module above, a first user input is received that includes at least one request information for generating at least one content associated with at least one of the at least one object included in the first data, and Through the input module above, a second user input for specifying at least one of the at least one object in the first data is received, and Based on at least one of the intervals of the first user input and the second user input received at the time when each of the first user input and the second user input is received or at a designated time, one or more request information associated with at least one of the at least one object among the at least one request information is identified, and Based on one or more of the identified request information, at least one prompt associated with at least one of the at least one object is generated, and An electronic device that generates content associated with at least one of the at least one object based on at least one prompt through a generative artificial intelligence (AI) model.
2. In Paragraph 1, The above at least one object includes a first object and a second object, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Identifying request information associated with a first segment of the first user input received while receiving a second user input for specifying the first object as one or more request information associated with the first object, and Identifying request information associated with a second segment of the first user input received while receiving a second user input for specifying the second object as one or more request information associated with the second object, and An electronic device that generates at least one prompt associated with each of the first object and the second object based on one or more request information associated with each of the first object and the second object.
3. In Paragraph 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: After receiving a second user input for specifying the first object, receiving a second user input for specifying the second object, and Identifying a third section of the received first user input as one or more request information associated with the first object before receiving a second user input for specifying the second object, and An electronic device that identifies a fourth segment of the first user input received during at least one period, such as while receiving a second user input for specifying the second object, during a specified time period after receiving the second user input for specifying the second object, or before receiving the second user input for specifying the third object included in the first data, as one or more request information associated with the second object.
4. In Paragraph 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: At least a portion of the first user input, wherein the similarity with the first segment is greater than or equal to a specified threshold value, is identified as a portion associated with the first segment, and An electronic device that identifies at least a portion of the first user input, the similarity to the second segment being greater than or equal to a specified threshold value, as a portion associated with the second segment.
5. In Paragraph 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Based on the above one or more request information, a content creation method including the number of contents to be created in association with at least one of the above at least one object is determined, and An electronic device that generates at least one prompt associated with at least one of the at least one object based on the above content generation method and the above one or more request information.
6. In Paragraph 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Among the first user inputs, each of the at least one first intervals received while receiving the second user input for specifying each of the at least one object is identified as a interval that corresponds temporally to each of the at least one object, and Calculate the first similarity between each of the at least one first intervals corresponding in time to each of the at least one object and each of the at least one object, and If the calculated first similarity is greater than or equal to a specified first threshold value, the request information associated with each of the at least one first interval is identified as one or more request information associated with each of the at least one object, and An electronic device that generates each of at least one prompt associated with each of the at least one object based on one or more identified request information.
7. In Paragraph 6, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: If the calculated first similarity is less than the specified first threshold value, one or more first intervals among the at least one first interval that do not temporally correspond to each of the at least one object are identified, and If the second similarity between each of the above one or more first segments and each of the above at least one object is greater than or equal to a specified second threshold, at least one request information associated with each of the above one or more first segments is identified as one or more request information associated with each of the above at least one object, and An electronic device that generates at least one prompt associated with each of the at least one object based on one or more identified request information.
8. In Paragraph 6, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: If the calculated first similarity is less than the specified first threshold value, at least one second segment other than the at least one first segment is identified in the first user input, and Calculate the second similarity between each of the at least one second segment and each of the at least one object, and If the calculated second similarity is greater than or equal to a specified second threshold, at least one request information associated with each of the at least one second interval is identified as one or more request information associated with each of the at least one object, and An electronic device that generates at least one prompt associated with each of the at least one object based on one or more identified request information.
9. In Paragraph 1, The above first user input is, An electronic device comprising at least one of voice input, text input, or natural language command, comprising request information including at least one sentence, word, or phrase requesting the creation of content associated with each of the at least one object.
10. In Paragraph 1, The above second user input is, An electronic device comprising at least one of a touch input, a gesture input, a drag operation, or an eye-tracking input for specifying each of the at least one object in the first data displayed through the display of the electronic device.
11. In a method for generating content of an electronic device, An operation to acquire first data including at least one object; The operation of receiving a first user input including at least one request information for generating at least one content associated with at least one of the at least one object included in the first data, through an input module of the electronic device; The operation of receiving a second user input to specify at least one of the at least one object in the first data through the input module above; An operation of identifying one or more request information associated with at least one of the at least one object among the at least one request information, based on at least one of the intervals of the first user input and the second user input received at the time when each of the first user input and the second user input is received or at a designated time; Based on the one or more identified request information above, an operation to generate at least one prompt associated with at least one of the at least one object; and A method comprising the operation of generating content associated with at least one of the at least one object based on at least one prompt through a generative artificial intelligence (AI) model.
12. In Paragraph 11, The above at least one object includes a first object and a second object, and The operation of identifying one or more request information associated with at least one of the above-mentioned at least one object is An operation of identifying request information associated with a first segment of the first user input received while receiving a second user input for specifying the first object as one or more request information associated with the first object; and The method includes an operation of identifying request information associated with a second section of the first user input received while receiving a second user input for designating the second object as one or more request information associated with the second object, and The operation of generating at least one prompt associated with at least one of the above-mentioned at least one object is, A method comprising the operation of generating at least one prompt associated with each of the first object and the second object based on one or more request information associated with each of the first object and the second object.
13. In Paragraph 11, The operation of generating at least one prompt associated with at least one of the above-mentioned at least one object is, An operation to determine a content generation method including the number of contents to be generated in association with at least one of the at least one object based on the above one or more request information; and A method comprising the operation of generating at least one prompt associated with at least one of the at least one object based on the above content generation method and the above one or more request information.
14. In Paragraph 11, The operation of identifying one or more request information associated with at least one of the above-mentioned at least one object is, Among the first user inputs, an operation of identifying each of the at least one first intervals received while receiving the second user input for specifying each of the at least one object as intervals that correspond temporally to each of the at least one object; An operation to calculate a first similarity between each of the at least one first intervals corresponding temporally to each of the at least one object and each of the at least one object; and A method comprising the operation of identifying request information associated with each of the at least one first interval as one or more request information associated with each of the at least one object when the calculated first similarity is greater than or equal to a specified first threshold value.
15. In a computer-readable non-transient recording medium, The above recording medium stores instructions, and When the above instructions are executed by at least one processor, the electronic device: An operation to acquire first data including at least one object; The operation of receiving a first user input including at least one request information for generating at least one content associated with at least one of the at least one object included in the first data, through an input module of the electronic device; The operation of receiving a second user input to specify at least one of the at least one object in the first data through the input module above; An operation of identifying one or more request information associated with at least one of the at least one object among the at least one request information, based on at least one of the intervals of the first user input and the second user input received at the time when each of the first user input and the second user input is received or at a designated time; Based on the one or more identified request information above, an operation to generate at least one prompt associated with at least one of the at least one object; and A recording medium configured to perform an operation of generating content associated with at least one of the at least one object based on at least one prompt through a generative artificial intelligence (AI) model.