Methods, apparatus, computer-readable storage media, and computer programs for image manipulation in effect creation tools.
The effect creation tool simplifies image manipulation by allowing users to input natural language requests to a model, receiving and importing image results directly, addressing the challenges of image import and manipulation in game engine-based tools.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- LEMON CO LTD
- Filing Date
- 2024-05-28
- Publication Date
- 2026-06-24
AI Technical Summary
Users of game engine-based tools face challenges in finding and importing desired images for video creation, requiring additional knowledge or artistic skills, limiting their ability to manipulate images within the tool.
An effect creation tool utilizing a model-based approach that allows users to input natural language strings for image acquisition or modification, receiving image results from a model, and importing them directly into the tool, simplifying the image manipulation process.
Enables users to easily obtain and modify images within the tool using natural language requests, reducing complexity and enhancing user experience without the need for external image retrieval or editing skills.
Smart Images

Figure 2026520697000001_ABST
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims priority to U.S. Application No. 63 / 505,346, filed May 31, 2023 (inventive title: Technology for model - based image manipulation in an effect creation tool), and U.S. Application No. 18 / 448,622, filed August 11, 2023 (inventive title: Technology for model - based image manipulation in an effect creation tool), the disclosures of which are hereby incorporated by reference in their entirety.
[0002] [Technical Field] The described aspects relate to an effect creation tool, more specifically, to performing image manipulation in an effect creation tool.
Background Art
[0003] Along with tools for facilitating the creation of video games or other video - based applications or functions, game engines exist to simplify the creation of video games or other video - based applications or functions by providing much of the video processing or display framework. The tools can include effect creation tools, such as tools based on game engines, and can include a user interface or other mechanisms that enable a user to identify layouts, images, etc. included in the creation of video applications, video effects, and / or corresponding functions. Tools based on game engines can generate corresponding instructions having syntax that can be processed by the game engine based on an interaction with the user interface. In response, the game engine can generate a corresponding video - based application, video effect, or feature based on the syntax generated by the tools.
[0004] Social media applications can also use game engines for specific functions. For example, a video recording and editing application can use a game engine to display recorded videos, display animations on the displayed videos (e.g., display a face mask on the face of the subject in the video), or enable interaction with the displayed videos. On the other hand, when importing images for use in video creation with a game engine-based tool, the user typically needs to find and / or retrieve the desired image and provide its storage location, from which the image is imported. Users are usually limited to using images supported by the game engine-based tool, or images that they can find or create using image editing software other than the game engine-based tool. This may require additional knowledge from the user to find the desired image, or the user's artistic ability to generate or edit the image. [Overview of the Initiative] [Problems that the invention aims to solve]
[0005] The following description provides a simplified overview of one or more realizations in order to offer a basic understanding of such realizations. This overview is not a comprehensive overview of all possible realizations, nor is it intended to identify the key or definitive elements of all realizations, nor to define the scope of any or all realizations. Its sole purpose is to present some simplified concepts of one or more realizations as a prelude to the more detailed descriptions that will follow. [Means for solving the problem]
[0006] In one example, a computer implementation method for image manipulation in an effect creation tool is provided, which includes receiving a natural language string requesting an operation related to acquiring or modifying an image via a user interface provided by the effect creation tool; providing an input to a model that includes at least a portion of the natural language string; receiving an image result output from the model based on the input; and importing the image result as an asset into the effect creation tool.
[0007] In another example, a device for image manipulation in an effect creation tool is provided, comprising one or more processors and one or more non-temporary memories having instructions. When the instructions are executed by the one or more processors, the one or more processors are caused to receive a natural language string requesting an operation related to image acquisition or modification via a user interface provided by the effect creation tool; to provide an input to a model that includes at least a portion of the natural language string; to receive an image result output from the model based on the input; and to import the image result as an asset into the effect creation tool.
[0008] In another example, when executed by one or more processors, one or more non-temporary computer-readable storage media are provided to store instructions causing the one or more processors to execute methods for image manipulation in the effect creation tool. The methods include receiving a natural language string requesting an operation related to acquiring or modifying an image via a user interface provided by the effect creation tool; providing an input to a model that includes at least a portion of the natural language string; receiving an image result output from the model based on the input; and importing the image result as an asset into the effect creation tool.
[0009] To achieve the above and related objectives, one or more implementations include features that are fully described below and, in particular, pointed out in the claims. The following description and accompanying drawings illustrate in detail specific exemplary features of one or more implementations. However, these features represent only a fraction of the various ways in which different implementation principles may be employed, and this specification is intended to include all such implementations and their equivalents. [Brief explanation of the drawing]
[0010] [Figure 1] This is a schematic diagram of an example system for using a model when performing image processing in an effect creation tool, as described in the example.
[0011] [Figure 2] This is a flowchart illustrating an example of how to use a model to perform image manipulation in an effect creation tool, as described in the example.
[0012] [Figure 3] This block diagram shows an example of interaction with the model, as described here.
[0013] [Figure 4] This is a schematic diagram of an example of a device for performing the functions described here. [Modes for carrying out the invention]
[0014] The detailed descriptions provided below in relation to the attached drawings are intended to describe various configurations and not to represent the only configurations that can implement the concepts described herein. These detailed descriptions include specific details to provide a thorough understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts can be implemented without these specific details. In some cases, well-known components are shown in block diagram form to avoid obscuring such concepts.
[0015] This disclosure describes various examples related to assisting in performing image manipulation in effect creation tools using models, such as artificial intelligence (AI) or machine learning (ML) models. Effect creation tools may include applications that enable the creation of video applications (e.g., games), visual or video effects, video features, and augmented reality (AR) or virtual reality (VR) effects (e.g., for social media applications). For example, effect creation tools, provided to work with a game engine, enable the creation of video applications, visual effects, AR / VR effects, etc., for rendering using the game engine. In one example, effect creation tools may include game engine-based tools developed in conjunction with a game engine to provide a mechanism for creating videos or video effects using user interface commands. A game engine can provide a platform for rendering video via a display device, and such game engine may include low-level instructions for rendering video via one or more processors (e.g., a central processing unit (CPU) and / or a graphics processing unit (GPU)), and may expose instructions or interfaces that other applications (e.g., tools based on the game engine) can use to cause the game engine to render specific graphics or videos.
[0016] According to the embodiments described herein, a model may be used to provide image results based on natural language requests entered into an effect creation tool, for example, via a user interface (UI), and the effect creation tool may be used to incorporate such image results into a video application, video effect, or other feature being developed. For example, the effect creation tool may include a UI that includes options for inputting natural language queries to perform specific image-related operations. For example, the UI may include options at specific steps in creating a video application, video effect, or feature that allows obtaining, editing, creating, etc., image variations using natural language requests or queries. To obtain image results, a natural language request may be provided to a model, and the effect creation tool may import such image results as part of the video application, video effect, or feature being created.
[0017] The embodiments described herein enable the effect creation tool to natively support querying models for image manipulation based on natural language requests. In this regard, users creating videos, video effects, or features using the effect creation tool can easily obtain and / or import image results within the tool, import images into the tool, and modify images to be compatible with the tool's import process, without having to download images from outside the tool. This improves the user experience when using the tool, allowing users who may not be familiar with image processing to perform image manipulation using natural language requests without having to understand and use other processes within the tool to create, import, or modify images.
[0018] As used herein, a processor, at least one processor, and / or one or more processors configured or operable to perform multiple actions, individually or in combination, means including at least two different processors capable of performing different, overlapping or non-overlapping subsets of the multiple actions, or a single processor capable of performing all of the multiple actions. In one non-limiting example of multiple processors capable of performing different actions of multiple actions in combination, the description of a processor, at least one processor, and / or one or more processors configured or operable to perform actions X, Y and Z may include at least a first processor configured or operable to perform a first subset of X, Y and Z (e.g., performing X) and at least a second processor configured or operable to perform a second subset of X, Y and Z (e.g., performing Y and Z). Alternatively, the first processor, the second processor, and the third processor may each be configured or operable to perform each of actions X, Y and Z, respectively. It should be understood that any combination of one or more processors may be configured or operable to perform any combination of one or more actions from a plurality of actions.
[0019] As used herein, at least one memory, and / or one or more memories, that is configured or already stores instructions that can be executed by one or more processors to perform multiple actions, either individually or in combination, means that at least two different memories are capable of storing different, overlapping or non-overlapping subsets of instructions for performing different, overlapping or non-overlapping subsets of multiple actions, or a single memory is capable of storing instructions for performing all of multiple actions. In a non-limiting example of one or more memories capable of storing, individually or in combination, different subsets of instructions for performing different actions among multiple actions, the description of at least one memory and / or one or more memories configured, operable, or already storing instructions for performing actions X, Y, and Z may include at least a first memory configured, operable, or already storing a first subset of instructions for performing a first subset of X, Y, and Z (e.g., instructions for performing X), and at least a second memory configured, operable, or already storing a second subset of instructions for performing a second subset of X, Y, and Z (e.g., instructions for performing Y and Z). Alternatively, the first memory, the second memory, and the third memory may each be configured to store, or already store, a first subset of instructions for performing X, a second subset of instructions for performing Y, and a third subset of instructions for performing Z, respectively. It should be understood that any combination of one or more memory locations may store, be configured to store, or be operable of, any one or more instructions that can be executed by one or more processors to perform any one or any combination of multiple actions.Furthermore, one or more processors may be coupled to at least one of one or more memories and configured or operable to execute instructions to perform multiple actions. For example, in the above non-limiting example of different subsets of instructions for performing actions X, Y, and Z, the first processor may be coupled to a first memory storing instructions for performing action X, and at least a second processor may be coupled to at least a second memory storing instructions for performing actions Y and Z, and the first and second processors may work together to execute their respective subsets of instructions to achieve the execution of actions X, Y, and Z. Alternatively, three processors may each access one of three different memories storing one of the instructions for performing X, Y, or Z, and the three processors may work together to execute their respective subsets of instructions to achieve the execution of actions X, Y, and Z. Alternatively, a single processor may execute instructions stored in a single memory or instructions distributed across multiple memories to achieve the execution of actions X, Y, and Z.
[0020] Moving from Figure 1 to Figure 4, examples are given with reference to one or more components and one or more methods that can perform the actions or operations described herein, where dashed components and / or actions / operations may be optional. The operations described below in Figure 2 are presented as being performed in a specific order and / or by exemplary components, but the order of actions and the components that perform the actions may vary in some examples depending on the implementation. Also, in some examples, one or more of the actions, functions and / or components described may be performed by a specially programmed processor, in particular a processor running programmed software or computer-readable media, or by any other combination of hardware components and / or software components capable of performing the described actions or functions.
[0021] Figure 1 is a schematic diagram of an example system for using a model when performing image processing in an effect creation tool according to the embodiments described herein. The system includes a device 100 (e.g., a computing device) having a processor 102 (e.g., one or more processors) and / or memory 104 (e.g., one or more memories). In one example, the device 100 may include a processor 102 and / or memory 104 configured to execute or store instructions or other parameters related to providing an operating system 106 capable of running one or more applications, services, etc. The one or more applications, services, etc. may include an effect creation tool 110 that includes or may include an application that facilitates the creation of video, a video (e.g., a game), video effects, or other video features, and a game engine 120 (e.g., similarly running via the operating system 106) can render the video to a display 108. For example, the processor 102 and the memory 104 may be separate components that are communicably coupled by a bus (for example, on the motherboard or other part of a computing device, on an integrated circuit, such as a system on a chip (SoC)), or components that are integrated with each other (for example, the processor 102 may include the memory 104 as an onboard component 101). In other examples, for example, the processor 102 may include multiple processors 102 of multiple devices 100, and the memory 104 may include multiple memories 104 of multiple devices 100. The memory 104 may store instructions, parameters, data structures, etc., for use / execution by the processor 102 to perform the functions described herein.
[0022] Furthermore, the apparatus 100 can include substantially any apparatus that can have a processor 102 and a memory 104, such as, for example, a computer (e.g., a workstation, a server, a personal computer, etc.), a personal device (e.g., a cellular phone such as a smartphone, a tablet, etc.), a smart device such as a smart TV, and the like. Also, in one example, various components or modules of the apparatus 100 may be within a single device as shown, or may be distributed among different devices communicatively coupled to each other (e.g., within a network).
[0023] The effect creation tool 110 may include a user interface module 112 for generating a user interface for output to the display 108 of the device 100 (or a display of another device). For example, the user interface module 112 may accept user input interactions for creating a video, an application containing a video (e.g., a game), or other video features (e.g., effects for a video), as described herein. For example, interactions may include selecting an image, video, or effect for display, modifying an image, video, or effect, such as adding an effect, modifying a part of an image or video, or overlaying an additional image. Furthermore, the user interface module 112 may output to the display 108 of the device 100, such as outputting a video being created, or outputting a menu with interactive options for creating a video (e.g., a preview of an image or video, a list of images or videos to import). Furthermore, in one example, the effect creation tool 110 may optionally include a model query module 114 for querying a model to obtain image results from a natural language request, an image import module 116 for importing image results into the effect creation tool 110 (e.g., within the current project for creating images, videos, or effects), and / or a model training module 118 for training the model based on a natural language request and the desired image results for the request. In addition, the device 100 can communicate with a model 128 (e.g., an AI model or an ML model), and in the case of a model located remotely, it can communicate via a network 122, and in some examples, the model 128 may instead be stored in memory 104.
[0024] In one example, the effect creation tool 110 can provide a user interface (UI), such as a graphical UI, that facilitates the creation of video applications, video features, etc. for execution using the game engine 120 via the user interface (UI) module 112. The game engine 120 can use one or more processors 102 (e.g., a central processing unit (CPU) and / or a graphics processing unit (GPU)) to provide a platform for rendering video on the display 108 or other display device. For example, the effect creation tool 110 can include a video creation studio application having options for creating features, such as a canvas for video, and inserting textures, overlays, etc. into the video, a video preview window for previewing the created video, and the like. In various examples, the effect creation tool 110 can support operations including obtaining or modifying an image of a video, and support operations in one or more UIs provided via the user interface module 112. The aspects described herein relate to using a model for performing one or more of the image-related operations that can reduce the complexity in using such functionality of the effect creation tool 110.
[0025] In one example, if options are selected or engaged through user interaction for one or more operations provided by the user interface module 112, the model query module 114 can query the model 128 with specific inputs to obtain one or more image results. For example, the user interface module 112 may receive a natural language request for an image or image manipulation from a user interaction. The model query module 114 may provide at least a portion of the natural language request as input to the model 128. In one example, the model query module 114 may further influence the output received from the model 128 (e.g., influence the format or context of the output) by adding specific terminology to the input to the model 128. Such terminology may be specific to a face mask request and / or include additional options provided by the user interface module 112 that are selected through user interaction when making the natural language request. For example, the user interface module 112 may receive a natural language request in the face mask application function of the effect creation tool 110. In this example, the model query module 114 may add the term "face mask," the size of the image to be considered for the face mask (e.g., resolution size or file size), and the file type of the image to be considered for the face mask to the input to the model 128.
[0026] In one example, the model query module 114 can receive one or more image results from the model 128 based on the input provided to the model 128. The image import module 116 can, in one example, import the received image results into the effect creation tool 110, which may include storing the image results in memory 104, loading the image results from memory 104 for display on the UI provided by the user interface module 112, and converting the images from their original file type, resolution, or format to a file type, resolution, or format supported by the effect creation tool 110. In one example, the user interface module 112 can provide the image results as selectable options to be incorporated into a video application or feature being created using the effect creation tool 110.
[0027] In one example, the model training module 118 may provide training data to model 128 to adjust the results received from subsequent queries to model 128. For example, the training data may include instructions for image result parameters or formats received from model 128 for a particular type of input query provided to model 128. In another example, the training data may be based on feedback received from the user (e.g., via the UI provided by the user interface module 112) indicating whether the results provided by model 128 in response to a natural language request accurately represent the intent of the user's request.
[0028] Figure 2 is a flowchart of an example of method 200 for using a model to perform image manipulation in an effect creation tool, according to the embodiments described herein. For example, method 200 may be performed by an apparatus 100 that performs an effect creation tool 110 and / or one or more of its components in order to provide intuitive image manipulation based on natural language requests by using a model.
[0029] In method 200, operation 202 can receive natural language strings requesting operations related to image acquisition or modification via a UI provided by the effect creation tool. For example, user interface module 112 can, in cooperation with, for example, processor 102, memory 104, operating system 106, effect creation tool 110, etc., receive natural language strings requesting operations related to image acquisition or modification via a UI provided by the effect creation tool. For example, user interface module 112 can display various UIs via display 108 that enable user interaction with the UI, and such user interaction is for selecting options to create videos, games, video effects, etc., using the effect creation tool. For example, the UI may include options or operations for acquiring an image as a texture for a video, acquiring variations of an image for a video, editing an image for a video, acquiring a face mask to overlay on a face in a video, etc. As part of one or more of these operations, user interface module 112 can provide a mechanism for the user to input natural language requests related to the operation, such as a text input box.
[0030] For example, on a UI for acquiring an image, the user interface module 112 may include a text box for requesting an image using natural language. For example, on a UI for acquiring variations of an image, the user interface module 112 may allow the selection of an imported image and may also include a text box for requesting variations of the imported image using natural language. For example, on a UI for editing an image, the user interface module 112 may include a text box for requesting how to edit the image (for example, "add" some features to the image). For example, on a UI for acquiring a face mask, the user interface module 112 may include a text box for requesting the type or description of the face mask using natural language.
[0031] Method 200 allows, optionally in operation 204, to generate input that includes at least a portion of a natural language string and context parameters associated with the operation. For example, the model query module 114 can work in conjunction with, for example, the processor 102, memory 104, operating system 106, effect creation tool 110, etc., to generate input that includes at least a portion of a natural language string and context parameters associated with the operation. For example, when retrieving an image, the model query module 114 may include, along with the natural language request, context parameters indicating that an image result is desired, image specifications (which may be shown on the UI), such as size (e.g., resolution or file size), file type, whether it is 2D or 3D, the number of generation steps for the query, and a prompt intensity indicator for the natural language string request. For example, context parameters may be added to the natural language string as separate values.
[0032] For example, to obtain variations of an image, the model query module 114 may include a context parameter containing the original image from which variations are desired, along with the natural language request. In some examples, the context parameter may further include a string indicating that variations of the original image are desired, one or more of the aforementioned context parameters for obtaining the image, etc. In another example, when applying a face mask, the model query module 114 may include a context parameter indicating that the image will be used as a face mask, along with the natural language request (therefore, the result should have some relation to the use of a face mask, such as having already been used as a face mask by another user / application, or having properties that allow it to be overlaid in a visible position, such as eyes).
[0033] In method 200, in operation 206, an input containing at least a portion of a natural language string can be provided to the model. For example, the model query module 114 can work in conjunction with, for example, the processor 102, memory 104, operating system 106, effect creation tool 110, etc., to provide an input containing at least a portion of a natural language string to a model (e.g., model 128). For example, the model query module 114 can provide at least a portion of a natural language string as input to model 128, which may or may not further include the context parameters described above. As described, model 128 may be located remotely or stored in device 100 (e.g., in memory 104). For example, the model query module 114 can query multiple models based on the input.
[0034] In method 200, operation 208 can receive image result output from a model based on the input. In one example, the model query module 114 can receive image results output from a model (e.g., model 128) based on the input, in cooperation with, for example, a processor 102, memory 104, an operating system 106, an effect creation tool 110, etc. For example, the image result output may be based on a natural language string and / or any additional context parameters. In some examples, the context parameters may specify a desired set of output criteria that model 128 can use when providing the image result output. The image result output may include one or more images found by model 128 based on the natural language string and / or context parameters. In one example, the image result output may include the image file, the location of the image file (e.g., a universal resource locator (URL) for the image file), etc. In one example, the model query module 114 can receive image results from multiple models (e.g., if multiple models are queried).
[0035] In method 200, in operation 210, the image result can be imported as an asset into the effect creation tool. For example, the image import module 116 can import the image result as an asset into the effect creation tool 110 in cooperation with, for example, the processor 102, memory 104, operating system 106, and effect creation tool 110. For example, the image import module 116 can import the image result to generate a video effect using the image result, for example, to overlay the image result on a video. For example, the image import module 116 can import the image result as a display image by storing the display image in memory 104, or by displaying an option to select the display image as a selectable asset (for example, as a texture, feature, etc.) within the tool for incorporating it into the video or effect being created. When the display image is imported into the effect creation tool 110, the user interface module 112 can display a UI with one or more selectable options for incorporating the image result into the video application or effect being created. In one example, the image import module 116 may modify the image result as part of the import to the effect creation tool 110, and this modification may include modifying the image size (e.g., resolution or file size), modifying the image file type, and so on.
[0036] In another example, the image import module 116 may further or alternatively import the image result as a prefab of a face mask component that can be applied to the video or effect being created. For example, the image import module 116 may create the prefab in memory as an asset within the effect creation tool 110 that can be used on the video or effect being created (e.g., via user interaction). In one example, the image import module 116 may create an empty prefab in memory and then add a face mask component to the prefab that may contain face-related information, such as materials, components, or other parameters. In this example, the image import component 116 can import the image result into the face mask component, thereby allowing the image result to be applied to the face for the effect. In this regard, the generation process may be automated so that the user can search for the desired effect and select it for import / application without requiring, for example, image retrieval, face mask object creation, image modification, manual image import, or asset storage.
[0037] For example, with respect to an image being acquired or a variation of an image being acquired, the user interface module 112 can display an option to add the image as a texture to the video being created. For example, with respect to a face mask image being acquired, the user interface module 112 can display an option to add the face mask to a face in the video being created (e.g., a detected human face), and / or display the face mask applied to the face in the video preview window.
[0038] In one example, in the case of image editing, method 200 optionally allows prompting via the UI in operation 212 regarding the location on the image for modification. In one example, user interface module 112 can work in conjunction with, for example, processor 102, memory 104, operating system 106, effect creation tool 110, etc., to prompt via the UI regarding the location on the image for modification. For example, user interface module 112 can enable the selection or instruction of the location on the image for modification. Based on the indicated location and natural language strings and / or context parameters, model query module 114 can query model 128 to obtain image results in order to modify the image. For example, image import module 116 can import the image results into effect creation tool 110, and user interface module 112 can modify the image using the image results (for example, placing the image results at the indicated location on the image, or overlaying the image results at the indicated location on the image (or merging the image results with the indicated location on the image)).
[0039] In one example, the image import module 116 may import a single image result or a variety of image results. In one example, the context parameter may indicate the number of desired results from model 128. In one example, the user interface module 112 displays the image results and allows user interaction to select a subset of the image results and import them into the effect creation tool 110. In another example, the model query module 114 can query multiple models 128 for image results. In one example, the user interface module 112 can display image results for each model, thereby allowing user interaction for the UI to select a model, and image results from the selected model may be displayed via the UI. In one example, a variety of image results may be selected from one or more models for import into the effect creation tool 110.
[0040] In method 200, optionally in operation 214, instructions for multiple image results and / or multiple models can be displayed. In one example, the user interface module 112 can display instructions for multiple image results and / or multiple models in cooperation with, for example, the processor 102, memory 104, operating system 106, effect creation tool 110, etc. For example, importing image results can be based on a selection of one or more of multiple image results for import. In another example, if multiple models are queried, the user interface module 112 can display instructions for multiple models, and a selection of one of the models (for example, via user interaction) can cause the user interface module 112 to display the image results associated with the selected model.
[0041] In method 200, training data can be provided to the model as an option in operation 216. In one example, the model training module 118 can provide training data to the model (e.g., model 128) in cooperation with, for example, the processor 102, memory 104, operating system 106, effect creation tool 110, etc. For example, the model training module 118 can provide feedback received from user interaction with the user interface, such as whether the image result is related to a natural language string received via the user interface, the degree or rating of the image result with respect to the natural language string, etc. In another example, the model training module 118 can provide the model with training data including specifications of output parameters desired for a particular image operation, as described herein.
[0042] Figure 3 is a block diagram showing an example of an interaction 300 with model 128 according to the embodiments described herein. For example, the interaction 300 may include an input interaction with model 128 or an output interaction with model 128. For example, user interface module 112 may provide a UI having a user prompt 302 for creating an image. In user prompt 302, the user can input a natural language string for creating an image, for example, "jumping cat". Model query module 114 may provide model 128 with at least a portion of the natural language string to query about the image based on the natural language string. In another example, user interface module 112 may provide a UI showing an image asset 304 already imported into effect creation tool 110 and an option to create an image variation. Model query module 114 may provide model 128 with an instruction for the image asset or the location of the image asset (e.g., URL) to query about the image variation based on the original image asset.
[0043] In yet another example, the user interface module 112 may provide a UI with options for editing or modifying an image, such as a mask image 306, an image asset 308, and / or a user prompt 310. The options for the mask image 306 may include a prompt for a natural language string describing the desired mask. In this example, the model query module 114 may provide the model 128 with at least a portion of the natural language string and context parameters indicating the mask, etc., for querying the image and obtaining image results related to the mask. As described above, the image asset 308 or user prompt 310 may also be possible input interactions for editing or modifying the image. Such input interactions may include, for example, providing the original image to obtain variations for editing, or obtaining an image for editing based on a natural language string.
[0044] In one example, as described herein, model 128 may output one or more images based on a query (e.g., generate images (312)), and the images may be imported into UI panel 314. For example, image results may be imported, and the associated UI may include an option to apply the image results (e.g., as a texture, mask, etc.) to the video being created. In this regard, the effect creation tool 110 can natively support the creation and / or modification of images by using the model to obtain images or modifications thereto. This allows the user to create or modify images using the model without, for example, separately searching for images or modifications, converting images or modifications for use in the effect creation tool 110, or manually importing images or modifications.
[0045] In either case, the effect creation tool 110 enables the creation of video effects using a small number of interactive steps. For example, a natural language request can be entered in user prompt 302 or 310. After the request is entered, the effect creation tool 110 can automate the search of model 128 for the corresponding image result and return the image result for import into UI panel 314. In one example, this may include displaying the image result and enabling user interaction with one or more of the image results to import the image resource into the effect creation tool 110 and apply the image result as a video effect. In another example, the effect creation tool 110 may display a single option to perform a combined step of importing the image result and applying it as a video effect.
[0046] Figure 4 shows an example of a device 400 identical or similar to device 100 (Figure 1), including details of additional optional components as shown in Figure 1. In one implementation, device 400 may include a processor 402 similar to processor 102 for performing processing functions associated with one or more of the components and functions described herein. The processor 402 may comprise a single or multiple sets of processors or a multi-core processor. The processor 402 may also be implemented as an integrated processing system and / or a distributed processing system.
[0047] The device 400 may further include a memory 404, which may be similar to memory 104, for storing, for example, a local version of an application being executed by processor 402, such as an effect creation tool 110, associated modules, instructions, parameters, etc. Memory 404 may include types of memory available to the computer, such as random access memory (RAM), read-only memory (ROM), tape, magnetic disk, optical disk, volatile memory, non-volatile memory, and any combination thereof.
[0048] Furthermore, the device 400 may include a communication module 406 that enables it to establish and maintain communication with one or more other devices, parties, entities, etc., using the hardware, software, and services described herein. The communication module 406 may communicate between modules on the device 400, and between the device 400 and external devices, such as devices located via a communication network, and / or devices connected in series or locally to the device 400. For example, the communication module 406 may include one or more buses and further include a transmit chain module and a receive chain module associated with wireless or wired transmitters and receivers, respectively, which are operable for interfacing with external devices.
[0049] Furthermore, the device 400 may include a data store 408, which may be any suitable combination of hardware and / or software enabling large-capacity storage of information, databases, and programs employed in connection with the implementation described herein. For example, the data store 408 may be or include a data repository for applications and / or associated parameters (e.g., the effect creation tool 110, associated modules, instructions, parameters, etc.) that are being executed by the processor 402 or are not currently being executed. Furthermore, the data store 408 may be a data repository for the effect creation tool 110, associated modules, instructions, parameters, etc., and / or for one or more other modules of the device 400.
[0050] The device 400 may include a user interface module 410 that is operable to receive input from a user of the device 400 and is also operable to generate output for presentation to the user. The user interface module 410 may include one or more input devices, including but not limited to a keyboard, numeric keypad, mouse, touch sensor display, navigation keys, function keys, microphone, voice recognition component, gesture recognition component, depth sensor, eye-tracking sensor, switch / button, any other mechanism capable of receiving input from a user, or any combination thereof. Furthermore, the user interface module 410 may include one or more output devices, including but not limited to a display, speaker, haptic feedback mechanism, printer, any other mechanism capable of presenting output to a user, or any combination thereof. The user interface module 410 may include or communicate with the user interface module 112 to enable input via the user interface module 112 or to receive output via the user interface module 112 for display.
[0051] For example, an element, any part of an element, or any combination of elements may be implemented using a “processing system” comprising one or more processors. Examples of processors include microprocessors, microcontrollers, digital signal processors (DSPs), field programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various functions described throughout this disclosure. One or more processors in the processing system may run software. Software should be interpreted broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, etc.
[0052] Therefore, in one or more implementations, one or more of the described functions may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, these functions may be stored or encoded on a computer-readable medium as one or more instructions or codes. The computer-readable medium includes computer storage media. The storage medium may be any available medium accessible by a computer. Such computer-readable media may include, but not limited to, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium accessible by a computer that is used to carry or store desired program code in the form of instructions or data structures. Magnetic disks and optical disks as used herein include compact discs (CDs), laserdiscs, optical disks, digital versatile discs (DVDs), and floppy disks, where magnetic disks typically reproduce data magnetically, and optical disks reproduce data optically using a laser. The above combinations should also be included within the scope of computer-readable media.
[0053] The foregoing explanations are provided to enable those skilled in the art to implement the various realizations described herein. Various modifications to these realizations are obvious to those skilled in the art, and the general principles defined herein may apply to other realizations. Therefore, the claims are not intended to be limited to the realizations shown herein, but are given the full scope consistent with the language of the claims, where a singular reference to an element means "one or more" unless specifically stated as "only." Unless specifically stated, the term "several" means one or more. All structural and functional equivalents of the elements of the various realizations described herein, known or hereafter known to those skilled in the art, are intended to be encompassed within the claims. Furthermore, nothing disclosed herein is intended to be transferred to the public, whether expressly stated in the claims or not. No element shall be construed as means + function unless expressly referenced using the phrase "means for."
Claims
1. A computer implementation method for image manipulation in an effect creation tool, The effect creation tool receives natural language strings requesting operations related to image acquisition or modification via a user interface provided by the tool. Providing the model with input that includes at least a portion of the aforementioned natural language string, The image result output based on the aforementioned input is received from the model, Importing the aforementioned image results as assets into the effect creation tool, Computer implementation methods including
2. To generate the input to include the natural language string and the context parameters associated with the operation, The computer implementation method according to claim 1, further comprising:
3. The operation includes acquiring the image, the image result includes the display image, and importing the image result includes creating the asset for the display image within the effect creation tool. The computer implementation method according to claim 1.
4. The operation includes obtaining the image as a variation of the original image, the image result includes the display image, and importing the image result includes creating the asset for the display image within the effect creation tool. The computer implementation method according to claim 1.
5. The further includes prompting for the location of the modification on the image via the user interface, The operation includes the modification to the image, and importing the image result includes applying the image result as the modification to the image. The computer implementation method according to claim 1.
6. The operation includes applying the image as a mask to an object, the natural language string indicates the image to be applied as a mask, and importing the image result includes applying the image result as a mask onto the image. The computer implementation method according to claim 1.
7. Providing the input to the model includes providing the input to multiple models, and receiving the image results includes receiving multiple image results from the multiple models. The computer implementation method according to claim 1.
8. Display instructions for the multiple models, and based on the selection of one of the multiple models, display a portion of the multiple image results corresponding to the one of the multiple models. The computer implementation method according to claim 7, further comprising:
9. An image manipulation device in an effect creation tool, comprising one or more processors and one or more non-temporary memories having instructions, wherein when an instruction is executed by the one or more processors, the one or more processors, The effect creation tool receives natural language strings requesting operations related to image acquisition or modification via a user interface provided by the tool. Providing the model with input that includes at least a portion of the aforementioned natural language string, The image result output based on the aforementioned input is received from the model, A device that imports the aforementioned image results as assets into the effect creation tool and performs the following actions.
10. The apparatus according to claim 9, wherein, when the instruction is executed by the one or more processors, it causes the one or more processors to generate the input, which includes the natural language string and the context parameters associated with the operation.
11. The apparatus according to claim 9, wherein the operation includes acquiring the image, the image result includes a display image, and the instruction, when executed by the one or more processors, causes the one or more processors to import the image result, which includes creating the asset for the display image in the effect creation tool.
12. The apparatus according to claim 9, wherein the operation includes acquiring the image as a variation of the original image, the image result includes a display image, and the instruction, when executed by the one or more processors, causes the one or more processors to import the image result, which includes creating the asset for the display image in the effect creation tool.
13. The apparatus according to claim 9, wherein, when the instruction is executed by the one or more processors, the one or more processors prompt the user interface to specify the location of a modification to the image on the image, the operation of which includes the modification to the image, and when the instruction is executed by the one or more processors, the one or more processors import the image result, which includes applying the image result as the modification to the image.
14. The apparatus according to claim 9, wherein the operation includes applying the image as a mask to an object, the natural language string indicates the image to be applied as a mask, and the instruction, when executed by the one or more processors, causes the one or more processors to import the image result, which includes applying the image result to the image as a mask.
15. When the instruction is executed by the one or more processors, it causes the one or more processors to perform the task of providing the input to the multiple models. The apparatus according to claim 9, wherein when the instruction is executed by the one or more processors, it causes the one or more processors to perform the task of receiving a plurality of image results from the plurality of models.
16. The apparatus according to claim 15, wherein when the instruction is executed by the one or more processors, the one or more processors cause the one or more processors to display instructions for the plurality of models and, based on the selection for one of the plurality of models, to display a portion of the plurality of image results corresponding to one of the plurality of models.
17. When executed by one or more processors, one or more non-temporary computer-readable storage media storing instructions for causing the one or more processors to execute a method for image manipulation in an effect creation tool, wherein the method is The effect creation tool receives natural language strings requesting operations related to image acquisition or modification via a user interface provided by the tool. Providing the model with input that includes at least a portion of the aforementioned natural language string, The image result output based on the aforementioned input is received from the model, One or more non-temporary computer-readable storage media, including importing the aforementioned image results as assets into the effect creation tool.
18. The method further comprises generating the input to include the natural language string and context parameters associated with the operation, one or more non-temporary computer-readable storage media according to claim 17.
19. The operation includes acquiring the image, the image result includes a display image, and importing the image result includes creating the asset for the display image within the effect creation tool, one or more non-temporary computer-readable storage media according to claim 17.
20. The operation comprises applying the image as a mask to an object, the natural language string indicates the image to be applied as a mask, and importing the image result comprises applying the image result as a mask to the image, one or more non-temporary computer-readable storage media according to claim 17.