Automatic special effect generation
By using artificial intelligence to generate special effects, and utilizing machine learning models to automatically create special effects, the problem of complex operation of existing tools has been solved, achieving low-threshold and efficient special effects generation, and meeting users' needs for rapid creation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-12
- Publication Date
- 2026-03-13
AI Technical Summary
Existing special effects generation tools have complex user interfaces, resulting in a high barrier to entry and long processing times, making it difficult to meet the creative expression needs of beginners and intermediate users.
Employing artificial intelligence technology, the system receives user text input through a prompting word system, uses machine learning models to generate executable code and scripts, and automatically creates special effects, including AR filters and beauty filters, reducing the complexity of user operations during the effect creation process.
It reduces the difficulty and time cost of generating special effects, allowing users to quickly create high-quality special effects without needing extensive domain knowledge, thus improving content creation efficiency.
Smart Images

Figure CN121666599A_ABST
Abstract
Description
Cross-reference to related applications
[0001] This disclosure claims priority to U.S. Utility Patent Application No. 18 / 233,262, filed August 11, 2023, with the U.S. Patent and Trademark Office, which is incorporated herein by reference in its entirety. Background Technology
[0002] Communication is increasingly conducted using internet-based tools. These tools can be any software or platform. Users can create content and design features using these tools. There is a pressing need to improve the technologies used for content creation and feature design through these tools. Attached Figure Description
[0003] A better understanding will be achieved by reading the following detailed description in conjunction with the accompanying drawings. For ease of illustration, exemplary embodiments of various aspects of this disclosure are shown in the drawings; however, the invention is not limited to the specific methods and means disclosed.
[0004] Figure 1 An example system for automatic special effects generation according to this disclosure is shown.
[0005] Figure 2 An example user interface for automatic effects generation according to this disclosure is shown.
[0006] Figure 3 An example user interface for automatic effects generation according to this disclosure is shown.
[0007] Figure 4 An example user interface for automatic effects generation according to this disclosure is shown.
[0008] Figure 5 An example flowchart generated according to the depiction script of this disclosure is shown.
[0009] Figure 6 An example process for automatic special effects generation is shown.
[0010] Figure 7 Another example process for automatic effects generation is shown.
[0011] Figure 8 Another example process for automatic effects generation is shown.
[0012] Figure 9 Another example process for automatic effects generation is shown.
[0013] Figure 10 Another example process for automatic effects generation is shown.
[0014] Figure 11 Another example process for automatic effects generation is shown.
[0015] Figure 12 An example computing device is shown that can be used to perform any of the techniques disclosed herein. Detailed Implementation
[0016] Communication can be conducted using internet-based tools that allow users to create content (e.g., images and / or video content) and distribute that content to other users for consumption. These internet-based tools can offer users various effects (e.g., filters) to use when creating content. Effects are features that can be used to enhance or modify content. In this way, effects can provide content creators with more capabilities to express their creativity. Effects can include, for example, beauty filters and / or augmented reality (AR) filters. Beauty filters can be configured to enhance facial features, smooth skin, add makeup, change facial characteristics, and / or similar functions. AR filters can be configured to overlay digital elements onto the real-world environment depicted in the content. AR filters can be used to add objects, masks, backgrounds, interactive elements, and / or similar elements to content.
[0017] Special effects can be created using special effects creation tools such as EffectHouse and Lens Studio. When creating effects using these tools, users (e.g., effect creators) may need to perform complex user interface (UI) operations. However, effect creators may require significant domain knowledge or experience before they can perform such complex UI operations. This complexity of UI operations can create a high barrier to entry for effect creation. Furthermore, effect creation can be excessively time-consuming. For most beginner to intermediate effect creators, building effects to realize their ideas can take too much time. Therefore, there is a need to improve the technologies used for effect generation.
[0018] This paper describes an improved technique for special effects generation. The technique utilizes artificial intelligence (AI) to automatically generate special effects upon receiving text input from a user (e.g., a special effects creator). Thus, the technique can be used by special effects creators with little to no domain knowledge or experience, and it is more efficient (e.g., less time-consuming) than existing special effects generation techniques.
[0019] Figure 1An example system 100 for automatically generating special effects is shown, featuring effects that can be used to enhance or modify content (e.g., image content, video content, etc.). System 100 may include a prompting word system 104, an effects creation tool 106, and an effects iteration 108. The prompting word system 104, the effects creation tool 106, and the effects iteration 108 may communicate with each other through one or more networks.
[0020] The prompt word system 104 can receive user input 102. User input 102 can include any text input. User input 102 can include voice (e.g., audio) input. If user input 102 includes voice input, the voice input can be converted into text input using any suitable speech-to-text conversion technique. User input 102 can be received through a user interface. User input 102 received through the user interface can be sent (e.g., forwarded) to the prompt word system 104.
[0021] For example, a user can enter one or more letters, symbols, words, phrases, or sentences in one or more text boxes on the user interface. For instance, a user can enter the word "birthday" in a text box. In response to filling the text box, the user can click one or more "Run" buttons. The letters, symbols, words, phrases, or sentences entered by the user can indicate the effects the user wants to generate. For example, if the user enters "birthday" in a text box, this can indicate that the user wants to generate birthday-themed effects. Based on (e.g., in response to) the user selecting (multiple) run buttons on the user interface, the letters, symbols, words, phrases, or sentences entered by the user can be sent (e.g., forwarded) to the prompt word system 104.
[0022] The prompt word system 104 can receive user input (e.g., letters, symbols, words, phrases, or sentences). The prompt word system 104 can generate executable code and / or scripts based on (e.g., using) user input 102. The prompt word system 104 can include one or more machine learning models, such as large language models. Each machine learning model can include an artificial intelligence (AI) algorithm. The AI algorithm can utilize deep learning techniques and large-scale datasets to process and understand language (e.g., text). Each machine learning model can be trained using massive amounts of data to learn language patterns, enabling the machine learning model to perform tasks. For example, each machine learning model in the prompt word system 104 can be trained to process and understand language (e.g., text, user input 102) to generate executable code and / or scripts. The machine learning model can be configured to receive user input 102 as input. The machine learning model can be configured to generate executable code and / or scripts based on user input 102. The machine learning model can output the generated executable code and / or scripts.
[0023] In an embodiment, to generate executable code and / or scripts, the machine learning model can generate one or more special effects creatives corresponding to user input 102. For example, if user input 102 includes the word "birthday," the machine learning model can generate a "smash cake" special effects creative. This "smash cake" special effects creative may include an AR filter that allows the user to virtually "smash" a cake on their face (or another area of the image or video content item). The "smash cake" special effects creative may also include an AR filter that displays colored confetti and balloon bursting effects over at least one area of the image or video content item. In addition to the "smash cake" special effects creative, the machine learning model can also generate other (e.g., additional or alternative) special effects creatives corresponding to the user input "birthday." If the machine learning model generates more than one special effects creative corresponding to user input 102, the user can select their favorite special effects creative. For example, a list of special effects creatives corresponding to the user input can be output (e.g., displayed) on the user interface. The user can select the desired special effects creative from the list of special effects creatives corresponding to user input 102.
[0024] In an embodiment, to generate executable code and / or scripts, the machine learning model can decompose one special effects creative into special effects descriptions (e.g., components). The special effects creative decomposed by the machine learning model may include a special effects creative selected by the user. Alternatively, the special effects creative decomposed by the machine learning model may include a special effects creative selected by the machine learning model. For example, the machine learning model may automatically select a special effects creative from multiple special effects creatives, such as a first special effects creative. The special effects description (e.g., component) may include one or more of a scene description, resource description, or interaction description.
[0025] For example, a "smash the cake" effect creative can be broken down into effect descriptions (e.g., components). The effect description associated with the "smash the cake" effect creative can include a scene description. This scene description can include details indicating that the scene associated with the effect creative includes three-dimensional (3D) objects, such as a cake, confetti, and balloons exploding around it. The scene description can also include details indicating that facial tracking is used to track the user's face when the virtual cake is smashed.
[0026] The effect description associated with the "smashing cake" effect concept can include a resource description. This resource description can include images of the cake, confetti, and balloons. It can also include 3D meshes of the cake and confetti. Furthermore, it can include a sequence of images of exploding balloons. The resource description can include materials used to make the cake look realistic. Finally, it can include visual effects (VFX) shaders used to add particles to the scene.
[0027] The description of the "smash the cake" effect concept may include an interaction description. This interaction description may include details indicating that the user can interact with the effect by virtually smashing the cake against their face. The interaction description may include details indicating that the user can interact with the effect via a touchscreen. The interaction description may include details indicating that the colored confetti and balloon explosions may be timed or triggered by the smashing action.
[0028] In an embodiment, the prompting word system 104 (e.g., a machine learning model configured to process language and perform language-related tasks) can determine components associated with the scene description. The prompting word system 104 can determine a detailed list of components associated with the scene description. This detailed list of components may include one or more 3D objects, a facial tracker, a screen image, etc. For example, the detailed list of components associated with a scene description for a "smash the cake" effect idea may include 3D objects such as a cake, confetti, balloons, etc. These 3D objects may be placed in the 3D space of the "smash the cake" effect. The detailed list of components associated with a scene description for a "smash the cake" effect idea may include a facial tracker configured to track the user's face when the virtual cake is "smashed." The detailed list of components associated with a scene description for a "smash the cake" effect idea may include a screen image. This screen image may be an image depicting a balloon exploding. This screen image may be placed on the screen during the use of the effect.
[0029] In an embodiment, the prompt word system 104 can determine the resources associated with the resource description. The prompt word system 104 can determine a detailed list of resources associated with the resource description. This detailed list of resources may include one or more images, animation keyframes, materials, textures, 3D meshes, diffuse, normal maps and / or specular maps, particle systems, image sequences, shaders, and / or the like. One or more machine learning models (such as AI-generated content models) may be used to generate the resources in this list. The AI-generated content model(s) may include trained machine learning models(s) that assist or replace manual content generation by generating content based on user-input keywords or needs.
[0030] For example, the resource list associated with the resource description for the "cake smashing" effect creative may include 3D meshes of the cake and confetti. The resource list associated with the resource description for the "cake smashing" effect creative may include diffuse, normal maps, and / or specular maps configured to make the cake look more realistic. The resource list associated with the resource description for the "cake smashing" effect creative may include particle systems to make the confetti scattering effect look more realistic. The resource list associated with the resource description for the "cake smashing" effect creative may include one or more textures (e.g., makeup textures, smear textures, drip textures, etc.) for the face tracker component to realistically simulate the effect of the cake on the user's face. The resource list associated with the resource description for the "cake smashing" effect creative may include an image sequence of exploding balloons for the screen image component to add dynamic elements to the scene. The resource list associated with the resource description for the "cake smashing" effect creative may include VFX shaders for the screen image component to add particles, making the balloons look more lively and fun.
[0031] In an embodiment, the prompt word system 104 can determine the interactions associated with the interaction description. The prompt word system 104 can determine a detailed list of interactions associated with the interaction description. This detailed list of interactions may include one or more user interactions associated with the interaction description (e.g., how a user can interact with a component). For example, the detailed list of interactions associated with an interaction description for a "smash the cake" effect idea may include screen-tapping interactions. A user can tap the screen to delay, show, and / or hide object interactions for a 3D cake object.
[0032] In an embodiment, the prompting word system 104 can generate component categories. These component categories may include scene components, post-processing components, facial effect components, and / or similar categories. The prompting word system 104 can generate these component categories by categorizing a detailed list of components into multiple categories. The prompting word system 104 can generate component parameters associated with each component in the component category. These component parameters may include component transformations associated with scene components. These component parameters may include post-processing parameter fillers associated with post-processing components. These component parameters may include facial effect parameter fillers associated with facial effect components. For example, component parameters for a "smash the cake" effect creative may include parameters configured to position the cake in the center of the user's face. The prompting word system 104 can generate executable code and / or scripts based on the component parameters associated with each component in the component category.
[0033] In this embodiment, the prompting system 104 can send the executable code and / or script to the effects creation tool 106. The effects creation tool 106 can receive the executable code and / or script. The effects creation tool 106 can execute the code and / or script in an execution environment. To execute the code and / or script, the effects creation tool 106 can define (e.g., design, create, generate) a set of application programming interfaces (APIs) for generating effects based on the script. The effects creation tool 106 can call the APIs to assemble effects corresponding to the desired effect. The effects creation tool 106 can build an execution environment for executing the script. The effects creation tool 106 can output the generated effects by running the generated script. The generated effects may include filters used when creating content. The generated effects can be used to enhance or modify content. The generated effects may include, for example, beauty filters and / or AR filters.
[0034] In this embodiment, the effect iteration 108 allows the user to iterate over the generated effect. The effect iteration 108 allows the user to input more text to iterate over the generated effect. For example, the user might want to add new elements to the generated effect. Alternatively, the user might want to remove one or more elements from the generated effect. Or, the user might want to replace one or more elements in the generated effect. The effect iteration 108 can utilize one or more AI-generated content models to generate new elements. The effect iteration 108 can add new elements to the generated effect. The effect iteration 108 can remove one or more elements from the generated effect. The effect iteration 108 can replace one or more elements in the generated effect with new elements (e.g., elements generated by an AI-generated content model).
[0035] In this embodiment, the generated special effects can be output. The generated special effects may include a final special effect (e.g., after all iterations (if any) have been performed in effect iteration 108). The output generated special effects can be used to modify content, such as image or video content. For example, the output generated special effects can be used to modify content before it is distributed by a content service or made available to subscribers of that content service. The content may include short videos. The duration of the short video may be less than or equal to a predetermined time limit, such as one minute, five minutes, or other predetermined minutes. By way of example and not limitation, the short video may include at least one (but no more than four) 15-second clips strung together. The short duration of the short video can provide viewers with a fast-paced entertainment experience, allowing users to watch a large amount of video within a short timeframe. This fast-paced entertainment experience may be popular on social media platforms.
[0036] Figure 2 This is example UI 200. UI 200 can be configured to receive user text input. UI 200 can include one or more text boxes. These text boxes can include prompt text boxes. Users can enter user input through UI 200. Users can enter user input in the prompt text boxes. This user input can include any text input. The user input can include one or more letters, symbols, words, phrases, or sentences from the prompt text boxes. Then, the user can click one or more "Run" buttons. The letters, symbols, words, phrases, or sentences entered by the user can indicate the effects the user wants to generate. For example... Figure 2 As shown, the user has entered "archi!" in the prompt text box. The user can then select the "Run" button. In response to the user's selection of the "Run" button, the user input can be sent (e.g., forwarded) to the prompt system (e.g., prompt system 104). The prompt system can then utilize the user input to generate executable code and / or scripts based on (e.g., using) the user input. The executable code and / or scripts can then be sent to an effects creation tool (e.g., effects creation tool 106) to generate effects.
[0037] Figure 3 This is example UI 300. UI 300 can be configured to display generated effects. UI 300 can display effect 302. Effect 302 can be generated using executable code and / or scripts generated by a prompting word system based on user input (e.g., generated by effect creation tool 106). For example, effect 302 can be generated using executable code and / or scripts generated by a prompting word system 104 based on user input “archi!” (e.g., generated by effect creation tool 106). Effect 302 can include augmented reality (AR) filters. Effect 302 can be configured to apply digital elements (e.g., ... Figure 3 The archetype / cartoon image shown is overlaid onto the real-world environment depicted in the image or video content.
[0038] Figure 4This is example UI 400. UI 400 can be used by the user to perform effect iterations. UI 400 can include iteration settings 402. The user can adjust these iteration settings 402 to modify effect 302. For example, the user can adjust iteration settings 402 to adjust one or more of the following associated with effect 302: position, rotation, scaling, texture, stretching mode, blending mode, color, horizontal flip, or vertical flip. The generated effect can be output. The generated effect can include a final effect (e.g., after performing all iterations, if any). The output generated effect can be used to modify content, such as image or video content.
[0039] Figure 5 This is example flowchart 500. This flowchart describes a process for generating executable code and / or scripts. This process can be executed by one or more trained machine learning models to generate executable code and / or scripts based on user input. At 502, user input can be received. This user input can include arbitrary text input (e.g., one or more letters, symbols, words, phrases, or sentences). This user input can also include speech (e.g., audio) input. If the user input includes speech input, it can be converted to text input using any suitable speech-to-text conversion technique. This user input can be received through a user interface.
[0040] At point 504, creative output can be performed. A list of special effects creatives can be output. These special effects creatives can correspond to user input. For example, each special effects creative in the list can correspond to an effect associated with user input. At point 506, a special effects creative can be selected. This special effects creative can be selected from the list of special effects creatives. This special effects creative can be selected by the user. For example, the user can select their favorite special effects creative. Alternatively, the special effects creative can be selected automatically (e.g., in the absence of user input), such as by a prompting system (e.g., prompting system 104).
[0041] At point 508, the effect concept can be broken down. The effect concept can be broken down into effect descriptions (e.g., components). The effect description (e.g., components) can include one or more of the following: scene description 512, resource description 514, or interaction description 510. At point 516, a detailed list of interactions can be determined. This detailed list of interactions can include one or more user interactions associated with the interaction descriptions (e.g., how a user can interact with the components). At point 518, a detailed list of components associated with the scene description can be determined. The detailed list of components associated with the scene description can include one or more 3D objects, face trackers, screen images, etc. At point 520, a detailed list of resources associated with the resource description can be determined. This detailed list of resources can include one or more images, animation keyframes, materials, textures, 3D meshes, diffuse, normal maps and / or specular maps, particle systems, image sequences, shaders, and / or similar elements.
[0042] The detailed component list can be categorized into several types. At 522, some components in the detailed component list can be categorized as scene components. At 524, some components in the detailed component list can be categorized as post-processing components. At 526, some components in the detailed component list can be categorized as facial effects components. Component parameters can be generated for each component in the component category. At 528, component transformations can be generated. These transformations can be associated with scene components. At 530, post-processing parameter fills can be generated. These fills can be associated with post-processing components. At 532, facial effect parameter fills can be generated. These fills can be associated with facial effect components. At 534, executable code and / or scripts can be generated. These can be generated based on the component parameters associated with each component in the component category. At 536, the executable code and / or scripts can be sent to effects creation tools such as Effect House.
[0043] Figure 6 It shows the result of Figure 1 Example process 600 is shown, executed by one or more components. For example, process 600 may be executed at least partially by system 100. Process 600 may be executed to automatically generate special effects. Although Figure 6 The process is described as a sequence of operations, but those skilled in the art will understand that various embodiments may add, delete, reorder, or modify the described operations.
[0044] At point 602, multiple special effects creatives can be generated. These multiple special effects creatives can be generated by at least one large language model. The multiple special effects creatives can be generated in response to received text input from a user. User input can include any text input (e.g., one or more letters, symbols, words, phrases, or sentences). User input can be received through a user interface. Each of the multiple special effects creatives can correspond to the text input. Special effects creatives can be selected from the multiple special effects creatives. For example, special effects creatives can be selected by the user. Alternatively, special effects creatives can be selected automatically (e.g., in the absence of user input).
[0045] At point 604, the special effects creative can be decomposed. The special effects creative can be decomposed into multiple components. The special effects creative can be decomposed in response to selection of the special effects creative (e.g., by user selection). The special effects creative can be one of these multiple special effects creatives. The multiple components can include one or more of the following: 3D objects, facial trackers, screen images, etc. At point 606, executable code can be generated. This executable code can be generated based on the decomposition of the special effects creative.
[0046] At point 608, code can be executed. This code can be executed by a pre-defined special effects creation tool. The code can be executed to generate special effects. The special effects creation tool can execute the code and / or script within an execution environment. To execute the code and / or script, the special effects creation tool can define (e.g., design, create, generate) a set of application programming interfaces (APIs) within itself for generating special effects based on the script. The special effects creation tool can call the APIs to assemble special effects corresponding to the special effects concept. The special effects creation tool can build an execution environment for executing the script. The special effects creation tool can output the generated special effects by running the generated script.
[0047] At point 610, the generated effects can be output. The output generated effects can be used to modify content, such as images or video content. For example, the output generated effects can be used to modify content before it is distributed by a content service or provided to its subscribers. This content can include short videos. The duration of the short video can be less than or equal to a predetermined time limit, such as one minute, five minutes, or other predetermined minutes. As an example and not a limitation, the short video can include at least one (but no more than four) fifteen-second clips strung together. The short duration of the short video can provide viewers with a fast-paced entertainment experience, allowing users to watch a large amount of video in a short period of time. This fast-paced entertainment experience may be widely popular on social media platforms.
[0048] Figure 7 It shows the result of Figure 1Example process 700 is shown, executed by one or more components. For example, process 700 may be executed at least partially by system 100. Process 700 may be executed to automatically generate special effects. Although Figure 7 The process is described as a sequence of operations, but those skilled in the art will understand that various embodiments may add, delete, reorder, or modify the described operations.
[0049] At point 702, the special effects creative can be decomposed. The special effects creative can be decomposed into multiple components. The special effects creative can be decomposed in response to selection of the special effects creative (e.g., by user selection). The special effects creative can be one of these multiple special effects creatives. The multiple components can include one or more of the following: 3D objects, facial trackers, screen images, etc.
[0050] At point 704, a scene for the special effect can be defined. This effect can correspond to a special effect concept. The scene can include multiple components. The scene can include facial tracking. At point 706, multiple components can be determined. A detailed list of components associated with the scene can be determined. This list of components can include one or more 3D objects, facial trackers, screen images, etc. For example, the list of components associated with a scene description for a "smash the cake" special effect concept can include 3D objects such as cake, confetti, balloons, etc. These 3D objects can be placed in the 3D space of the "smash the cake" special effect. The list of components associated with a scene description for a "smash the cake" special effect concept can include a facial tracker configured to track faces when a virtual cake is "smashed." The list of components associated with a scene description for a "smash the cake" special effect concept can include a screen image. This screen image can be an image depicting a balloon exploding. This screen image can be placed on the screen during the use of the special effect.
[0051] At point 708, parameters can be generated. These parameters can be associated with multiple components. These components can be categorized into different component classes, such as scene components, post-processing components, and facial effects components. Component parameters associated with each component in the component class can be generated. These component parameters can include component transformations associated with scene components. These component parameters can include post-processing parameter fills associated with post-processing components. These component parameters can include facial effect parameter fills associated with facial effects components. For example, component parameters for a "cake-smashing" effect could include parameters configured to position the cake in the center of the face. Executable code and / or scripts can be generated based on these component parameters. The effects creation tool can execute this code and / or scripts in an execution environment to generate the effects. The generated effects can be used to modify content, such as image or video content.
[0052] Figure 8 It shows the result of Figure 1Example process 800 is shown, executed by one or more components. For example, process 800 may be executed at least partially by system 100. Process 800 may be executed to automatically generate special effects. Although Figure 8 The process is described as a sequence of operations, but those skilled in the art will understand that various embodiments may add, delete, reorder, or modify the described operations.
[0053] At point 802, the special effects creative can be decomposed. This special effects creative can be decomposed into multiple components. The special effects creative can be decomposed in response to a user selection. This special effects creative can be one of multiple special effects creatives. These multiple components can include one or more of the following: 3D objects, facial trackers, screen images, etc. At point 804, multiple resources can be identified. These multiple resources can be used for the special effects. The special effects can correspond to the special effects creative. A detailed list of resources associated with the special effects can be identified. This list of resources can include one or more images, animation keyframes, materials, textures, 3D meshes, diffuse, normal maps and / or specular maps, particle systems, image sequences, shaders, and / or similar elements. At point 806, multiple resources can be generated. This resource can be generated using one or more AI-generated content models.
[0054] Figure 9 It shows the result of Figure 1 Example process 900 is shown, executed by one or more components. For example, process 900 may be executed at least partially by system 100. Process 900 may be executed to automatically generate special effects. Although Figure 9 The process is described as a sequence of operations, but those skilled in the art will understand that various embodiments may add, delete, reorder, or modify the described operations.
[0055] At point 902, the special effects creative can be decomposed. This special effects creative can be decomposed into multiple components. The special effects creative can be decomposed in response to a user's selection of this special effects creative. This special effects creative can be one of multiple special effects creatives. These multiple components can include one or more of the following: 3D objects, facial trackers, screen images, etc. At point 606, executable code can be generated.
[0056] At 904, at least one interaction can be identified. This at least one interaction can be identified in relation to an effect. The effect can correspond to an effect concept. This at least one interaction can include one or more user interactions (e.g., how a user can interact with multiple components). For example, the one or more interactions can include showing or hiding (multiple) objects in response to receiving user input (e.g., screen tap gesture, finger swipe gesture, voice command, body movement, etc.). For example, a user can perform one or more gestures to delay, show, and / or hide one or more of the multiple components. At 906, the interaction can be configured. The interaction can be configured to be triggered by a predetermined user gesture. For example, the interaction can be configured to be triggered by a screen tap gesture, finger swipe gesture, voice command, body movement, etc.
[0057] Figure 10 It shows the result of Figure 1 The example process 1000 is executed by one or more components shown. For example, process 1000 may be executed at least partially by system 100. Process 1000 may be executed to automatically generate special effects. Although Figure 10 The process is described as a sequence of operations, but those skilled in the art will understand that various embodiments may add, delete, reorder, or modify the described operations.
[0058] At point 1002, a pre-defined special effects creation tool can be configured. This tool can be configured to include a set of application programming interfaces (APIs). These APIs can be created within the special effects creation tool 106 to generate special effects based on a script. The tool can call the APIs to assemble special effects corresponding to the desired effect. At point 1004, an execution environment can be built. The tool can receive executable code and / or scripts associated with the special effects. This execution environment is built to execute the code / script.
[0059] At point 1006, special effects can be generated. These effects can be generated by calling this set of APIs to execute code and / or scripts within an execution environment. A pre-defined effects creation tool can output the generated effects by running code / scripts generated by a prompting word system (e.g., prompting word system 104). The generated effects can be used to modify content, such as images or video content. For example, the generated effects can be used to modify content before it is distributed by a content service or made available to its subscribers. This content can include short videos. The duration of the short video can be less than or equal to a pre-defined time limit, such as one minute, five minutes, or other pre-defined minutes. As an example and not a limitation, the short video can include at least one (but no more than four) fifteen-second clips strung together. The short duration of the video can provide viewers with a fast-paced entertainment experience, allowing users to watch a large amount of video in a short period. This fast-paced entertainment experience may be widely popular on social media platforms.
[0060] Figure 11 It shows the result of Figure 1 Example process 1100 is shown, executed by one or more components. For example, process 1100 may be executed at least partially by system 100. Process 1100 may be executed to automatically generate special effects. Although Figure 11 The process is described as a sequence of operations, but those skilled in the art will understand that various embodiments may add, delete, reorder, or modify the described operations.
[0061] At point 1102, executable code / script can be generated. This executable code / script can be generated based on the decomposition of the effect creative based on a prompting word system (e.g., prompting word system 106). The effect creative can be decomposed into multiple components. The effect creative can be decomposed in response to a user's selection of the effect creative. The effect creative can be one of multiple effect creatives generated by the prompting word system. The multiple components can include one or more of the following: 3D objects, facial trackers, screen images, etc.
[0062] At point 1104, the code / script can be executed. This code / script can be executed by a predefined special effects creation tool. The code / script can be executed to generate special effects. The special effects creation tool can execute the code and / or script within an execution environment. To execute the code and / or script, the special effects creation tool can define (e.g., design, create, generate) a set of application programming interfaces (APIs) for generating special effects based on the script. The special effects creation tool can call the APIs to assemble special effects corresponding to the special effects concept. The special effects creation tool can build an execution environment for executing the code / script. The special effects creation tool can output the generated special effects by running the generated code / script. At point 1106, the generated special effects can be output. The generated special effects may include a scene. The generated special effects may include multiple resources. The generated special effects may include at least one interaction.
[0063] At point 1108, the generated effect can be iterated. The generated effect can be iterated based on user input. Iterating the generated effect can include removing one or more elements from the generated effect. Iterating the generated effect can include adding one or more new elements to the generated effect. Iterating the generated effect can include replacing at least one element in the generated effect with at least one new element. Iterating the generated effect can include: generating a new element using one or more AI-generated content models and adding that new element to the generated effect, removing one or more elements from the generated effect, and / or replacing one or more elements in the generated effect. A final effect can be output. This final effect is the effect after all iterations (if any). The output final effect can be used to modify content, such as image or video content.
[0064] Figure 12 The diagram shows a computing device that can be used in various aspects, such as... Figure 1 The services, networks, modules, and / or devices described herein. About Figure 1 In the example architecture, any component in system 100 can independently... Figure 12 This is implemented using one or more instances of the computing device 1200 shown. Figure 12 The computer architecture shown illustrates conventional server computers, workstations, desktop computers, laptop computers, tablet computers, network devices, PDAs, e-readers, digital cellular phones, or other computing nodes, and can be used to perform any aspect of the computer described herein, such as to implement the methods described herein.
[0065] The computing device 1200 may include a substrate or “motherboard”, which is a printed circuit board, to which multiple components or devices may be connected via a system bus or other electrical communication paths. One or more central processing units (CPUs) 1204 may operate in conjunction with a chipset 1206. The CPUs 1204 may be standard programmable processors that perform arithmetic and logical operations required for the operation of the computing device 1200.
[0066] Multiple CPUs 1204 can perform desired operations by manipulating switching elements (distinguishing and changing these states) to transition from one discrete physical state to another. Switching elements typically include electronic circuitry (such as flip-flops) that maintains one of two binary states, and electronic circuitry that provides an output state based on a logical combination of the states of one or more other switching elements (such as logic gates). These basic switching elements can be combined to create more complex logic circuits, including registers, adders / subtractors, arithmetic logic units, floating-point units, etc.
[0067] The CPU 1204 can be enhanced or replaced by other processing units, such as the GPUs 1205. The GPUs 1205 may include processing units specifically designed for (but not limited to) highly parallel computing, such as graphics processing and other visualization-related processing.
[0068] Chipset 1206 can provide an interface between CPU(s) 1204 and the remaining components and devices on the substrate. Chipset 1206 can provide an interface with random access memory (RAM) 1208, which serves as the main memory in computing device 1200. Chipset 1206 can also provide an interface with computer-readable storage media, such as read-only memory (ROM) 1220 or non-volatile RAM (NVRAM, not shown), for storing basic routines that may facilitate booting computing device 1200 and transferring information between various components and devices. ROM 1220 or NVRAM can also store other software components required for the operation of computing device 1200 according to the aspects described herein.
[0069] Computing device 1200 can operate in a network environment using logical connections to remote computing nodes and computer systems via a local area network (LAN). Chipset 1206 may include the ability to provide network connectivity via a network interface controller (NIC) 1222 (such as a Gigabit Ethernet adapter). NIC 1222 enables computing device 1200 to connect to other computing nodes via network 1216. It should be understood that multiple NICs 1222 may be present in computing device 1200 to connect the computing device to other types of networks and remote computer systems.
[0070] Computing device 1200 can be connected to mass storage device 1228, which provides non-volatile storage for the computer. Mass storage device 1228 can store system programs, application programs, other program modules, and data, as described in detail herein. Mass storage device 1228 can be connected to computing device 1200 via storage controller 1224 connected to chipset 1206. Mass storage device 1228 can consist of one or more physical storage units. Mass storage device 1228 may include management component 1212. Storage controller 1224 can interface with physical storage units via Serial Attached SCSI (SAS) interface, Serial Advanced Technology Attachment (SATA) interface, Fibre Channel (FC) interface, or other types of interfaces for physical connection and data transfer between the computer and physical storage units.
[0071] The computing device 1200 can reflect the storage of information by changing the physical state of the physical storage units, storing the data on the mass storage device 1228. The specific changes in physical state may depend on various factors and different implementations of this specification. Examples of these factors include, but are not limited to, the technology used to implement the physical storage units and whether the mass storage device 1228 is primary or secondary storage, etc.
[0072] For example, computing device 1200 can issue instructions via storage controller 1224 to change the magnetic properties of a specific location within a disk drive unit, the reflection or refraction properties of a specific location within an optical storage unit, or the electrical properties of a specific capacitor, transistor, or other discrete component in a solid-state storage unit, to store information in mass storage device 1228. Other transformations of the physical medium are also possible without departing from the scope and spirit of this specification; the foregoing examples are provided for illustrative purposes only. Computing device 1200 can also read information from mass storage device 1228 by detecting the physical state or characteristics of one or more specific locations within the physical storage unit.
[0073] In addition to the aforementioned high-capacity storage device 1228, the computing device 1200 can also access other computer-readable storage media to store and retrieve information, such as program modules, data structures, or other data. Those skilled in the art will understand that a computer-readable storage medium can be any available medium for storing non-transient data and accessible by the computing device 1200.
[0074] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile, transient computer-readable storage media and non-transient computer-readable storage media, as well as removable and non-removable media implemented in any method or technology. Computer-readable storage media include, but are not limited to, RAM, ROM, erasable programmable ROM (“EPROM”), electrically erasable programmable ROM (“EEPROM”), flash memory or other solid-state memory technologies, optical disc ROM (“CD-ROM”), digital versatile optical disc (“DVD”), high-definition DVD (“HD-DVD”), Blu-ray or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, other magnetic storage devices, or any other medium that can be used to store desired information in a non-transient manner.
[0075] High-capacity storage devices (such as) Figure 12 The mass storage device 1228 shown can store an operating system used to control the operation of the computing device 1200. This operating system may include a version of the Linux operating system. It may also include a version of the Microsoft Windows Server operating system. Depending on other aspects, it may include a version of the UNIX operating system. Various mobile phone operating systems, such as iOS and Android, may also be used. It should be understood that other operating systems may also be used. The mass storage device 1228 can store other systems, applications, and data used by the computing device 1200.
[0076] Mass storage device 1228 or other computer-readable storage medium may also be encoded with computer-executable instructions that, when loaded into computing device 1200, can transform the computing device from a general-purpose computing system into a special-purpose computer capable of implementing the aspects described herein. These computer-executable instructions transform computing device 1200 by specifying how CPU(s) 1204 transition between states, as described above. Computing device 1200 can access the computer-readable storage medium storing the computer-executable instructions, which, when executed by computing device 1200, can perform the methods described herein.
[0077] Computing devices (such as Figure 12 The computing device 1200 shown may also include an input / output controller 1232 for receiving and processing input from multiple input devices, such as a keyboard, mouse, touchpad, touchscreen, electronic stylus, or other types of input devices. Similarly, the input / output controller 1232 may provide output to a display, such as a computer monitor, flat panel display, digital projector, printer, plotter, or other types of output device. It should be understood that the computing device 1200 may not include... Figure 12All components shown may include Figure 12 Other components not explicitly shown, or those that may be used with Figure 12 The architecture shown is completely different.
[0078] As described in this article, a computing device can be a physical computing device, such as... Figure 12 The computing device 1200 in the system. A computing node may also include virtual machine host processes and one or more virtual machine instances. Computer-executable instructions can be indirectly executed by the physical hardware of the computing device by interpreting and / or executing instructions stored and executed in the context of the virtual machine.
[0079] It should be understood that these methods and systems are not limited to any particular method, component, or implementation. It should also be understood that the terminology used herein is for describing specific embodiments only and is not intended to limit the invention.
[0080] As used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly specifies otherwise. Herein, a range may be expressed as beginning “about” a particular value and / or ending “about” another particular value. When such a range is expressed, another embodiment includes beginning and / or ending at that particular value. Similarly, when the antecedent “about” is used to express a value as an approximation, it should be understood that the particular value constitutes another embodiment. It should also be understood that the endpoints of each range are significant in relation to another endpoint, and also significant independently of the other endpoint.
[0081] "Optional" or "optionally" means that the event or situation described below may or may not occur, and the description includes instances where the event or situation occurs and instances where the event or situation does not occur.
[0082] Throughout the description and claims of this specification, the word “comprising” and its variations, such as “including” and “consisting of”, means “including, but not limited to” and is not intended to exclude, for example, other components, integers, or steps. “Exemplary” means “example” and is not intended to convey indications of preferred or ideal embodiments. “Like” is not used in a limiting sense but is for illustrative purposes only.
[0083] Components that can be used to perform the methods and systems are described. When describing combinations, subsets, interactions, groups, etc., of these components, it should be understood that while specific references to various individual and collective combinations and arrangements of these components may not be explicitly described, specific conjectures and descriptions are made herein for all methods and systems. This applies to all aspects of this application, including but not limited to operations within the methods. Therefore, if various additional operations are available, it should be understood that each of these additional operations can be performed using any particular embodiment or combination of embodiments of the method.
[0084] The method and system can be more readily understood by referring to the following detailed description of preferred embodiments and examples included therein, as well as the accompanying drawings and their descriptions.
[0085] Those skilled in the art will understand that these methods and systems can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware. Furthermore, these methods and systems can take the form of a computer program product on a computer-readable storage medium containing computer-readable program instructions (e.g., computer software). More specifically, the methods and systems can take the form of network-implemented computer software. Any suitable computer-readable storage medium can be used, including hard disks, CD-ROMs, optical storage devices, or magnetic storage devices.
[0086] Embodiments of the methods and systems will be described below with reference to block diagrams and flowcharts of methods, systems, apparatuses, and computer program products. It should be understood that each block in the block diagrams and flowcharts, as well as combinations of blocks in the block diagrams and flowcharts, can be implemented, respectively, by computer program instructions. These computer program instructions can be loaded onto a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute on the computer or other programmable data processing apparatus, create components for implementing the functions specified in the flowchart blocks(s).
[0087] These computer program instructions may also be stored in a computer-readable storage medium that can instruct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of writing comprising computer-readable instructions for implementing the functions specified in the flowchart block(s). The computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in the flowchart block(s).
[0088] The various features and processes described above can be used independently of each other or combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of this disclosure. Furthermore, certain method or process blocks may be omitted in some implementations. The methods and processes described herein are not limited to any particular order, and the blocks or states associated with them may be executed in other suitable orders. For example, the described blocks or states may be executed in an order different from the specifically described order, or multiple blocks or states may be combined into a single block or state. Example blocks or states may be executed serially, in parallel, or in some other manner. Blocks or states may be added or removed in the example embodiments. The configuration of the example systems and components described herein may differ from those described. For example, elements may be added, removed, or rearranged compared to the example embodiments described.
[0089] It should also be understood that various items are shown as being stored in memory or on storage devices during use, and these items, or portions thereof, may be transferred between memory and other storage devices for memory management and data integrity purposes. Alternatively, in other embodiments, some or all of the software modules and / or systems may be executed in memory on another device and communicate with the illustrated computing system via inter-computer communication. Furthermore, in some embodiments, some or all of the systems and / or modules may be implemented or provided in other ways, such as at least partially in firmware and / or hardware, including, but not limited to, one or more application-specific integrated circuits (“ASICs”), standard integrated circuits, controllers (e.g., by executing appropriate instructions, and including microcontrollers and / or embedded controllers), field-programmable gate arrays (“FPGAs”), complex programmable logic devices (“CPLDs”), etc. Some or all of the modules, systems, and data structures may also be stored (e.g., as software instructions or structured data) on computer-readable media, such as hard disks, memory, networks, or portable media articles, for retrieval by appropriate devices or via appropriate connections. The system, modules, and data structures can also be transmitted as generated data signals (e.g., as part of a carrier or other analog or digital propagation signal) over various computer-readable transmission media, including wireless and wired / cable media, and can take many forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital data packets or frames). Such computer program products can also take other forms in other embodiments. Therefore, the invention can be practiced using other computer system configurations.
[0090] While methods and systems have been described in conjunction with preferred embodiments and specific examples, their scope is not intended to be limited to the particular embodiments illustrated, as the embodiments herein are intended to be illustrative rather than restrictive in all respects.
[0091] Unless otherwise expressly stated, no method described herein is intended to be construed as requiring its operations to be performed in a particular order. Therefore, if a method claim does not actually describe the order in which its operations should be followed, or if the claims or specification do not expressly state that the operations should be limited to a particular order, then such an order is not intended to be inferred in any way. This applies to any possible basis for non-express interpretation, including: logical problems related to the arrangement of steps or operational flows; plain meanings derived from grammatical organization or punctuation; and the number or type of embodiments described in the specification.
[0092] Those skilled in the art will understand that various modifications and variations can be made without departing from the scope or spirit of this disclosure. Other embodiments will be conceived by those skilled in the art upon consideration of the specification and the practices described herein. The specification and example figures should be considered exemplary only; the true scope and spirit are indicated by the appended claims.
Claims
1. A method for automatically generating special effects, comprising: Multiple special effects ideas are generated by at least one trained machine learning model in response to received text input; In response to the selection of a special effects creative, the special effects creative is decomposed into multiple components, wherein the special effects creative is one of the multiple special effects creatives; Based on the decomposition of the aforementioned special effects concept, executable code is generated; The code is executed by a pre-designed special effects creation tool to generate the special effects; as well as Output the generated special effects.
2. The method according to claim 1, wherein decomposing the special effects concept further includes: Define the scene for the special effect; Identify the plurality of components; as well as Generate parameters associated with the plurality of components.
3. The method according to claim 1, wherein decomposing the special effects idea further includes determining multiple resources for the special effects.
4. The method according to claim 3, further comprising: The plurality of resources are generated using at least one other machine learning model, wherein the at least one other machine learning model is trained to generate content based on user input.
5. The method of claim 1, wherein decomposing the special effects idea further includes determining at least one interaction for the special effects.
6. The method according to claim 5, further comprising: Configure the at least one interaction to be triggered by a predetermined user gesture.
7. The method of claim 1, wherein the generated special effects include a scene, multiple resources, and at least one interaction.
8. The method of claim 1, wherein the code is executed by the predetermined special effects creation tool to generate the special effects, further comprising: Configure the predetermined special effects creation tool to include a set of application programming interfaces (APIs). Build the execution environment; as well as The special effects are generated by calling the set of APIs to execute the code in the execution environment.
9. The method according to claim 1, further comprising: Based on user input, the generated special effect is iterated, wherein iterating the generated special effect includes removing one or more elements from the generated special effect, adding one or more new elements to the generated special effect, or replacing at least one element in the generated special effect with at least one new element.
10. A system comprising: At least one processor; as well as At least one memory includes computer-readable instructions that, when executed by the at least one processor, cause the system to perform operations, said operations including: Multiple special effects ideas are generated by at least one trained machine learning model in response to received text input; In response to the selection of a special effects creative, the special effects creative is decomposed into multiple components, wherein the special effects creative is one of the multiple special effects creatives; Based on the decomposition of the aforementioned special effects concept, executable code is generated; and The code is executed by a pre-designed special effects creation tool to generate special effects; and Output the generated special effects.
11. The system of claim 10, wherein decomposing the special effects concept further includes: Define the scene for the special effect; Identify the plurality of components; as well as Generate parameters associated with the plurality of components.
12. The system of claim 10, wherein decomposing the special effects idea further includes determining multiple resources for the special effects, and wherein the operation further includes: The plurality of resources are generated using at least one other machine learning model, wherein the at least one other machine learning model is trained to generate content based on user input.
13. The system of claim 10, wherein decomposing the special effects idea further includes determining at least one interaction for the special effects, and wherein the operation further includes: Configure the at least one interaction to be triggered by a predetermined user gesture.
14. The system of claim 10, wherein the generated special effects include a scene, multiple resources, and at least one interaction.
15. The system of claim 10, wherein executing the code by the predetermined special effects creation tool to generate the special effects further comprises: Configure the predetermined special effects creation tool to include a set of application programming interfaces (APIs). Build the execution environment; as well as The special effects are generated by calling the set of APIs to execute the code in the execution environment.
16. The system of claim 10, further comprising: Based on user input, the generated special effect is iterated, wherein iterating the generated special effect includes removing one or more elements from the generated special effect, adding one or more new elements to the generated special effect, or replacing at least one element in the generated special effect with at least one new element.
17. A non-transient computer-readable storage medium storing computer-readable instructions that, when executed by a processor, cause the processor to perform operations, the operations including: Multiple special effects ideas are generated by at least one trained machine learning model in response to received text input; In response to the selection of a special effects creative, the special effects creative is decomposed into multiple components, wherein the special effects creative is one of the multiple special effects creatives; Based on the decomposition of the aforementioned special effects concept, executable code is generated; as well as The code is executed by a pre-designed special effects creation tool to generate the special effects; as well as Output the generated special effects.
18. The non-transient computer-readable storage medium of claim 17, wherein decomposing the special effects concept further comprises: Define the scene for the special effect; Identify the plurality of components; as well as Generate parameters associated with the plurality of components.
19. The non-transient computer-readable storage medium of claim 17, wherein decomposing the special effects concept further comprises: Determine at least one interaction for the special effect, and said operation further includes: Configure the at least one interaction to be triggered by a predetermined user gesture.
20. The non-transient computer-readable storage medium of claim 17, wherein executing the code by the predetermined special effects creation tool to generate the special effects further comprises: Configure the predetermined special effects creation tool to include a set of application programming interfaces (APIs). Build the execution environment; as well as The special effects are generated by calling the set of APIs to execute the code in the execution environment.