Method and apparatus for generating media content, device, and storage medium
The editing interface, which includes material input components, text input components, and style selection components, allows users to determine the target generation mode, thus addressing their needs for personalized media content generation and enabling the efficient generation of media content that meets specific styles and themes.
Patent Information
- Application Number
- PCT/CN2024/096400
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2025-12-04
AI Technical Summary
Users are increasingly demanding personalized editing of generated media content, and existing technologies are struggling to efficiently generate media content that meets specific styles and themes.
A method for generating media content is provided, which determines the target generation mode through the editing interface of the material input component, text input component, and style selection component, and generates the target media content based on the generation control information.
It reduces the complexity of media content editing, improves the efficiency and effectiveness of media content generation, and meets users' needs for personalized editing.
Smart Images

Figure CN2024096400_04122025_PF_FP_ABST
Abstract
Description
Media content generation method, apparatus, device, and storage medium TECHNICAL FIELD
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a media content generation method, apparatus, device, and computer-readable storage medium. BACKGROUND
[0002] With the development of social applications, more and more users publish pictures or videos for showing their daily life or browse the content published by other publishers in the application through the social applications. Thus, the personalized editing demand of the users for the media data to be published is also increasing.
[0003] SUMMARY
[0004] In a first aspect of the present disclosure, a media content generation method is provided. The method comprises: presenting an editing interface, the editing interface comprising a material input component, a text input component, and a style selection component; in response to receiving a media generation request, determining a target generation mode based on input states of the material input component, the text input component, and the style input component; and providing target media content generated based on the target generation mode based on generation control information received in the editing interface.
[0005] In a second aspect of the present disclosure, a media content generation apparatus is provided. The apparatus comprises a presentation module configured to present an editing interface, the editing interface comprising a material input component, a text input component, and a style selection component; a determination module configured to, in response to receiving a media generation request, determine a target generation mode based on input states of the material input component, the text input component, and the style input component; and a providing module configured to provide target media content generated based on the target generation mode based on generation control information received in the editing interface.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The device comprises at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon a computer program, which, when executed by a processor, implements the method of the first aspect.
[0008] It is to be understood that the description of the background of the disclosure contained herein does not constitute an admission that any of the embodiments of the disclosure and / or the BRIEF DESCRIPTION OF DRAWINGS
[0009] The above and other features, aspects, and advantages of various embodiments of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings. In the drawings, like reference numerals refer to like elements, wherein:
[0010] FIG. 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0011] FIGS. 2A to 2G show schematic diagrams of example interfaces for interaction, according to some embodiments of the present disclosure;
[0012] FIG. 3 shows a flowchart of a process of media content generation, according to some embodiments of the present disclosure;
[0013] FIG. 4 shows a block diagram of a media content generation apparatus, according to some embodiments of the present disclosure; and
[0014] FIG. 5 shows a block diagram of a device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION
[0015] Embodiments of the present disclosure will be described in more detail with reference to the drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It is to be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and should not be construed as limiting the scope of protection of the present disclosure.
[0016] In the description of embodiments of the present disclosure, the term "comprising" and its conjugations should be understood to encompass the meanings of "consisting of" and "consisting essentially of". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "an embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions can also be included below.
[0017] It can be understood that the image data (including but not limited to the data itself, the obtaining, use, storage or deletion of the data) involved in the technical solution should comply with the requirements of the relevant laws and regulations and relevant provisions.
[0018] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of data involved in the present disclosure, the use range, the use scenario, etc. should be informed to the relevant users and the authorization of the relevant users should be obtained through appropriate means, wherein the relevant users can include any type of right subject, such as an individual, an enterprise, or a group.
[0019] As used herein, the term“model” can learn an association between respective inputs and outputs from training data, such that after training is complete, a corresponding output can be generated for a given input. The generation of a model can be based on machine learning techniques. Deep learning is a type of machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is one example of a model based on deep learning. In this document, a“model” can also be referred to as a“machine learning model,” a“learning model,” a“machine learning network,” or a“learning network,” which terms are used interchangeably herein.
[0020] As described above, users have increasingly high requirements for the presentation effect of generated media. For example, users expect that media content satisfying a specific style and theme can be generated by a media content generation application on the basis of provided original media material or without the basis of original media material, so that the generated media content can exhibit a presentation effect that conforms to the aesthetic standards of the user.
[0021] The technical solutions of the present disclosure provide a media content generation scheme. In the media content generation scheme of the embodiments of the present disclosure, an editing interface including a material input component, a text input component, and a style selection component is presented. If a media generation request is received, a target generation mode is determined based on the input state of the material input component, the text input component, and the style input component. The target media content generated based on the target generation mode is provided based on the generation control information received in the editing interface.
[0022] In this way, the complexity of media content editing can be reduced and the efficiency of media content generation can be enhanced.
[0023] Example Environment
[0024] Referring first to FIG. 1, a schematic diagram of an example environment 100 in which example implementations consistent with the present disclosure can be implemented is shown schematically.
[0025] FIG. 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In this example environment 100, a target application 120 is installed in an electronic device 110. A user 102 can interact with the target application 120 via the electronic device 110 and / or an attached device of the electronic device 110.
[0026] Target application 120 can be an application capable of providing services to user 102 associated with media content generation, including authoring (e.g., shooting and / or editing), publishing, browsing, etc. of media data. In this document, "media data" can be a variety of forms of content, including video, audio, images, image sets, text, etc.
[0027] For example, electronic device 110 can edit the received raw media material through target application 120. In some embodiments, the raw media material can be, for example, media material input by user 102. In some other embodiments, the raw media material can also be media material captured by a capturing device based on instructions of user 102 through electronic device 110. In some embodiments, the capturing device can be configured to be connected with electronic device 110 with each other. In some other embodiments, the capturing device can also be integrated inside electronic device 110.
[0028] In environment 100 of FIG. 1, if target application 120 is in an active state, electronic device 110 can present a page 140 of target application 120 to user 102. Page 140 can be a variety of pages that target application 120 can provide, such as a presentation page of media data, a content authoring page, a content editing page, etc.
[0029] In some embodiments, implementation of at least part of the functionalities of target application 120 can be implemented based on model 131. Model 131 can be deployed in server 130, for example. For instance, electronic device 110 communicates with server 130 to implement the provision of services of target application 120. That is, during the running of target application 120, the capability of one or more models (e.g., model 131) can be invoked.
[0030] In some embodiments, electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a game device, or any combination thereof, including accessories and peripherals of such devices, or any combination thereof. In some embodiments, electronic device 110 can also be capable of supporting any type of interface to user (such as "wearable" circuitry, etc.). Server 130 can be any type of computing system / server capable of providing computational capability, including but not limited to mainframe, edge computing node, computing device in a cloud environment, etc.
[0031] It should be appreciated that the structure and function of the environment 100 are described for illustrative purposes only and without implying any limitation of the scope of the present disclosure.
[0032] Example interaction process for media content generation
[0033] FIGS. 2A-2G illustrate schematic diagrams of example interfaces for interaction, in accordance with some embodiments of the present disclosure. The interaction process of embodiments of the present disclosure is described below in conjunction with FIGS. 2A-2G. The process can be implemented at least in part at the electronic device 110 shown in FIG. 1. For ease of discussion, the process will be described with reference to the environment 100 of FIG. 1.
[0034] As shown in FIG. 2A, the electronic device 110 can present an editing interface 200A. The editing interface 200A is presented with a material input component 210. If a selection is received for the material input component 210, the electronic device 110 can present a material selection interface including a plurality of selectable reference materials. The selectable reference materials can include, for example, pictures, videos, and the like.
[0035] In some embodiments, if the electronic device 110 receives a selection for a selectable reference material, the selected reference material is presented at a region 211 of the editing interface 200B, as shown in FIG. 2B. In some other embodiments, the electronic device 110 can not receive a selected reference material. In this case, the input state of the material input component 210 is no reference material received, i.e., the input state is empty.
[0036] In addition, the editing interfaces 200A and 200B are also presented with a text input component 220. Via the text input component 220, the electronic device 110 can obtain at least one text content for the media content generation.
[0037] In some embodiments, if the electronic device 110 has received a selected reference material, upon receiving a selection for the text input component 220, the editing interface 200C can present the selected reference material at the region 211 and present an entry 221 for typing text, as shown in FIG. 2C. Via the entry 221, the electronic device 110 can obtain at least one text content for the media content generation. The input text content can be presented at the text input component 220. If the electronic device 110 receives an indication of confirmation for the input text content (e.g., via a confirmation control 222), the electronic device 110 determines the input text content as the text content for the media content generation.
[0038] In some embodiments, if the electronic device 110 does not receive the selected reference material, upon receiving a selection for the text input component 220, the editing interface 200D presents an entry 221 for typing text, as shown in FIG. 2D. Via the entry 221, the electronic device 110 can obtain at least one text content for the media content generation. The inputted text content can be presented at the text input component 220. If the electronic device 110 receives an indication of confirmation for the inputted text content (e.g., via the confirmation control 222), the electronic device 110 determines the inputted text content as the text content for the media content generation.
[0039] If the electronic device 110 receives at least one text content for the media content generation via the text input component 220, the input state of the text input component 220 is that the text content has been received, and if the electronic device 110 does not obtain at least one text content for the media content generation, the input state of the text input component 220 is that no text content has been received, i.e., the input state is empty.
[0040] The editing interfaces 200A and 200B also present a plurality of style input components 230. Via the plurality of style input components 230, the electronic device 110 can obtain a target style for the media content generation.
[0041] In some embodiments, the presentation order of the plurality of style input components 230 can be determined based on the selected reference material. For example, the style input components that are more suitable or match the reference material can be presented more preferentially.
[0042] If the electronic device 110 receives a selection for the target style via one of the plurality of style input components 230, the input state of the style input component 230 is that the target style has been received, and if the electronic device 110 does not receive a selection for the target style, the input state of the text input component 220 is that no target style has been obtained, i.e., the input state is empty.
[0043] If a media generation request is received, e.g., the electronic device 110 receives a selection for the media content generation control 240 in the editing interfaces 200A and 200B, the generation mode of the media content is determined based on the input states of the material input component 210, the text input component 220, and the plurality of style input components 230.
[0044] As has been described above, the input state indicates whether the material input component 210, the text input component 220, or the style input component 230 has received the respective input content associated with generating the media content, i.e., whether its input content is empty.
[0045] In some embodiments, for a generation mode of media content, the electronic device 110 can determine a generation model of target media content and / or at least one control parameter for the generation model based on the input states of the material input component 210, the text input component 220, and the plurality of style input components 230.
[0046] In some embodiments, the generation model can be at least partially provided by the server 130, for example, can be provided by the model 131 deployed on the server 130. Based on the input states of the material input component 210, the text input component 220, and the plurality of style input components 230, the electronic device 110 can determine with which generation model to perform a plurality of processing procedures for corresponding input content associated with the generation of media content.
[0047] For example, if the electronic device 110 has acquired reference material via the material input component 210, at least one text content (e.g., a prompt word) for media content generation via the text input component 220, and a target style via the style input component 230, the electronic device 110 can determine that a generation model that supports the generation of media content based on the reference material, the text content, and the target style is needed. In this case, the generation model can be a model that can support the generation of media content based on the reference material (e.g., a picture), the text, and the target style together.
[0048] For another example, if the electronic device 110 has acquired at least one text content (e.g., a prompt word) for media content generation via the text input component 220 and a target style via the style input component 230, the electronic device 110 can determine that a generation model that supports the generation of media content based on the text content and the target style is needed. In this case, the generation model can be a model that can support the generation of media content based on the text and the target style together.
[0049] It should be understood that the electronic device 110 in the process of determining the generation mode, i.e., determining the generation model of target media content and / or at least one control parameter for the generation model, should determine the corresponding generation mode based on the plurality of cases involved by whether the input state of the material input component 210 is empty, whether the input state of the text input component 220 is empty, and whether the input state of the plurality of style input components 230 is empty. All possible combinations of input content that can exist can fall within the scope of the present disclosure, which are not listed one by one here.
[0050] After the generation mode of the media content is determined, the electronic device 110 can provide the target media content generated based on the generation mode of the media content according to the generation control information received from the editing interface (e.g., from at least one of the editing interfaces 200A to 200D), i.e., the above-mentioned corresponding input content associated with the generation of the media content.
[0051] For example, after the electronic device 110 acquires the reference material via the material input component 210, acquires at least one text content for the generation of the media content via the text input component 220, and acquires the target style via the style input component 230, if a selection of the media content generation control 240 is received, the generation of the media content is performed according to the determined generation mode of the media content and the acquired generation control information, e.g., the above-mentioned corresponding input content associated with the generation of the media content. For example, as shown in FIG. 2E, the electronic device 110 is performing the generation process of the media content, and the area 250 presented on the intermediate interface 200E can display the completion progress of the generation process.
[0052] Once the generation of the media content is completed, the generated target media content is presented on the area 260 of the viewing interface 200E of the electronic device 110, as shown in FIG. 2F.
[0053] In some embodiments, during the generation process, if the electronic device 110 has acquired the reference material via the material input component 210 (e.g., the picture presented at the area 211 of the editing interface 200B), has acquired at least one text content for the generation of the media content via the text input component 220 (e.g., the prompt word: wings of the angel), and has acquired the target style via the style input component 230 (e.g., the clay style), the electronic device 110 can generate and present the generated target media content on the area 260 of the viewing interface 200E according to the determined generation mode and these generation control information. In this case, the generation mode determined by the electronic device 110 is determined based on the condition that the input content of the material input component 210, the text input component 220, and the style input component 230 are all not empty.
[0054] More specifically, in some embodiments, during the generation process of the media content, if the electronic device 110 receives a selection of a first style via the style input component 230, the electronic device 110 provides a set of fine-tuning parameters corresponding to the first style to the generation model and acquires the target media content generated by the generation model based on the set of fine-tuning parameters. The set of fine-tuning parameters can correspond to a fine-tuning model. The fine-tuning model can be trained based on a set of sample media corresponding to the first style.
[0055] If the electronic device 110 receives a selection for a second style via the style input component 230, the electronic device 110 can generate prompt information based on the second style. The prompt information can include description information corresponding to the second style. The electronic device 110 can obtain target media content generated by the generation model based on the prompt information.
[0056] In some embodiments, in addition to the generated target media content, at least a portion of the generation control information can also be presented in the viewing interface. As shown in FIG. 2F, at least one text content (e.g., a prompt word) obtained by the electronic device 110 for media content generation and an indication of the target style can be presented at region 270 of the viewing interface 200F.
[0057] After the generation process of the target media content is completed, if the electronic device 110 receives an indication for publishing the media content (e.g., receives a selection for the publishing control 270), the generated target media content is published.
[0058] According to the media generation scheme described in detail in connection with the above embodiments of the present disclosure, the complexity of the interactive process of media content editing can be reduced, thereby enhancing the efficiency of media content generation.
[0059] In some other embodiments, it is also possible that the electronic device 110 presents one or more candidate prompt words according to the text content (e.g., a prompt word) received via the text input component 220. These candidate prompt words can be text content expanded based on the received text content, so as to make the text content for media content generation more abundant.
[0060] The electronic device 110 can receive a recommendation request for the text content received via the text input component 220. In some embodiments, the recommendation request is received if the electronic device 110 receives a trigger for the recommendation control. As shown in the editing interface 200C of FIG. 2C and the editing interface 200D of FIG. 2D, the recommendation request for the received text content can be determined if the electronic device 110 receives a selection for the text optimization control 223.
[0061] After receiving the recommendation request, the electronic device 110 can provide a plurality of candidate prompt words based on the text content (e.g., a prompt word) that has been received via the text input component 220, for example. As shown in the editing interface 200G of FIG. 2G, the plurality of candidate prompt words 292 to 295 presented in the candidate prompt word presentation region 290 can be generated based on the text content (e.g., a prompt word 291) that has been received via the text input component 220. If the electronic device 110 receives a selection for one of the plurality of candidate prompt words 292 to 295, the previously received text content can be updated to the selected candidate prompt word.
[0062] In addition, it is also possible that the plurality of candidate prompt words can also be generated according to at least one reference material acquired via the material input component 210. In some embodiments, the plurality of candidate prompt words can be generated based on an object detection result of the at least one reference material. The object detection result indicates whether the at least one reference material includes an object of a predetermined type. For example, if the reference material acquired via the material input component 210 does not include an object of a predetermined type (e.g., the reference material does not present a subject of a person, an animal, a plant, etc.), the plurality of candidate prompt words can involve a description of the subject to enrich the media content to be generated.
[0063] It should be understood that the plurality of candidate prompt words can be generated locally by the electronic device 110. It is also possible that the plurality of candidate prompt words can also be generated at least partially by the server 130, for example, by the model 131 deployed on the server 130. The optimization recommendation for the text content received via the text input component 220 can further improve the refinement of the prompt words, thereby bringing better media content generation effect.
[0064] Example process
[0065] FIG. 3 shows a flowchart of a media content generation process 300 according to some embodiments of the present disclosure. The process 300 can be implemented at the electronic device 110. It should be understood that the process 300 can also be implemented at the server 130.
[0066] At block 310, the electronic device 110 presents an editing interface including a material input component, a text input component, and a style selection component.
[0067] At block 320, if the electronic device 110 receives a media generation request, at block 330, the electronic device 110 determines a target generation mode based on input states of the material input component, the text input component, and the style input component.
[0068] At block 340, the electronic device 110 provides target media content generated based on the target generation mode based on the generation control information received in the editing interface.
[0069] In some embodiments, the input state indicates whether the input content of the material input component, the text input component, or the style input component is empty.
[0070] In some embodiments, the target generation mode indicates at least one of a generation model used for the target media content and / or at least one control parameter for the generation model.
[0071] In some embodiments, the style input component indicates a selection of a first style, the target media content can be generated based on a process that includes: providing a set of fine-tuning parameters corresponding to the first style to a generation model, the set of fine-tuning parameters corresponding to a fine-tuning model trained based on a set of sample media corresponding to the first style; and obtaining the target media content generated by the generation model based on the set of fine-tuning parameters.
[0072] In some embodiments, the style input component indicates a selection of a second style, the target media content can be generated based on a process that includes: generating prompt information based on the second style, the prompt information including description information corresponding to the second style; and obtaining the target media content generated by a generation model based on the prompt information.
[0073] In some embodiments, the generation control information includes at least one of: reference material obtained via the material input component, text content obtained via the text input component, and / or a target style determined via the style selection component.
[0074] In some embodiments, providing the target media content generated based on the target generation mode includes: presenting the target media content and at least a portion of the generation control information in a viewing interface.
[0075] In some embodiments, the electronic device 110 can further receive a first prompt word input by a user via the text input component; in response to receiving a recommendation request, present a set of candidate prompt words associated with the first prompt word; and based on a selection of a second prompt word from the set of candidate prompt words, update the first prompt word.
[0076] In some embodiments, the electronic device 110 can further receive the recommendation request in response to a triggering of a recommendation control in the text input component.
[0077] In some embodiments, the electronic device 110 can further generate the set of candidate prompt words based on at least one reference material obtained via the material input component.
[0078] In some embodiments, the generation of the set of candidate prompt words is further based on an object detection result of the at least one reference material, the object detection result indicating whether the at least one reference material includes an object of a predetermined type.
[0079] Example apparatuses and devices
[0080] Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above method or process. FIG. 4 shows a schematic structural block diagram of an apparatus 400 for media content generation according to some embodiments of the present disclosure.
[0081] As shown in FIG. 4, the apparatus 400 can include a presentation module 410 configured to present an editing interface including a material input component, a text input component, and a style selection component; a determination module 420 configured to determine a target generation mode based on input states of the material input component, the text input component, and the style input component in response to receiving a media generation request; and a provision module 430 configured to provide target media content generated based on the target generation mode based on generation control information received in the editing interface.
[0082] In some embodiments, the input state indicates whether input content of the material input component, the text input component, or the style input component is empty.
[0083] In some embodiments, the target generation mode indicates at least one of: a generation model used for the target media content and / or at least one control parameter for the generation model.
[0084] In some embodiments, the style input component indicates a selection of a first style, and the target media content can be generated based on a process of: providing a set of fine-tuning parameters corresponding to the first style to a generation model, the set of fine-tuning parameters corresponding to a fine-tuning model trained based on a set of sample media corresponding to the first style; and obtaining the target media content generated by the generation model based on the set of fine-tuning parameters.
[0085] In some embodiments, the style input component indicates a selection of a second style, and the target media content can be generated based on a process of: generating prompt information based on the second style, the prompt information including description information corresponding to the second style; and obtaining the target media content generated by a generation model based on the prompt information.
[0086] In some embodiments, the generation control information includes at least one of: reference material obtained via the material input component, text content obtained via the text input component, and / or a target style determined via the style selection component.
[0087] In some embodiments, the provision module 430 is further configured to present the target media content and at least a portion of the generation control information in a viewing interface.
[0088] In some embodiments, the apparatus 400 can further include a receiving module configured to receive, via the text input component, a first prompt word input by a user; a second presenting module configured to present, in response to the received recommendation request, a set of candidate prompt words associated with the first prompt word; and an updating module configured to update the first prompt word based on a selection of a second prompt word from the set of candidate prompt words.
[0089] In some embodiments, the electronic device 110 can further include a second receiving module configured to receive, in response to a triggering of a recommendation control in the text input component, the recommendation request.
[0090] In some embodiments, the electronic device 110 can further include a generating module configured to generate, in response to obtaining at least one reference material via the material input component, the set of candidate prompt words based on the at least one reference material.
[0091] In some embodiments, the generation of the set of candidate prompt words is further based on an object detection result of the at least one reference material, the object detection result indicating whether the at least one reference material includes an object of a predetermined type.
[0092] FIG. 5 illustrates a block diagram of a computing device / server 500 in which one or more embodiments of the present disclosure can be implemented. It should be understood that the computing device / server 500 illustrated in FIG. 5 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein.
[0093] As shown in FIG. 5, the computing device / server 500 is in the form of a general-purpose computing device. The components of the computing device / server 500 can include, but are not limited to, one or more processors or processing units 510, memory 520, storage 530, one or more communication units 540, one or more input devices 560, and one or more output devices 560. The processing unit 510 can be a real or virtual processor and is capable of executing various processing in accordance with programs stored in the memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device / server 500.
[0094] The computing device / server 500 typically includes multiple computer storage media. Such media can be any available media accessible to the computing device / server 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data (e.g., training data for training) and accessible within the computing device / server 500.
[0095] The computing device / server 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0096] The communication unit 540 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device / server 500 can be implemented as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device / server 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0097] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. The computing device / server 500 can also communicate as needed with one or more external devices (not shown) via communication unit 540. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with the computing device / server 500, or with any device (e.g., network card, modem, etc.) that enables the computing device / server 500 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interfaces (not shown).
[0098] According to an example implementation of the present disclosure, a computer-readable storage medium is provided having stored thereon one or more computer instructions that, when executed by a processor, implement the method described above.
[0099] Various aspects of the disclosure are now described with reference to the drawings. In general, the drawings described below are diagrammatic and schematic representations of actual or conceptual structures and processes, and are not limiting of the scope of the present disclosure. In the drawings, the same reference numerals are used to represent similar or like items. Unless otherwise noted, the drawings are not to scale and are intended to conceptually illustrate the structures and procedures described herein. In the description that follows, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent, however, to one of ordinary skill in the art that the specific detail need not be used to practice the present disclosure. In general, the drawings and description are not intended to confine the scope of the present disclosure to the specific embodiments disclosed. The scope of the present disclosure is intended to encompass all alternatives, modifications and equivalents of the specific embodiments disclosed. The drawings and description are intended to serve as a complete description of the present disclosure and are not intended to manifest the scope of the present disclosure. The present disclosure is defined only by the appended claims, which are to be interpreted in the light of the proper construction of a patent claim, and any definitions in the claims are to be understood according to the rules of claim construction set forth in 37 C.F.R. § 1.57 and 37 C.F.R. § 1.84.
[0100] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can be coupled to a computer or other programmable data processing apparatus, such that the computer readable storage medium can provide computer readable program instructions to the computer or other programmable data processing apparatus, which can implement the functions / acts specified in the flowchart and / or block diagram block or blocks. The computer readable program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process, such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0101] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can be coupled to a computer or other programmable data processing apparatus, such that the computer readable storage medium can provide computer readable program instructions to the computer or other programmable data processing apparatus, which can implement the functions / acts specified in the flowchart and / or block diagram block or blocks. The computer readable program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process, such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0102] The flow diagrams and block diagrams in the drawings are representative of the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various implementations of the present disclosure. In this regard, each block in the flow diagrams and block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical functions ("instructions"). In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or combinations of special-purpose hardware and computer instructions.
[0103] Having described various implementations of the disclosure above, the descriptions are not exhaustive and do not limit the implementations to the disclosed implementations. Numerous modifications and adaptations thereof will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The scope of the implementations includes any other implementation from which the following teachings can be practised. The language used in the specification has been principally selected for readability and instructional purposes and it can not have been selected to delineate or circumscribe the inventive subject matter. Accordingly, the disclosure is intended to be illustrative, but not limiting, of the described implementations.
Claims
1. A method for generating media content, comprising: An editing interface is presented, which includes a material input component, a text input component, and a style selection component; In response to receiving a media generation request, a target generation mode is determined based on the input states of the material input component, the text input component, and the style input component; as well as Based on the generation control information received in the editing interface, target media content generated based on the target generation mode is provided.
2. The method according to claim 1, wherein the input state indicates whether the input content of the material input component, the text input component, or the style input component is empty.
3. The method of claim 1, wherein the target generation mode indicates at least one of the following: A generation model for the target media content; At least one control parameter for the generated model.
4. The method of claim 1, wherein the style input component indicates the selection of a first style, and the target media content is generated based on the following process: Provide the generative model with a set of fine-tuning parameters corresponding to the first style, the set of fine-tuning parameters corresponding to the fine-tuning model, the fine-tuning model being trained based on a set of sample media corresponding to the first style; and Obtain the target media content generated by the generation model based on the set of fine-tuning parameters.
5. The method of claim 1, wherein the style input component indicates the selection of a second style, and the target media content is generated based on the following process: Based on the second style, a prompt message is generated, the prompt message including descriptive information corresponding to the second style; and Obtain the target media content generated by the generation model based on the prompt information.
6. The method of claim 1, wherein the generation of control information includes at least one of the following: Reference materials obtained via the aforementioned material input component, The text content obtained through the text input component The target style determined by the style selection component.
7. The method of claim 1, wherein providing the target media content generated based on the target generation mode comprises: The target media content and at least a portion of the generation control information are presented in the viewing interface.
8. The method according to claim 1, further comprising: The text input component receives the first prompt word entered by the user; In response to the received recommendation request, a set of candidate suggestions associated with the first suggestion is presented; as well as The first prompt word is updated based on the selection of the second prompt word from the set of candidate prompt words.
9. The method according to claim 1, further comprising: In response to the triggering of the recommendation control in the text input component, the recommendation request is received.
10. The method of claim 8, further comprising: In response to obtaining at least one reference material via the material input component, the set of candidate prompt words is generated based on the at least one reference material.
11. The method of claim 10, wherein the generation of the set of candidate prompts is further based on the object detection result of the at least one reference material, the object detection result indicating whether the at least one reference material includes an object of a predetermined type.
12. An apparatus for generating media content, comprising: The presentation module is configured to present an editing interface, which includes a material input component, a text input component, and a style selection component. The determination module is configured to, in response to receiving a media generation request, determine a target generation mode based on the input states of the material input component, the text input component, and the style input component; as well as The module is configured to provide target media content generated based on the target generation mode, based on the generation control information received in the editing interface.
13. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 11.
14. A computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Media content display method and device, equipment and storage medium
CN116704078A
Rich media-based digital human report video generation method and system
CN117131210A
Method and device for displaying works, equipment and storage medium
CN117492601A
Method of and system for controlling the qualities of musical energy embodied in and expressed by digital music to be automatically composed and generated by an automated music composition and generation engine
US20190237051A1