Method and apparatus for generating media content, device, and storage medium
By acquiring and processing visual information from descriptive text to generate media content, the problem of low generation efficiency in existing platforms is solved, enabling efficient generation of media content while maintaining consistency with existing content.
Patent Information
- Application Number
- PCT/CN2025/102613
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-21
- Filing Date
- 2025-06-20
- Publication Date
- 2025-12-26
AI Technical Summary
Existing application platforms struggle to efficiently generate media content and fail to meet user needs.
By acquiring the first descriptive text, generating multiple second descriptive texts using the first processing entity to describe the visual information of the media content to be generated, and generating media materials using the second processing entity, the target media content is finally generated.
It improves the efficiency of media content generation, meets user needs, and maintains consistency with existing media content in terms of plot and character design.
Smart Images

Figure CN2025102613_26122025_PF_FP_ABST
Abstract
Description
Methods, apparatus, devices and storage media for generating media content
[0001] This application claims priority to Chinese Patent Application No. 202410813920.9, filed on June 21, 2024, entitled “Method, Apparatus, Device and Storage Medium for Generating Media Content”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] The exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to methods, apparatus, devices, and computer-readable storage media for generating media content. Background Technology
[0003] In recent years, with the rapid development of the internet, more and more users are publishing or viewing media content on various application platforms. However, existing application platforms are struggling to meet the needs of users generating media content. Summary of the Invention
[0004] In a first aspect of this disclosure, a method for generating media content is provided. The method includes: obtaining a first descriptive text about the media content to be generated; generating multiple second descriptive texts based on the first descriptive text using a first processing entity, the multiple second descriptive texts describing visual information of the media content to be generated; generating multiple media materials corresponding to the multiple second descriptive texts using a second processing entity; and generating target media content based on the multiple media materials.
[0005] In a second aspect of this disclosure, an apparatus for generating media content is provided. The apparatus includes: an acquisition module configured to acquire first descriptive text about media content to be generated; a first processing module configured to generate multiple second descriptive texts based on the first descriptive text using a first processing entity, the multiple second descriptive texts describing visual information of the media content to be generated; a second processing module configured to generate multiple media materials corresponding to the multiple second descriptive texts using a second processing entity; and a generation module configured to generate target media content based on the multiple media materials.
[0006] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
[0008] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method of the first aspect.
[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented;
[0012] Figure 2 illustrates a flowchart of an example process for generating media content according to some embodiments of the present disclosure;
[0013] Figure 3 illustrates a schematic diagram of an example framework for a media content generation system according to some embodiments of the present disclosure;
[0014] Figures 4A and 4B illustrate example interfaces according to some embodiments of the present disclosure;
[0015] Figure 5 shows a schematic structural block diagram of an example apparatus for generating media content according to some embodiments of the present disclosure; and
[0016] Figure 6 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure. Detailed Implementation
[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0018] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0019] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0020] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0021] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0022] As mentioned above, with the rapid development of the internet, more and more users are publishing or viewing media content on various application platforms. However, the efficiency of generating media content on existing application platforms is relatively low, making it difficult to meet user needs.
[0023] Embodiments of this disclosure propose a scheme for generating media content. According to this scheme, a first descriptive text about the media content to be generated can be obtained; a first processing entity can generate multiple second descriptive texts based on the first descriptive text, the multiple second descriptive texts describing the visual information of the media content to be generated; a second processing entity can generate multiple media materials corresponding to the multiple second descriptive texts; and a target media content can be generated based on the multiple media materials.
[0024] In this way, the embodiments of this disclosure can generate media content corresponding to a media generation request using multiple processing entities based on the media generation request, thereby improving the efficiency of media content generation and meeting user needs.
[0025] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.
[0026] Example Environment
[0027] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In environment 100, an electronic device 110 and a media content generation system 136 are deployed. In some embodiments, the electronic device 110 receives a media generation request 130 from a user 140. Subsequently, the electronic device 110 invokes the media content generation system 136 to generate target media content 120 based on the media generation request 130. In some embodiments, the media content generation system 136 may run on a local device or a remote device.
[0028] In some embodiments, electronic device 110 may include various types of computing systems / servers providing computing capabilities, and electronic device 110 may include terminal devices. Such terminal devices may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 may also support any type of user-facing interface (such as "wearable" circuitry). Electronic device 110 may, for example, include computing systems / servers such as mainframes, edge computing nodes, computing devices in cloud environments, virtual machines, etc. Although shown as a single device, electronic device may include multiple physical devices.
[0029] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0030] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.
[0031] Generate media content
[0032] Figure 2 shows a flowchart of an example process 200 for generating media content according to some embodiments of the present disclosure. Process 200 can be implemented at electronic device 110. Process 200 will now be described with reference to Figure 1.
[0033] As shown in Figure 2, in box 210, electronic device 110 acquires a first descriptive text about the media content to be generated.
[0034] In some embodiments, the electronic device 110 may acquire a first descriptive text (e.g., script content) input by the user. In some embodiments, the first descriptive text may also be determined based on the user's selection of existing text. For example, the user may select existing script content as the first descriptive text.
[0035] Additionally, the first descriptive text can also be generated based on user input or user selection.
[0036] The following describes an example process for generating media content according to an embodiment of the present disclosure with reference to the media content generation system 136 shown in FIG3.
[0037] Figure 3 illustrates a schematic diagram of an example framework 300 of a media content generation system 136 according to some embodiments of the present disclosure. As shown in the example framework 300 of Figure 3, the media content generation system 136 may include a text expansion entity 310 (optional), a target description entity 320 (e.g., a plot description entity), a picture description entity 330, an audio generation entity 340 (optional), and a picture generation entity 350. As an example, the multiple entities mentioned herein can be implemented as appropriate machine learning models to generate corresponding output results based on the received content. As an example, the text expansion entity 310, the target description entity 320, and the picture description entity can be implemented as language models with different text processing capabilities. The audio generation entity 340 can be implemented as a speech model that generates audio content based on text. The picture generation entity 350 can be implemented as a visual model that generates image content based on text. This disclosure is not intended to limit the specific construction method of the entities. The processing capabilities of each entity are described in the embodiments disclosed later, and will not be repeated here.
[0038] In some embodiments, referring to FIG3, the electronic device 110 can utilize the target description entity 320 to generate a first descriptive text about the media content to be generated based on a media generation request. As an example, the electronic device 110 can receive a media generation request 130 from a user 140, and invoke the target description entity 320 to generate a first descriptive text (e.g., a script) of the media content to be generated corresponding to the media generation request 130. As an example, the media generation request 130 may include prompts, such as requirements for the storyline, character design, and art style (e.g., ink painting style, realistic style, etc.) of the media content to be generated. As an example, the first descriptive text may include, for example, a description of the storyline and character features of the media content to be generated. As an example, the media content may be, for example, video content, graphic content (e.g., comics), etc.
[0039] In some embodiments, continuing to refer to FIG3, the electronic device 110 may, in response to receiving a media generation request 130, expand the prompt item of the media generation request using the text expansion entity 310 to generate expanded text content associated with the media generation request 130. It is understood that the expanded text content is derived from the media generation request 130, and is more complete and specific than the prompt item in the media generation request 130.
[0040] In some embodiments, continuing to refer to FIG3, the electronic device 110 may utilize the target description entity 320 to process extended text content and generate a first description text about the media content to be generated.
[0041] In some embodiments, continuing to refer to FIG3, the electronic device 110 may also obtain reference information associated with the media generation request.
[0042] In some embodiments, the reference information may include reference media content associated with the media generation request 130. As an example, the media generation request 130 may be a continuation request for the reference media content. That is, the media generation request 130 may create a media content generation request that continues the plot of the reference media content based on the storyline and / or character portrayals corresponding to the reference media content.
[0043] In some embodiments, the media generation request 130 may be a request issued by user 140 in a video continuation scenario. User 140 may, for example, request continuation of reference media content. The reference media content may be included in a media content set, which stores multiple media content items, including the reference media content.
[0044] In some embodiments, a media content set is associated with multiple media contents, which are organized based on a tree structure of a directed acyclic graph. The tree structure includes multiple nodes corresponding to the multiple media contents, and the edges in the tree structure indicate the content association between the media contents corresponding to two nodes (e.g., plot development relationship, such as the plot development of media content A being the plot development of media content B). As an example, the electronic device 110 can traverse the tree structure based on a depth-first search algorithm to determine at least one media content associated with a reference media content. In some embodiments, the reference information may include a reference media content associated with the media generation request 130 and at least one media content associated with the reference media content.
[0045] In some embodiments, continuing to refer to FIG3, the electronic device 110 can utilize the text extension entity 310 to generate extended text content based on reference information and the media generation request 130. As an example, the extended text content may include character descriptions (e.g., physical features, clothing features, etc.) and scene descriptions (e.g., architectural features, etc.) from the reference media content, thereby ensuring that the subsequently generated media content remains consistent with the relevant descriptions in the reference media content. In this way, the extended text content can take into account both the media generation request and the reference media content, enabling the generated media content to continue the character characteristics, scene features, and plot development of the previous media content while satisfying the media generation request 130.
[0046] As shown in Figure 2, in box 220, the electronic device uses a first processing entity (e.g., screen description entity 330) to generate multiple second description texts based on the first description text. These multiple second description texts can describe the visual information of the media content to be generated (e.g., a description of the screen). As an example, compared to the first description text, the multiple second description texts contain more detailed descriptions of the screen. For instance, the first description text may mention a first scene, and the second description text can add more detailed descriptions of the first scene. For example, the second description text may include elements present in the first scene (e.g., trees, buildings, etc.), the overall screen style, etc., so that the subsequent model can generate the corresponding screen based on the multiple second description texts.
[0047] In some embodiments, the first descriptive text may include plot description text for a first chapter, and the second descriptive text may include scene description text corresponding to multiple storyboard segments of the first chapter. As an example, the multiple second descriptive texts may, for example, describe multiple storyboard segments of the media content to be generated. Multiple storyboard segments can, for example, be understood as visual representations of multiple scenes in the media content to be generated.
[0048] In some embodiments, continuing to refer to FIG3, the electronic device 110 can acquire context information associated with the first chapter, the context information being determined based on at least one second chapter associated with the first chapter. As an example, the at least one second chapter can be a preceding chapter of the first chapter. That is, the first chapter can be a new creation based on the plot content of at least one second chapter.
[0049] In some embodiments, continuing to refer to FIG3, the electronic device 110 can utilize a first processing entity to generate multiple second description texts based on the plot description text and context information. As an example, the second description texts can be obtained by supplementing the scene description content based on the plot description text and context information.
[0050] In some embodiments, continuing to refer to FIG3, the contextual information may indicate: character setting information associated with at least one second chapter. As an example, character setting information may include a visual content description of the character (e.g., the character's physical characteristics, clothing characteristics, etc.), a personality content description of the character (e.g., kind, brave, etc.), or an information content description of the character (e.g., gender, age, story, etc.).
[0051] In some embodiments, continuing to refer to FIG3, the context information may indicate scene setting information associated with at least one second chapter. As an example, scene setting information may include screen background, lighting effects, or scene elements (e.g., buildings, props, etc.).
[0052] In some embodiments, continuing to refer to FIG3, the contextual information may indicate: plot description text associated with at least one second chapter. As an example, the plot description text may include storyline, background information, or major events, etc.
[0053] As shown in Figure 2, in box 230, electronic device 110 uses a second processing entity to generate multiple media materials corresponding to multiple second descriptive texts. As an example, as shown in Figure 3, the third processing entity can be screen generation entity 350.
[0054] In some embodiments, continuing to refer to FIG3, the first descriptive text includes a character description. The electronic device 110 can generate a visual representation of at least one character based on the character description. Further, the electronic device 110 can utilize a second processing entity (e.g., image generation entity 350) to generate multiple media materials corresponding to multiple second descriptive texts based on the visual representation of at least one character. By pre-generating the visual representation of the character, the visual representation of the character can remain consistent across multiple media materials.
[0055] In some embodiments, the multiple media materials may include: multiple video clips corresponding to multiple second descriptive texts, or multiple images corresponding to multiple second descriptive texts.
[0056] As shown in Figure 2, in box 240, electronic device 110 generates target media content based on multiple media materials. As an example, electronic device 110 can merge multiple video clips into a complete video content as the target media content.
[0057] In some embodiments, continuing to refer to FIG3, the electronic device 110 may also utilize the audio generation entity 340 to generate audio content of the media content to be generated based on the first descriptive text. As an example, the audio content may include background music (e.g., background music) of the media content to be generated. As an example, the audio content may also include voice-over (e.g., dialogue voices of characters in the media content, etc.) of the media content to be generated.
[0058] In some embodiments, continuing to refer to FIG3, the electronic device 110 can generate target media content 360 based on audio content and multiple media materials. As an example, the electronic device 110 can match audio content with corresponding media materials and merge the audio content and multiple media materials to generate target media content 360.
[0059] In some embodiments, continuing to refer to FIG3, the electronic device 110 may also create a node corresponding to the target media content 360 in the media content set to store the target media content 360 in the media content set.
[0060] Figures 4A and 4B illustrate example interfaces 400A to 400B according to some embodiments of the present disclosure. Interfaces 400A to 400B may, for example, be provided by a client.
[0061] In some embodiments, as shown in interface 400A of FIG4A, the client can obtain a media generation request initiated by user 140 based on input control 405. This media generation request may, for example, indicate prompts entered by the user, such as script description text about the chapter content to be generated.
[0062] As an example, the media content to be generated can correspond to the first chapter of a collection of works. Further, electronic device 110 (e.g., client or server) can use text extension entity 310 to extend the prompts and generate extended text content. Electronic device 110 can use target description entity 320 to generate a first description text for the first chapter based on the extended text content and previous chapters associated with the first chapter. Further, scene description entity 330 can generate multiple second description texts corresponding to multiple storyboard scenes based on the first description text. Further, scene generation entity 350 can generate multiple media materials (e.g., multiple video clips or multiple images) based on the multiple second description texts. Additionally, audio content (e.g., voice-over, background music) can be generated based on the first description text using audio generation entity 340. Finally, electronic device 110 can generate target media content based on multiple media materials and audio content.
[0063] In some embodiments, as shown in interface 400B of FIG4B, the client can present target media content 360 on interface 400B. Such target media content 360 corresponds, for example, to a new chapter in a collection. As an example, a chapter may include comic content consisting of multiple images (generated based on multiple storyboard description texts) corresponding to multiple storyboards. Alternatively, the chapter may include video content consisting of multiple video clips (generated based on multiple storyboard description texts) corresponding to multiple storyboards.
[0064] Based on the interaction process described above, embodiments of this disclosure can generate media content that meets the user's media generation request. Furthermore, embodiments of this disclosure can also expand existing media content by combining it with existing media content (e.g., reference media content and / or existing chapter content), ensuring that the generated media content continues the existing media content in terms of plot and maintains consistency in character design and scene features. In addition, embodiments of this disclosure can incorporate user interaction content, making the generated target media content better meet the user's needs. Therefore, the efficiency of media content generation is improved.
[0065] Example devices and equipment
[0066] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 5 shows a schematic structural block diagram of an example interactive apparatus 500 according to certain embodiments of this disclosure. Apparatus 500 may be implemented as or included in electronic device 110. The various modules / components in apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.
[0067] As shown in Figure 5, the device 500 includes an acquisition module 510 configured to acquire a first descriptive text about the media content to be generated; a second processing module 520 configured to generate multiple second descriptive texts based on the first descriptive text using a first processing entity, wherein the multiple second descriptive texts describe the visual information of the media content to be generated; a third processing module 530 configured to generate multiple media materials corresponding to the multiple second descriptive texts using a third processing entity; and a generation module 540 configured to generate target media content based on the multiple media materials.
[0068] In some embodiments, the apparatus 500 further includes a target description module, which is configured to: generate a first description text about the media content to be generated based on a media generation request using a target description entity.
[0069] In some embodiments, the target description module is further configured to: generate a first descriptive text about the media content to be generated based on the media generation request using a target description entity, including: expanding the prompt item of the media generation request using a text expansion entity to generate expanded text content; and processing the expanded text content using a first processing entity to generate the first descriptive text about the media content to be generated.
[0070] In some embodiments, the target description module is further configured to: extend the prompts of the media generation request with text extension entities to generate extended text content, including: obtaining reference information associated with the media generation request; and generating extended text content based on the reference information and prompts using text extension entities.
[0071] In some embodiments, the target description module is further configured to include reference information including reference media content associated with the media generation request.
[0072] In some embodiments, the target description module is further configured such that: the prompt describes first plot content, and the extended text content includes second plot content generated by extending the first plot content.
[0073] In some embodiments, the target description module is further configured to: the first description text includes the character description content of the character, and the second processing module 530 is further configured to: generate multiple media materials corresponding to multiple second description texts using the second processing entity, including: generating a visual representation of at least one character based on the character description content; and generating multiple media materials corresponding to multiple second description texts based on the visual representation of at least one character using the second processing entity.
[0074] In some embodiments, the second processing module 530 is further configured to include: multiple media materials including: multiple video clips corresponding to multiple second descriptive texts; or multiple images corresponding to multiple second descriptive texts.
[0075] In some embodiments, the apparatus 500 further includes an audio processing module configured to generate audio content of the media content to be generated based on the first descriptive text using an audio generation entity. The generation module 540 is further configured to generate target media content based on the audio content and multiple media materials.
[0076] In some embodiments, the first processing module 510 is further configured to: the first descriptive text includes the plot description text of the first chapter; the second processing module 520 is further configured to: the second descriptive text includes the scene description text corresponding to the multiple storyboard segments of the first chapter.
[0077] In some embodiments, the second processing module 520 is further configured to: generate multiple second description texts based on the first description text using the first processing entity, including: obtaining context information associated with the first chapter, the context information being determined based on at least one second chapter associated with the first chapter; and generating multiple second description texts based on the plot description text and the context information using the first processing entity.
[0078] In some embodiments, the second processing module 520 is further configured to: indicate at least one of the following: character setting information associated with at least one second chapter; scene setting information associated with at least one second chapter; and plot description text associated with at least one second chapter.
[0079] In some embodiments, the device 500 further includes a storage module configured to: create nodes corresponding to target media content in a media content set, the media content set being associated with multiple media contents, the multiple media contents being organized based on a tree structure of a directed acyclic graph, the tree structure including multiple nodes corresponding to the multiple media contents, and the edges in the tree structure indicating the content association between media contents corresponding to two nodes.
[0080] Figure 6 shows a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 600 shown in Figure 6 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 600 shown in Figure 6 can be used to implement the electronic device 110 of Figure 1.
[0081] As shown in Figure 6, the electronic device 600 is in the form of a general-purpose electronic device. Components of the electronic device 600 may include, but are not limited to, one or more processing units or processors 610, memory 620, storage devices 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processor 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 620. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 600.
[0082] Electronic device 600 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 600.
[0083] Electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0084] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0085] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 600, or with any device that enables electronic device 600 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0086] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0087] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0088] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0089] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0091] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1.A method for generating media content, comprising: obtaining a first description text about a media content to be generated; generating, by a first processing entity, a plurality of second description texts based on the first description text, the plurality of second description texts describing visual information of the media content to be generated; generating, by a second processing entity, a plurality of media materials corresponding to the plurality of second description texts; and generating a target media content based on the plurality of media materials. 2.The method of claim 1, wherein the obtaining a first description text about a media content to be generated comprises: generating, by a target description entity, the first description text about the media content to be generated based on a media generation request. 3.The method of claim 2, wherein the generating, by a target description entity, the first description text about the media content to be generated based on a media generation request comprises: extending, by a text extension entity, a prompt item of the media generation request to generate extended text content; and processing, by the target description entity, the extended text content to generate the first description text about the media content to be generated. 4.The method of claim 3, wherein the extending, by a text extension entity, a prompt item of the media generation request to generate extended text content comprises: obtaining reference information associated with the media generation request; and generating, by the text extension entity, the extended text content based on the reference information and the prompt item. 5.The method of claim 4, wherein the reference information comprises reference media content associated with the media generation request. 6.The method of claim 3, wherein the prompt item describes a first plot content, and the extended text content comprises a second plot content generated by extending the first plot content. 7.The method of claim 1, wherein the first description text comprises a role description content of a role, and the generating, by a second processing entity, a plurality of media materials corresponding to the plurality of second description texts comprises: generating a visual representation of at least one role based on the role description content; and generating, by the second processing entity, the plurality of media materials corresponding to the plurality of second description texts based on the visual representation of the at least one role. 8.The method of claim 1, wherein the plurality of media materials comprises: a plurality of video clips corresponding to the plurality of second description texts; or a plurality of pictures corresponding to the plurality of second description texts. 9.The method of claim 1, wherein the generating a target media content based on the plurality of media materials comprises: generating, by an audio generation entity, an audio content of the media content to be generated based on the first description text; and generating the target media content based on the audio content and the plurality of media materials. 10.The method of claim 1, wherein the first description text comprises a plot description text of a first chapter, and the second description text comprises a picture description text corresponding to a plurality of split shot segments of the first chapter. 11.The method of claim 10, wherein generating, by the first processing entity, the second description texts based on the first description text comprises: obtaining context information associated with the first chapter, the context information being determined based on at least one second chapter associated with the first chapter; and generating, by the first processing entity, the second description texts based on the scenario description text and the context information. 12.The method of claim 11, wherein the context information indicates at least one of: character setting information associated with the at least one second chapter; scene setting information associated with the at least one second chapter; scenario description text associated with the at least one second chapter. 13.The method of claim 1, further comprising: creating a node corresponding to the target media content in a media content set, the media content set being associated with a plurality of media contents, the plurality of media contents being organized based on a tree structure of a directed acyclic graph, the tree structure comprising a plurality of nodes corresponding to the plurality of media contents, and edges in the tree structure indicating content association between media contents corresponding to two nodes. 14.An apparatus for generating media content, comprising: an obtaining module configured to obtain a first description text about a media content to be generated; a first processing module configured to generate, by a first processing entity, a plurality of second description texts based on the first description text, the plurality of second description texts describing visual information of the media content to be generated; a second processing module configured to generate, by a second processing entity, a plurality of media materials corresponding to the plurality of second description texts; and a generating module configured to generate a target media content based on the plurality of media materials. 15.An electronic device, comprising: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the electronic device to perform the method according to any one of claims 1 to 13. 16.A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 13. 17.A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Video generation method and video generation system
CN112312189A
Generation method and device of interactive multimedia content, electronic equipment and storage medium
CN117633258A
Method and device for generating video describing entity, equipment and medium
CN117793482A
Method and device for generating media content, equipment and storage medium
CN118870143A
Style-based dynamic content generation
US20230154082A1