Method and device for generating media content, equipment and storage medium

CN121753071APending Publication Date: 2026-03-27BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies struggle to generate media content that users expect based on virtual avatars, and thus cannot effectively match user needs.

Method used

By receiving a generation request, the system generates descriptive text for reference media content and determines the screen structure. Combined with the image data of virtual objects, it generates target media content.

Benefits of technology

This allows the generated media content to better match user needs, and by combining virtual objects and reference media content, the relevance and aesthetics of the content are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121753071A_ABST
    Figure CN121753071A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method and device for generating media content, equipment and a storage medium. The method provided herein includes: receiving a generation request indicating a reference media content and a virtual object associated with a target user, the virtual object being created based on a configuration operation of the target user; generating a description text of the reference media content and determining a picture structure of the reference media content; and generating the target media content based on the description text, the picture structure and the image data of the virtual object. In this way, according to the embodiment of the invention, the generated target media content can be better matched with user requirements.
Need to check novelty before this filing date? Find Prior Art

Description

A method, apparatus, device, and storage medium for generating media content. Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to a method, apparatus, device, and computer-readable storage medium for generating media content. Background Technology

[0002] In recent years, with the development of the internet, more and more people are engaging in online activities on online platforms. For example, people use media platforms to view various media works and create virtual avatars. People expect to generate desired media content based on virtual avatars.

[0003] Summary of the Invention

[0004] In a first aspect of this disclosure, a method for generating media content is provided, comprising: receiving a generation request, the generation request indicating reference media content and a virtual object associated with a target user, the virtual object being created based on configuration operations of the target user; generating descriptive text of the reference media content and determining the screen structure of the reference media content; and generating target media content based on the descriptive text, the screen structure, and image data of the virtual object.

[0005] In a second aspect of this disclosure, an apparatus for generating media content is provided. The apparatus includes: a receiving module configured to receive a generation request, the generation request indicating reference media content and a virtual object associated with a target user, the virtual object being created based on configuration operations of the target user; a parsing module configured to generate descriptive text of the reference media content and determine the screen structure of the reference media content; and a first generation module configured to generate target media content based on the descriptive text, the screen structure, and image data of the virtual object.

[0006] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1 shows a schematic diagram of an example environment in which some embodiments of the present disclosure can be implemented;

[0011] Figure 2 illustrates a flowchart of a process for generating media content according to some embodiments of the present disclosure;

[0012] Figures 3A to 3C show renderings of generated media content according to some embodiments of the present disclosure;

[0013] Figure 4 shows a schematic structural block diagram of an example apparatus for generating media content according to some embodiments of the present disclosure; and

[0014] Figure 5 shows a block diagram of an apparatus capable of implementing several embodiments of the present disclosure. Detailed Implementation

[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0016] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0017] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0018] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0019] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0020] As briefly mentioned above, with the development of the internet, more and more people are engaging in online activities on online platforms. For example, people use media platforms to view various media works and create virtual avatars. People expect to generate desired media content based on virtual avatars.

[0021] Embodiments of this disclosure propose a scheme for generating media content. According to various embodiments of this disclosure, a generation request is received, indicating reference media content and a virtual object associated with a target user, the virtual object being created based on the target user's configuration operations; descriptive text of the reference media content is generated and the screen structure of the reference media content is determined; and target media content is generated based on the descriptive text, the screen structure, and the image data of the virtual object.

[0022] In this way, embodiments of this disclosure can generate media content associated with user-provided reference media content and user-created virtual objects. Based on this, embodiments of this disclosure can combine user-associated virtual objects and reference media content, enabling the generated media content to better match user needs.

[0023] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0024] Example Environment

[0025] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in Figure 1, the example environment 100 may include an electronic device 130 and a terminal device 110.

[0026] In this example environment 100, terminal device 110 may run an application 120 that supports the generation of media content. Application 120 may be any suitable type of application for generating media content, and examples may include, but are not limited to, video applications, social applications, or other suitable applications. User 140 may interact with application 120 via terminal device 110 and / or its attached devices.

[0027] In environment 100 of Figure 1, if application 120 is active, electronic device 130 can present interface 150 for supporting the generation of media content through application 120.

[0028] In environment 100 of Figure 1, electronic device 130 can receive a generation request sent by user 140 via terminal device 110. Furthermore, electronic device 130 can generate target media content corresponding to the generation request based on the generation request.

[0029] In some embodiments, terminal device 110 communicates with electronic device 130 to provide services to application 120. Terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 can also support any type of user-facing interface (such as "wearable" circuitry).

[0030] Electronic device 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Electronic device 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Electronic device 130 can provide backend services for the application 120 in terminal device 110 that supports generating media content.

[0031] A communication connection can be established between electronic device 130 and terminal device 110. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections; the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, electronic device 130 and terminal device 110 can achieve signaling interaction through their communication connection.

[0032] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0033] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0034] Example process for generating media content

[0035] Figure 2 shows a flowchart of a process 200 for generating media content according to some embodiments of the present disclosure. Process 200 can be implemented at an electronic device 130. Process 200 will now be described with reference to Figure 1.

[0036] As shown in Figure 2, in box 210, electronic device 130 receives a generation request from user 140 via terminal device 110. This generation request is for generating media content. The generation request received by electronic device 130 indicates reference media content and a virtual object associated with the target user, which is created based on the target user's configuration operations. The reference media content can be image content as shown in Figure 3A.

[0037] In some examples, virtual objects can be generated based on visual data and user-input descriptive text. The user-input descriptive text may include character setting information, knowledge information, visual information, etc., enabling the virtual object to possess certain interactive capabilities and personality traits based on this configuration information. Virtual objects can include their visual representation and their capabilities (e.g., dialogue capabilities, data processing capabilities). For example, users on a platform can create corresponding virtual objects by configuring the platform. In some scenarios, such virtual objects may also be referred to as digital avatars or virtual clones.

[0038] In other examples, obtaining configuration information for a virtual object may include: electronic device 130 acquiring at least one image based on a user's configuration operation. Configuration information for the virtual object is then determined based on the acquired at least one image. In some scenarios, the configuration information for the virtual object may be referred to as the virtual object's image data. As an example, the configuration operation may include selecting an image from a photo album or capturing an image using a camera component. As an example, the at least one image may be an image from a photo album or an image captured by the camera component. As an example, the visual representation of the virtual object may be generated based on the image data and descriptive text input by the user.

[0039] In box 220, electronic device 130 generates descriptive text for reference media content and determines the screen structure of the reference media content.

[0040] Specifically, after receiving a generation request, the electronic device 130 can obtain reference media content corresponding to the generation request. Based on the reference media content, the electronic device 130 can generate descriptive text corresponding to the reference media content.

[0041] In some embodiments, the electronic device 130 can perform edge detection processing on the reference media content to determine the edge information of the reference media content. Further, the electronic device 130 can determine the screen structure of the reference media content based on the edge information. It should be noted that the electronic device 130 can input the reference media content into a screen processing model to perform edge detection processing on the reference media content. As an example, the screen processing model can be implemented as a machine learning model capable of acquiring image edge information for image content. This invention does not intend to limit the specific content and training process of the screen processing model.

[0042] In some embodiments, the electronic device 130 may provide reference media content to the first model to generate descriptive text, the descriptive text indicating at least the screen information of the reference media content.

[0043] In some scenarios, the electronic device 130 can input a reference image as shown in Figure 3A into the first model to generate descriptive text corresponding to the image in Figure 3A. The descriptive text may include: image information, color information, resolution information, environmental information, etc. As an example, the descriptive text could be: "a man," "short hair," "black hair," "looking at the audience," "hoodie," "white clothes," "4K" (resolution), and "high contrast," etc.

[0044] In frame 230, electronic device 130 generates target media content based on descriptive text, screen structure, and image data of virtual objects.

[0045] As an example, as shown in Figures 3A and 3B, electronic device 130 can generate target image content (i.e. target media content) as shown in Figure 3B based on the descriptive text and screen structure and image data of virtual objects in Figure 3A.

[0046] In some embodiments, the process of generating target media content may further include: a second model generating intermediate media content based on the image data of the virtual object, the descriptive text of the reference media content, and the screen structure. As an example, the intermediate media content can also be understood as the initial media content. Further, the second model may adjust the intermediate media content based on the image data of the virtual object to generate the target media content.

[0047] As an example, electronic device 130 can provide the second model with descriptive text and screen structure of an image as shown in Figure 3A for the second model to learn. After the second model has completed learning, intermediate media content can be generated based on the image data of the virtual object, the descriptive text of the reference media content, and the screen structure.

[0048] Furthermore, the second model can adjust the image of the characters in the intermediate media content based on the image data of the virtual objects associated with the target user, such as the face shape, eye size, lip thickness, and position of facial features, to obtain the target media content.

[0049] In some embodiments, the second model can also generate intermediate media content based on prompts indicated by the generation request input by user 140 and style description information associated with the reference media content. As an example, the second model can be implemented as a machine learning model capable of generating images. This invention is not intended to limit the specific content or training process of the second model.

[0050] In some scenarios, the user-input prompt could be something like, "The person is wearing a white shirt, and the background is dark." The second model can refer to these prompts when generating the target media content to produce the image content (i.e., the target media content) shown in Figure 3C.

[0051] In some embodiments, the process of generating target media content may further include: a third model redrawing the target area of ​​intermediate media content based on image data to generate the target media content. As an example, the third model may be a machine learning model capable of redrawing images. In some embodiments, this disclosure may also use a second model to redraw the target area of ​​intermediate media content. This invention is not intended to limit the specific content and training process of the third model.

[0052] In some embodiments, the process of generating media content may further include: in response to receiving a regeneration request, the electronic device 130 generates additional media content based on descriptive text, screen structure, and image data of the virtual object. As an example, upon receiving a regeneration request, the electronic device 130 may generate additional media content to make the generated target media content more similar to the virtual object.

[0053] In some scenarios, after generating target media content, electronic device 130 can also perform secondary repairs on the media content to improve the similarity between the target media content and the reference content.

[0054] Specifically, the electronic device 130 can determine the similarity between a virtual object and a target portion of the target media content corresponding to the virtual object. Furthermore, the electronic device 130 can further adjust the target portion based on the virtual object's image data to make the target portion more similar to the virtual object. It should be noted that the electronic device 130 can adjust the target portion based on facial features in the image data, or it can adjust the target portion based on body features in the image data.

[0055] In other scenarios, the electronic device 130 can also process the tone and image quality of the target media content. Processing the tone and image quality of the target media content can make the generated target media content more aesthetically pleasing.

[0056] In some embodiments, the target media content includes visual content corresponding to the virtual object. As an example, the target media content may include clothing, appearance (e.g., shape, skin tone, etc.), background content, and / or accessories from the visual content of the virtual object.

[0057] Based on the process described above, embodiments of this disclosure are able to generate media content associated with user-provided reference media content and user-created virtual objects. Therefore, embodiments of this disclosure can combine user-associated virtual objects and reference media content to generate target media content that better matches user needs.

[0058] Example devices and equipment

[0059] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 4 shows a schematic structural block diagram of an example apparatus 400 for generating media content according to certain embodiments of this disclosure. Apparatus 400 may be implemented as or included in an electronic device. The various modules / components in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0060] As shown in Figure 4, the device 400 includes a receiving module 410 configured to receive a generation request, the generation request indicating reference media content and a virtual object associated with a target user, the virtual object being created based on the target user's configuration operations; a parsing module 420 configured to generate descriptive text of the reference media content and determine the screen structure of the reference media content; and a first generation module 430 configured to generate target media content based on the descriptive text, screen structure, and image data of the virtual object.

[0061] In some embodiments, the parsing module 420 is further configured to: determine edge information of the reference media content; and determine the screen structure of the reference media content based on the edge information.

[0062] In some embodiments, the parsing module 420 is further configured to provide reference media content to the first model to generate descriptive text, the descriptive text indicating at least the screen information of the reference media content.

[0063] In some embodiments, the first generation module 430 is further configured to: provide image data, descriptive text, and screen structure to the second model to generate intermediate media content; and adjust the intermediate media content based on the image data of the virtual object to generate target media content.

[0064] In some embodiments, the intermediate media content is also generated based on at least one of the following: a prompt indicated by the generation request; or style description information associated with the reference media content.

[0065] In some embodiments, the first generation module 430 is further configured to: redraw the target area of ​​the intermediate media content based on the image data to generate the target media content.

[0066] In some embodiments, the apparatus 400 further includes a second generation module configured to generate additional media content based on descriptive text, screen structure, and image data of virtual objects in response to receiving a regeneration request.

[0067] Figure 5 illustrates a block diagram of a computing device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the computing device 500 shown in Figure 5 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The computing device 500 shown in Figure 5 can be used to implement the electronic device of Figure 1.

[0068] As shown in Figure 5, the computing device 500 is in the form of a general-purpose computing device. Components of the computing device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 500.

[0069] Computing device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to computing device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within computing device 500.

[0070] The computing device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0071] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the components of the computing device 500 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0072] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, remote control, etc. Output device 560 can be one or more output devices, such as a monitor, projector, television, etc. Computing device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices, such as storage devices, can communicate with one or more devices that enable user interaction with computing device 500, or with any device (e.g., network card, modem, etc.) that enables computing device 500 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interfaces (not shown).

[0073] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0074] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0075] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0076] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0077] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0078] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for generating media content, comprising: receiving a generation request, the generation request indicating a reference media content and a virtual object associated with a target user, the virtual object being created based on a configuration operation of the target user; generating a description text of the reference media content and determining a picture structure of the reference media content; and generating a target media content based on the description text, the picture structure and image data of the virtual object.

2. The method of claim 1, wherein determining the picture structure of the reference media content comprises: determining edge information of the reference media content; and determining the picture structure of the reference media content based on the edge information.

3. The method of claim 1, wherein generating the description text of the reference media content comprises: providing the reference media content to a first model to generate the description text, the description text indicating at least picture information of the reference media content.

4. The method of claim 1, wherein generating a target media content based on the description text, the picture structure and image data of the virtual object comprises: providing the image data, the description text and the picture structure to a second model to generate an intermediate media content; and adjusting the intermediate media content based on the image data of the virtual object to generate the target media content.

5. The method of claim 4, wherein the intermediate media content is further generated based on at least one of: a hint item indicated by the generation request; style description information associated with the reference media content.

6. The method of claim 4, wherein adjusting the intermediate media content based on the image data of the virtual object to generate the target media content comprises: redrawing a target area of the intermediate media content based on the image data to generate the target media content.

7. The method of claim 1, further comprising: in response to receiving a re-generation request, generating an additional media content based on the description text, the picture structure and the image data of the virtual object.

8. An apparatus for generating media content, comprising: a receiving module configured to receive a generation request, the generation request indicating a reference media content and a virtual object associated with a target user, the virtual object being created based on a configuration operation of the target user; a parsing module configured to generate a description text of the reference media content and determine a picture structure of the reference media content; and a first generating module configured to generate a target media content based on the description text, the picture structure and image data of the virtual object.

9. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method according to any one of claims 1 to 7. ​ ​ ​ ​ ​ 10. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 7.