Media content generation method and apparatus, and device and storage medium

By acquiring reference media content and receiving media content generation requests, and instructing the selection of virtual objects and visual description information to generate target media content, the shortcomings of existing technologies in generating high-quality and convenient media content are solved, and the convenience and richness of media content generation are realized.

WO2025245783A1PCT designated stage Publication Date: 2025-12-04BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/096337
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-30
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

In existing technologies, it is difficult to generate high-quality and convenient media content during the media content generation process, especially in terms of extended image stylization.

Method used

By acquiring reference media content, receiving media content generation requests, instructing the selection of virtual objects, and creating target media content based on user configuration operations, the target media content is generated by combining visual description information.

Benefits of technology

It expands the ways in which media content is generated, improves the convenience and quality of media content generation, and enhances the richness of media content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024096337_04122025_PF_FP_ABST
    Figure CN2024096337_04122025_PF_FP_ABST
Patent Text Reader

Abstract

According to the embodiments of the present disclosure, provided are a media content generation method and apparatus, and a device and a storage medium. The method comprises: acquiring reference media content by means of an editing interface; receiving a media content generation request, wherein the media content generation request indicates the selection of a virtual object corresponding to a user, and the virtual object is created on the basis of a configuration operation of the user; and providing target media content, wherein the target media content is generated on the basis of the reference media content and visual description information that is associated with the virtual object.
Need to check novelty before this filing date? Find Prior Art

Description

Media content generation methods, apparatus, equipment and storage media Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and more particularly to methods, apparatuses, devices and computer-readable storage media for generating media content. Background Technology

[0002] In recent years, with the development of computer technology, media content has become an important medium for sharing or obtaining information. During the editing process of media content, people expect to obtain higher quality media content.

[0003] Summary of the Invention

[0004] In a first aspect of this disclosure, a method for generating media content is provided. The method includes: acquiring reference media content via an editing interface; receiving a media content generation request, the media content generation request indicating the selection of a virtual object corresponding to a user, the virtual object being created based on the user's configuration operations; and providing target media content, the target media content being generated based on the reference media content and visual descriptive information associated with the virtual object.

[0005] In a second aspect of this disclosure, a media content generation apparatus is provided. The apparatus includes an acquisition module configured to acquire reference media content via an editing interface; a receiving module configured to receive a media content generation request, the media content generation request indicating the selection of a virtual object corresponding to a user, the virtual object being created based on the user's configuration operations; and a providing module configured to provide target media content generated based on the reference media content and visual descriptive information associated with the virtual object.

[0006] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of the first aspect.

[0008] It should be understood that the content described in this summary section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0011] Figures 2A to 2F illustrate schematic diagrams of exemplary interfaces for interaction according to some embodiments of the present disclosure;

[0012] Figure 3 illustrates a flowchart of a media content generation process according to some embodiments of the present disclosure;

[0013] Figure 4 shows a block diagram of a media content generation apparatus according to some embodiments of the present disclosure; and

[0014] Figure 5 shows a block diagram of an apparatus capable of implementing several embodiments of the present disclosure. Detailed Implementation

[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0016] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0017] It is understood that the image data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0018] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, relevant users should be informed of the type, scope of use, and usage scenarios of the data involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and authorization from relevant users should be obtained. Among them, relevant users may include any type of rights holder, such as individuals, enterprises, and groups.

[0019] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0020] As mentioned above, during the media content editing process, users expect to obtain higher quality media content. Taking images as an example of media content, users may want to stylize and expand existing images.

[0021] This disclosure provides a media content generation scheme. In the media content generation scheme of this embodiment, reference media content is obtained via an editing interface. Next, a media content generation request is received and a target media content is provided. The media content generation request indicates the selection of a virtual object corresponding to the user, and the virtual object is created based on the user's configuration operations. The target media content is generated based on the reference media content and visual description information associated with the virtual object.

[0022] This approach can expand and enrich the ways in which media content is generated, and further enhance the convenience of media content generation.

[0023] Example Environment

[0024] Referring first to Figure 1, which schematically illustrates an example environment 100 in which exemplary implementations according to this disclosure may be implemented.

[0025] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In this example environment 100, a target application 120 is installed on an electronic device 110. A user 102 can interact with the target application 120 via the electronic device 110 and / or an attached device of the electronic device 110.

[0026] Target application 120 may be an application capable of providing services related to media content generation to user 102, including the creation (e.g., shooting and / or editing), publishing, browsing, and so on of media content. In this document, "media content" can be in various forms, including video, audio, images, image sets, text, and so on.

[0027] For example, electronic device 110 can edit received raw media material through target application 120. In some embodiments, the raw media material may be media material input by user 102. In some other embodiments, the raw media material may also be media material acquired by electronic device 110 based on instructions from user 102 via an acquisition device. In some embodiments, the acquisition device may be configured to be connected to electronic device 110. In some other embodiments, the acquisition device may also be integrated within electronic device 110.

[0028] In environment 100 of Figure 1, if the target application 120 is active, the electronic device 110 can present page 140 of the target application 120 to the user 102. Page 140 can be various types of pages that the target application 120 can provide, such as media content presentation pages, content creation pages, content editing pages, etc.

[0029] In some embodiments, at least some of the functionality of the target application 120 may be implemented based on model 131. Model 131 may, for example, be deployed in server 130. For instance, electronic device 110 communicates with server 130 to provide services to the target application 120. That is, during the operation of the target application 120, the capabilities of one or more models (e.g., model 131) may be invoked.

[0030] In some embodiments, electronic device 110 may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 may also support any type of user-facing interface (such as "wearable" circuitry). Server 130 is any type of computing system / server capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0031] It should be understood that the structure and function of environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0032] Example interactive process of media content generation

[0033] Figures 2A to 2F illustrate schematic diagrams of exemplary interfaces for interaction according to some embodiments of the present disclosure. The interaction process of embodiments of the present disclosure is described below with reference to Figures 2A to 2F. This process can be implemented at least in part at the electronic device 110 shown in Figure 1. For ease of discussion, the process will be described with reference to the environment 100 of Figure 1.

[0034] As shown in Figure 2A, the electronic device 110 can display an editing interface 200A of the target application 120. Through this editing interface 200A, the electronic device 110 can acquire reference media content. For example, the reference media content may include images and videos, etc. An entry point 210 for receiving reference media content can be displayed in the editing interface 200A. The electronic device 110 can receive input of reference media content through the entry point 210.

[0035] Electronic device 110 can receive a media content generation request that indicates the selection of a virtual object corresponding to a user. Such a virtual object may correspond to a target user (e.g., user X) and may be created based on the target user's configuration operations.

[0036] In some embodiments, the virtual object corresponds to a visual representation model created based on a set of media content associated with the user.

[0037] In some scenarios, such virtual objects may also be referred to as virtual avatars or digital clones, which may be driven by machine learning models to have the ability to interact with other users, for example, through text, voice, or video. This disclosure is not intended to limit the specific ways in which virtual objects are constructed.

[0038] Electronic device 110 can provide target media content generated based on reference media content and visual description information associated with virtual objects.

[0039] In some embodiments, the visual description information includes model information corresponding to a visual representation model. In some other embodiments, the visual description information also includes descriptive information about the visual appearance of the virtual object and / or at least one type of media content associated with the virtual object. The following describes in detail, with reference to the accompanying drawings, various ways in which the electronic device 110 receives an instruction to select a virtual object corresponding to a user.

[0040] Electronic device 110 can receive instructions for selecting a virtual object corresponding to the user via an editing interface. As shown in FIG2A, the editing interface 200A presents a control 240 for enabling the virtual object corresponding to the user. If electronic device 110 receives a selection for control 240, electronic device 110 receives an instruction for selecting the virtual object corresponding to the user.

[0041] The editing interface 200A also includes a text input component 220. In this case, the media content generation request received by the electronic device 110 also indicates the descriptive text obtained via the text input component 220, and the generated target media content is also generated based on the descriptive text.

[0042] Furthermore, the editing interface 200A also includes a style input component 230, which can correspond to multiple preset styles. In this case, the media content generation request received by the electronic device 110 also indicates a target style determined via the style input component 230, and the generated target media content corresponds to the specified target style.

[0043] Figure 2B's editing interface 200B illustrates another embodiment where a user-corresponding virtual object can be enabled to generate target media content. As shown in Figure 2B, in the editing interface 200B, a control 250 for enabling a user-corresponding virtual object is presented at the entry point 210 for receiving reference media content. If the electronic device 110 receives a selection for the control 250, the electronic device 110 receives an instruction for selecting a user-corresponding virtual object. In this case, the reference media content obtained from the entry point 210 may correspond to the user themselves corresponding to the virtual object, or it may correspond to an object other than that user.

[0044] The editing interface 200B may also include a text input component 220. In this case, the media content generation request received by the electronic device 110 also indicates the descriptive text obtained via the text input component 220, and the generated target media content is also generated based on the descriptive text.

[0045] Furthermore, the editing interface 200B may also include a style input component 230, which can correspond to multiple preset styles. In this case, the media content generation request received by the electronic device 110 also indicates a target style determined via the style input component 230, and the generated target media content corresponds to the specified target style.

[0046] Figure 2C's editing interface 200C illustrates another embodiment where a virtual object corresponding to the user can be enabled to generate target media content. In this case, the electronic device 110 can present a selection entry 260 for the virtual object and other reference media content. That is, in addition to presenting a selection control 261 corresponding to the user's virtual object, selection controls (e.g., control 262) for other reference media content can also be presented. This other reference media content can be photos from the electronic device 110's album or reference media materials provided by the target application 120. The selected other reference media content can relate to the user corresponding to the virtual object or to other objects besides the user corresponding to the virtual object.

[0047] Similarly, the editing interface 200C may also include a text input component 220. In this case, the media content generation request received by the electronic device 110 also indicates descriptive text obtained via the text input component 220, and the generated target media content is also generated based on the descriptive text.

[0048] The descriptive text received by the text input component 220 of the editing interface 200A, 200B or 200C can be used to indicate the fusion method of the virtual object and the reference media content.

[0049] For example, if the acquired reference media content corresponds to the user corresponding to the virtual object, the description file can integrate the virtual object and the user corresponding to the virtual object into the target media content to be generated. For instance, the description text can indicate the environment, scene, and / or the user's actions, clothing, makeup, hairstyle, etc., presented in the target media content. As shown in Figure 2D, the target media content generated by integrating the virtual object and the user corresponding to the virtual object is presented in area 280 of the viewing interface 200D. This target media content presents elements corresponding to the virtual object, as well as elements from the description file and / or preset styles. The description file received and / or the selected style used to generate this target media content can also be presented in area 290 of the viewing interface 200D.

[0050] For example, if the obtained reference media content corresponds to an object other than the user corresponding to the virtual object, the description file can instruct that the virtual object and the other object be merged into the target media content to be generated. For example, the description text can indicate the environment, scene, and / or the actions, clothing, makeup, hairstyle, etc., of the virtual object and the other object besides the user. As shown in Figure 2E, the target media content generated by merging the virtual object and the other object besides the user is presented in area 281 of the viewing interface 200E. This target media content presents elements corresponding to the virtual object, as well as elements from the description file and / or preset styles. The description file received and / or the selected style used to generate this target media content can also be presented in area 290 of the viewing interface 200E.

[0051] It should be understood that, in some embodiments, the generation of media content may be implemented at least in part by server 130, for example by model 131 deployed on server 130. Electronic device 110 may provide determined media generation control information to server 130 and obtain media content generated based on the generation control information from server 130.

[0052] In some other embodiments, the model can determine the execution order of multiple processes based on the relevance of media generation control information to generate the media content. That is, the model can change the execution order based on the associations between multiple media styles involved in the media generation control information, so that multiple processes corresponding to multiple media styles can be executed in an optimal order.

[0053] The solutions of this application described above in conjunction with different embodiments can expand and enrich the ways of generating media content and further enhance the convenience of media content generation.

[0054] In some other embodiments, the electronic device 110 may present one or more candidate prompts based on text content (e.g., prompt words) received via the text input component 220. These candidate prompts may be text content expanded from the received text content to enrich the text content used for media content generation.

[0055] Electronic device 110 can receive recommendation requests for text content received via text input component 220. In some embodiments, electronic device 110 receives the recommendation request if it receives a trigger on a recommendation control. As shown in the editing interface 200F of FIG2F, if electronic device 110 receives a selection for text optimization control 223, it can determine the recommendation request for the received text content.

[0056] Upon receiving the recommendation request, the electronic device 110 may provide multiple candidate suggestions, for example, based on text content (e.g., prompt words) already received via the text input component 220. As shown in the editing interface 200F of FIG2F, the multiple candidate recommendation words 272 to 275 presented in the candidate suggestion word presentation area 270 may be generated based on text content (e.g., prompt word 271) already received via the text input component 220. If the electronic device 110 receives a selection for one of the multiple candidate recommendation words 272 to 275, the previously received text content can be updated to the selected candidate recommendation word.

[0057] Furthermore, it is also possible that multiple candidate prompts can be generated based on at least one reference material obtained via the material input component 210. Additionally, multiple candidate prompts can be generated together based on a virtual object corresponding to the user and the obtained at least one reference material.

[0058] In some embodiments, multiple candidate prompts can be generated based on the detection results of objects in at least one reference material. The object detection results indicate whether the at least one reference material includes objects of a predetermined type. For example, it may indicate whether the reference material presents subjects such as people, animals, or plants. If the reference material presents subjects such as people, animals, or plants, the multiple candidate prompts may relate to descriptions of how the user's virtual object blends with the subjects presented in the reference material, thereby enriching the media content to be generated.

[0059] It should be understood that multiple candidate prompts can be generated locally by electronic device 110. Alternatively, multiple candidate prompts can also be generated at least partially by server 130, for example, through model 131 deployed on server 130. Optimized recommendations for the text content received via text input component 220 can further improve the granularity of the prompts, thereby resulting in better media content generation.

[0060] Example process

[0061] Figure 3 illustrates a flowchart of a media content generation process 300 according to some embodiments of the present disclosure. Process 300 may be implemented at electronic device 110. It should be understood that process 300 may also be implemented at server 130.

[0062] In frame 310, electronic device 110 obtains reference media content via an editing interface.

[0063] In box 320, electronic device 110 receives a media content generation request that indicates the selection of a virtual object corresponding to a user, the virtual object being created based on the user's configuration operations.

[0064] In box 330, electronic device 110 provides target media content generated based on the reference media content and visual descriptive information associated with the virtual object.

[0065] In some embodiments, the virtual object corresponds to a visual representation model created based on a set of media content associated with the user.

[0066] In some embodiments, the visual description information includes model information corresponding to the visual representation model.

[0067] In some embodiments, the visual description information includes: description information about the visual appearance of the virtual object; and at least one piece of media content associated with the virtual object.

[0068] In some embodiments, the editing interface further includes a text input component, the media content generation request further indicates descriptive text obtained via the text input component, and the target media content is also generated based on the descriptive text.

[0069] In some embodiments, the descriptive text is used to indicate how the virtual object is integrated with the reference media content.

[0070] In some embodiments, the electronic device 110 may also receive a first prompt word input by a user via the text input component; in response to a received recommendation request, present a set of candidate prompt words associated with the first prompt word; and update the first prompt word based on the selection of a second prompt word from the set of candidate prompt words.

[0071] In some embodiments, the electronic device 110 may also receive the recommendation request in response to a triggering of a recommendation control in the text input component.

[0072] In some embodiments, the electronic device 110 may also generate the set of candidate prompts based on the at least one reference material obtained via the material input component.

[0073] In some embodiments, the editing interface further includes a style input component, and the media content generation request further indicates a target style determined via the style input component, the target media content corresponding to the specified target style.

[0074] In some embodiments, the target media content includes elements corresponding to the virtual object.

[0075] Example devices and equipment

[0076] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 4 shows a schematic structural block diagram of an apparatus 400 for media content generation according to some embodiments of this disclosure.

[0077] As shown in Figure 4, the device 400 may include an acquisition module 410 configured to acquire reference media content via an editing interface; a receiving module 420 configured to receive a media content generation request, the media content generation request indicating the selection of a virtual object corresponding to a user, the virtual object being created based on the user's configuration operation; and a providing module 430 configured to provide target media content, the target media content being generated based on the reference media content and visual description information associated with the virtual object.

[0078] In some embodiments, the virtual object corresponds to a visual representation model created based on a set of media content associated with the user.

[0079] In some embodiments, the visual description information includes model information corresponding to the visual representation model.

[0080] In some embodiments, the visual description information includes: description information about the visual appearance of the virtual object; and at least one piece of media content associated with the virtual object.

[0081] In some embodiments, the editing interface further includes a text input component, the media content generation request further indicates descriptive text obtained via the text input component, and the target media content is also generated based on the descriptive text.

[0082] In some embodiments, the descriptive text is used to indicate how the virtual object is integrated with the reference media content.

[0083] In some embodiments, the apparatus 400 may further include a second receiving module configured to receive a first prompt word input by a user via the text input component; a presentation module configured to present a set of candidate prompt words associated with the first prompt word in response to a received recommendation request; and an update module configured to update the first prompt word based on the selection of a second prompt word from the set of candidate prompt words.

[0084] In some embodiments, the electronic device 110 may further include a third receiving module configured to receive the recommendation request in response to a triggering of a recommendation control in the text input component.

[0085] In some embodiments, the electronic device 110 may further include a generation module configured to generate the set of candidate prompts based on at least one reference material obtained via the material input component.

[0086] In some embodiments, the editing interface further includes a style input component, and the media content generation request further indicates a target style determined via the style input component, the target media content corresponding to the specified target style.

[0087] In some embodiments, the target media content includes elements corresponding to the virtual object.

[0088] Figure 5 shows a block diagram of a computing device / server 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the computing device / server 500 shown in Figure 5 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein.

[0089] As shown in Figure 5, the computing device / server 500 is in the form of a general-purpose computing device. Components of the computing device / server 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 560, and one or more output devices 560. The processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device / server 500.

[0090] The computing device / server 500 typically includes multiple computer storage media. Such media can be any available media accessible to the computing device / server 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data (e.g., training data for training) and accessible within the computing device / server 500.

[0091] The computing device / server 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0092] The communication unit 540 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device / server 500 can be implemented as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device / server 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0093] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. The computing device / server 500 can also communicate as needed with one or more external devices (not shown) via communication unit 540. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with the computing device / server 500, or with any device (e.g., network card, modem, etc.) that enables the computing device / server 500 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interfaces (not shown).

[0094] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores one or more computer instructions, wherein one or more computer instructions are executed by a processor to implement the methods described above.

[0095] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0096] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0097] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0098] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0099] Various implementations of this disclosure have been described above. The foregoing description is exemplary and not exhaustive, nor is it limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the implementations disclosed herein.

Claims

1. A method for generating media content, comprising: Access reference media content through the editing interface; Receive a media content generation request, the media content generation request indicating the selection of a virtual object corresponding to the user, the virtual object being created based on the user's configuration operations; as well as Provide target media content, which is generated based on the reference media content and visual description information associated with the virtual object.

2. The method of claim 1, wherein the virtual object corresponds to a visual representation model, the visual representation model being created based on a set of media content associated with the user.

3. The method according to claim 1, wherein the visual description information includes model information corresponding to the visual representation model.

4. The method according to claim 1, wherein the visual description information includes: Descriptive information about the visual appearance of the virtual object; At least one piece of media content associated with the virtual object.

5. The method of claim 1, wherein the editing interface further includes a text input component, the media content generation request further indicates descriptive text obtained via the text input component, and the target media content is further generated based on the descriptive text.

6. The method of claim 5, wherein the descriptive text is used to indicate the fusion method of the virtual object and the reference media content.

7. The method according to claim 5, further comprising: The text input component receives the first prompt word entered by the user; In response to the received recommendation request, a set of candidate suggestions associated with the first suggestion is presented; as well as The first prompt word is updated based on the selection of the second prompt word from the set of candidate prompt words.

8. The method according to claim 1, further comprising: In response to the triggering of the recommendation control in the text input component, the recommendation request is received.

9. The method according to claim 8, further comprising: In response to obtaining at least one reference material via the material input component, the set of candidate prompt words is generated based on the at least one reference material.

10. The method of claim 1, wherein the editing interface further includes a style input component, and the media content generation request further indicates a target style determined via the style input component, the target media content corresponding to the specified target style.

11. The method of claim 1, wherein the target media content includes elements corresponding to the virtual object.

12. A media content generation device, comprising: The acquisition module is configured to acquire reference media content via the editing interface; A receiving module is configured to receive a media content generation request, the media content generation request indicating the selection of a virtual object corresponding to a user, the virtual object being created based on the user's configuration operation; as well as A providing module is configured to provide target media content, which is generated based on the reference media content and visual description information associated with the virtual object.

13. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Media content generation method and device, electronic equipment and storage medium

    CN115098099A

  • Content providing method and device, equipment and storage medium

    CN117390237A

  • Method and device for creating virtual object, equipment and storage medium

    CN118012318A

  • Method and device for providing media content, electronic equipment and storage medium

    CN118092731A

  • KR20220148757A