Method and apparatus for generating media content, and device and storage medium
By presenting interfaces associated with virtual objects on media platforms, acquiring user requests, and using machine learning models to generate multimedia content, the problem of traditional platforms failing to meet user needs is solved, achieving high-quality generation of media content associated with virtual objects and improving the user interaction experience.
Patent Information
- Application Number
- PCT/CN2024/107986
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2026-01-29
AI Technical Summary
Traditional media platforms are unable to meet users' needs for generating media content associated with virtual objects, resulting in insufficient interactive experience.
By presenting a target interface associated with a virtual object, a media generation request is obtained, and target media content including images and text is generated. A machine learning model is used to analyze the image and text information input by the user to generate multimedia content associated with the virtual object.
It improves the interactive experience of users generating media content, meets users' needs for media content associated with virtual objects, and generates high-quality multimedia content.
Smart Images

Figure CN2024107986_29012026_PF_FP_ABST
Abstract
Description
A method, apparatus, device, and storage medium for generating media content. Technical Field
[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to a method, apparatus, device, and computer-readable storage medium for generating media content. Background Technology
[0002] In recent years, with the development of the internet, more and more people are engaging in online activities on online platforms. For example, people use media platforms to view various media works and create virtual avatars. People expect to generate media content using specific virtual avatars.
[0003] Summary of the Invention
[0004] In a first aspect of this disclosure, a method for generating media content is provided, comprising: presenting a target interface associated with a virtual object, the virtual object being created based on user configuration operations; obtaining a media generation request associated with the virtual object via the target interface, the media generation request indicating content description information and / or at least one reference image; and presenting a first portion of target media content generated based on the media generation request, the target media content including a plurality of images and target text associated with the virtual object, the first portion including a first image among the plurality of images and a first text content among the target text.
[0005] In a second aspect of this disclosure, an apparatus for generating media content is provided. The apparatus includes: a first presentation module configured to present a target interface associated with a virtual object, the virtual object being created based on user configuration operations; an acquisition module configured to acquire a media generation request associated with the virtual object via the target interface, the media generation request indicating content description information and / or at least one reference image; and a second presentation module configured to present a first portion of target media content generated based on the media generation request, the target media content including a plurality of images and target text associated with the virtual object, the first portion including a first image among the plurality of images and a first text content among the target text.
[0006] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0010] Figure 1 shows a schematic diagram of an example environment capable of real-time implementation of some embodiments of this disclosure;
[0011] Figure 2 shows a flowchart of an example process for generating content according to some embodiments of the present disclosure;
[0012] Figures 3A to 3D illustrate example interfaces according to some embodiments of the present disclosure;
[0013] Figures 4A to 4C show example interfaces according to further embodiments of the present disclosure;
[0014] Figure 5 illustrates a flowchart of an example process for generating media content according to some embodiments of the present disclosure;
[0015] Figure 6 shows a schematic structural block diagram of an example apparatus for generating media content according to some embodiments of the present disclosure; and
[0016] Figure 7 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure. Detailed Implementation
[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0018] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0019] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0020] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0021] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0022] As briefly mentioned above, with the development of the internet, more and more people are engaging in online activities on online platforms. For example, people use media platforms to view various media works and create virtual avatars. People expect to generate media content using specific virtual avatars. However, traditional media platforms cannot meet users' needs.
[0023] Embodiments of this disclosure propose a scheme for generating media content. According to this scheme, a target interface associated with a virtual object, which is created based on user configuration operations, can be presented; a media generation request associated with the virtual object can be obtained via the target interface, the media generation request indicating content description information and / or at least one reference image; and a first portion of target media content generated based on the media generation request can be presented, the target media content including multiple images and target text associated with the virtual object, the first portion including a first image among the multiple images and a first text content among the target text.
[0024] In this way, embodiments of the present disclosure can generate media content, including images and text, associated with a user-created virtual object, based on the user's media generation request. In this way, embodiments of the present disclosure can meet users' needs for generating media content based on virtual objects, improving the user's interactive experience.
[0025] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.
[0026] Example Environment
[0027] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in Figure 1, the example environment 100 may include an electronic device 110.
[0028] In this example environment 100, electronic device 110 may run an application 120 that supports user interface interaction. Application 120 may be any suitable type of application for user interface interaction, and examples may include, but are not limited to, video applications, social applications, or other suitable applications. User 140 may interact with application 120 via electronic device 110 and / or its attached devices.
[0029] In environment 100 of Figure 1, if application 120 is active, electronic device 110 can use application 120 to present interface 150 for supporting interface interaction.
[0030] In some embodiments, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).
[0031] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for applications 120 that support virtual scenarios in electronic devices 110.
[0032] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections; the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, server 130 and electronic device 110 can achieve signaling interaction through the communication connection between them.
[0033] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0034] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.
[0035] Generate media content
[0036] Figure 2 illustrates an example process 200 for generating media content according to some embodiments of the present disclosure. Process 200 can be implemented at electronic device 110. Process 200 will now be described with reference to Figure 1.
[0037] As shown in Figure 2, in box 205, electronic device 110 can provide story templates. As an example, electronic device 110 can display multiple preset story templates. These multiple preset story templates can be associated with multiple storylines.
[0038] [Revised according to Rule 91, 15.08.2024] An example process for generating media content according to an embodiment of the present disclosure will now be described with reference to the example interfaces shown in Figures 3A to 3D and Figures 4A to 4C.
[0039] Figures 3A to 3D illustrate example interfaces 300A to 300D according to some embodiments of the present disclosure. Interfaces 300A to 300D may be provided, for example, by the electronic device 110 shown in Figure 1.
[0040] In some embodiments, as shown in FIG3A, the electronic device 110 may present a target interface 300A associated with a virtual object. The virtual object is created based on user configuration operations. As an example, as shown in FIG3A, the visual representation of the virtual object may include, for example, virtual object 305.
[0041] In some embodiments, virtual objects can be generated based on image data and descriptive text input by the user. Virtual objects can include their visual representation and capabilities (e.g., dialogue capabilities, data processing capabilities, etc.). For example, a user on a platform can create a corresponding virtual object by configuring the platform. In some scenarios, such virtual objects may also be referred to as digital avatars or virtual clones. For example, the descriptive text input by the user may include character setting information, knowledge information, image information, etc., so that the virtual object can possess certain interactive capabilities and personality traits based on such configuration information. Acquiring the image data of the virtual object may include: the electronic device 110 acquiring at least one image based on the user's configuration operation. The image data of the virtual object is determined based on the acquired at least one image. For example, the configuration operation may include selecting an image from a photo album or taking an image using a camera component. For example, the at least one image may be an image from a photo album or an image taken using a camera component. For example, the visual representation of the virtual object can be generated based on image data and descriptive text input by the user.
[0042] In some embodiments, continuing to refer to FIG3A, electronic device 110 may obtain a media generation request associated with a virtual object via target interface 300A. The media generation request indicates content description information and / or at least one reference image.
[0043] In some embodiments, the content description information is used to describe the plot associated with the target media content. In some embodiments, the content description information may also be used to describe other textual content associated with the target content, such as travelogues, travel guides, plans, diaries, feelings, etc. This disclosure uses plot description as an example only and should not be considered as a limitation on the content description information of this disclosure.
[0044] In some embodiments, at least one reference image may include a second set of images input by the user.
[0045] In some embodiments, continuing to refer to FIG2, at block 210, electronic device 110 can receive an image input by a user. As an example, as shown in FIG2, reference image 215 may include an image input by a user.
[0046] In some embodiments, continuing to refer to FIG3A, the electronic device 110 may provide an image selection control 310 in a target interface 300A. Further, the electronic device 110 may present a plurality of preset images in response to a triggering of the selection control 310. The electronic device 110 may use the image selected by the user from the plurality of preset images as a second set of images input by the user.
[0047] In some embodiments, at least one reference image 215 may include image data of a virtual object and / or a first set of images generated based on the image data of the virtual object. As an example, the first set of images may include multiple image contents generated by expanding upon the image data and / or descriptive information (such as personality traits, knowledge, abilities, etc.) of the virtual object.
[0048] In some embodiments, continuing to refer to FIG3A, the electronic device 110 may provide a plot selection control 315 in the target interface 300A. Further, the electronic device 110 may present multiple plot entries in response to the triggering of the plot selection control 315.
[0049] In some embodiments, as shown in FIG3B, the electronic device 110 may present multiple plot entry points 320 in the target interface 300B. As an example, the multiple plot entry points correspond to different plot templates, and each plot template is associated with a content description information.
[0050] In some embodiments, continuing to refer to FIG2, in box 225, electronic device 110 can determine the story plot.
[0051] In some embodiments, continuing to refer to FIG3B, the electronic device 110 may present a target plot template associated with the target plot entry in the plot settings interface based on the target plot entry among multiple plot entries 320 (e.g., plot entry 320-1 corresponding to plot B).
[0052] In some embodiments, as shown in FIG3C, the electronic device 110 may present a target plot template (e.g., plot B) associated with a target plot entry in the plot setting interface 300C. As an example, as shown in FIG3C, the electronic device 110 may provide content description information (also referred to as first plot description information) corresponding to plot B in the plot setting interface 300C. As an example, the content description information may include a plot overview. The plot overview may include story background, event content, and / or character relationships, etc.
[0053] In some embodiments, continuing to refer to FIG3C, the electronic device 110 may provide an activation experience control 325 in the plot settings interface 300C. Further, the electronic device 110 may, in response to a triggering of the activation experience control 325 (e.g., a swipe operation), present target media content generated based on a media generation request associated with content description information.
[0054] In some embodiments, continuing to refer to FIG3C, the electronic device 110 may provide multiple preset characters associated with the target plot entry in the plot setting interface 300C. As an example, as shown in FIG3C, the electronic device 110 may provide multiple preset characters 330 in the target interface 300C. Further, the electronic device 110 may receive a selection of a target character option among multiple character options (e.g., receive a click operation on the target character option), and may determine the content description information (also known as the first plot description information) associated with the target plot entry based on the description information corresponding to the target character option. As an example, plot B may include multiple characters, such as character A, character B, character C, and character D. Each character has corresponding description information. As an example, the description information corresponding to character A may include character A's gender, identity characteristics, personality traits, experience, etc.
[0055] In some embodiments, continuing to refer to FIG2, in box 230, the first preset model 220 can generate content information (e.g., a text description of the image content) corresponding to the reference image 215. Further, the content information corresponding to the reference image 215 can be used as content description information (also referred to as second plot description information). As an example, the first preset model can be implemented as a machine learning model capable of analyzing and generating text from image content. This invention does not intend to limit the specific content and training process of the first preset model.
[0056] In some embodiments, continuing to refer to FIG2, in block 230, the electronic device 110 can acquire prompts input by the user. As an example, the prompts may include content requirements. Further, the second preset model can update content description information (e.g., first plot description information and / or second plot description information) based on the prompts input by the user. As an example, the second preset model can be implemented as a machine learning model capable of processing input text. This invention is not intended to limit the specific content and training process of the second preset model.
[0057] In some embodiments, continuing to refer to FIG2, in box 240, the second preset model can generate a complete story plot based on the first plot description information and / or at least one reference image 215 (e.g., the second plot description information associated with at least one reference image 215). In some embodiments, the complete story plot is also referred to as the target text.
[0058] In some embodiments, continuing to refer to FIG2, in box 245, a third preset model can supplement images according to the plot. In some embodiments, the third preset model can determine multiple images associated with virtual objects based at least on the plot description text.
[0059] In some embodiments, the third preset model can generate a third set of images associated with the virtual object based on the plot description text and / or at least one reference image. Further, multiple images can be determined based on the third set of images. As an example, the third set of images may include character images generated based on the virtual object's image data. As an example, the third preset model can be implemented as a machine learning model capable of generating images from text. This invention is not intended to limit the specific content and training process of the preset model.
[0060] In some embodiments, continuing to refer to FIG2, the plurality of images may include at least one reference image. In some embodiments, the plurality of images may include a third set of images, and a first set of images and / or a second set of images from at least one reference image. As an example, the plurality of images may include at least one image from a plurality of reference images.
[0061] In some embodiments, continuing to refer to FIG2, the plurality of images further includes at least one image generated based on the image data of the virtual object and the target text (e.g., at least one image in the third set of images). As an example, the at least one image here is a newly generated image other than the first set of images and the second set of images.
[0062] In some embodiments, as shown in FIG2, in box 250, the electronic device 110 can adjust the order of images according to the plot. In some embodiments, the order of multiple images in the target media content corresponds to the content order in the target text, for example, the development order of the plot.
[0063] In some embodiments, continuing to refer to FIG2, in box 255, the obtained target media content includes multiple images and target text associated with the virtual object.
[0064] In some embodiments, as shown in FIG3D, the electronic device 110 may provide target media content 330 generated based on plot B in the target interface 300D. As an example, as shown in FIG3D, a progress control 335 (e.g., 1 / 7) may indicate the update progress of the target media content corresponding to plot B. In the example 1 / 7, 7 may indicate that all content of the target media content corresponding to plot B will be updated in 7 days, and 1 may indicate that the current day is day 1, and the updated part of the target media content 330 corresponding to plot B is the first part.
[0065] In some embodiments, continuing to refer to FIG3D, the electronic device 110 may, in response to a trigger (e.g., a click operation) on the target media content 330 corresponding to plot B, present the updated portion of the target media content 330 in the target interface 300D.
[0066] In some embodiments, continuing to refer to FIG3D, the media generation request corresponding to the target media content 330 is a first request. The electronic device 110 can also acquire at least one additional media content 340, comprising multiple parts, generated based on at least one second request. The at least one additional media content 340 is generated based on a virtual object. In some embodiments, as shown in FIG3D, the electronic device 110, for example, presents a first portion (e.g., 1 / 7) of the first media content within a target time period, and presents a second portion (e.g., 2 / 7) of the second media content corresponding to the target time period.
[0067] Figures 4A to 4C illustrate example interfaces 400A to 400C according to some embodiments of the present disclosure. Interfaces 400A to 400C may be provided, for example, by the electronic device 110 shown in Figure 1.
[0068] In some embodiments, as shown in FIG4A, the electronic device 110 may present a first portion 405 of target media content generated based on a media generation request on a target interface 400A. The first portion 405 includes a first image 410 among a plurality of images and first text content 415 in target text. As an example, the first text content 410 is a portion of the target text, and the first text content 415 may be used to describe the first image 410.
[0069] In some embodiments, continuing to refer to FIG4A, the electronic device 110 may, in response to receiving an update request from a user for the first image 410, present an updated image corresponding to the first text content. As an example, the electronic device 110 may provide an update control 420 on the target interface 400A. Further, the electronic device 110 may, in response to triggering the update control 420 (e.g., a click operation), present an updated image corresponding to the first text content on the target interface 400A.
[0070] In some embodiments, continuing to refer to FIG4A, the target media content may include multiple portions corresponding to multiple time periods. The electronic device 110 may display the first portion 405 within the first time period corresponding to the first portion 405. As an example, the target media content may include seven portions corresponding to a seven-day update cycle.
[0071] In some embodiments, as shown in FIG4B, the electronic device 110 may present the second portion 425 within a second time period corresponding to the second portion 425. The second portion 425 may include a second image 430 among a plurality of images and second text content 435 in target text.
[0072] In some embodiments, the electronic device 110 may provide the complete target media content after multiple time periods. For example, the electronic device 110 may present the complete target media content seven days later. For example, the complete target media content may include comic strips, text and image content, or video content.
[0073] In some embodiments, as shown in FIG4C, the electronic device 110 may, in response to the user opening the target interface 400C, present a first portion 405 of the target media content in the target interface 400C within a first time period.
[0074] Based on the process described above, embodiments of this disclosure can, upon receiving a user's media generation request, generate media content associated with a user-configured virtual object based on the user-input images, prompts, and / or content selected by the user. In this way, this disclosure can provide users with media content associated with virtual objects (e.g., virtual avatars), thereby improving the quality of the generated media content.
[0075] Example process
[0076] Figure 5 shows a flowchart of an example process 500 for generating media content according to some embodiments of the present disclosure. Process 500 can be implemented at electronic device 110. Process 500 will now be described with reference to Figure 1.
[0077] As shown in the figure, in box 510, electronic device 110 presents a target interface associated with a virtual object, which is created based on the user's configuration operations.
[0078] In box 520, electronic device 110 obtains a media generation request associated with a virtual object via a target interface. The media generation request indicates content description information and / or at least one reference image.
[0079] In box 530, electronic device 110 presents a first portion of target media content generated based on a media generation request. The target media content includes multiple images and target text associated with a virtual object. The first portion includes a first image among the multiple images and a first text content among the target text.
[0080] In some embodiments, content description information is used to describe the plot associated with the target media content.
[0081] In some embodiments, at least one reference image includes a first set of images generated based on image data of a virtual object, and multiple images include the first set of images.
[0082] In some embodiments, the target text is generated based on a first set of images, and the order of the first set of images in the target media content corresponds to the content order in the target text.
[0083] In some embodiments, the plurality of images also includes at least one image generated based on the image data of the virtual object and the target text.
[0084] In some embodiments, at least one reference image includes a second set of images input by the user, and multiple images are generated based on the second set of images, image data of a virtual object, and target text.
[0085] In some embodiments, process 500 further includes: presenting multiple plot entry points in a target interface, the multiple plot entry points corresponding to different plot templates; and in response to receiving a selection of a target plot entry point among the multiple plot entry points, determining content description information based on the target plot template associated with the target plot entry point.
[0086] In some embodiments, determining content description information based on a target plot template associated with a target plot entry point includes: in response to the selection of a target plot entry point, presenting a plot settings interface, the plot settings interface displaying multiple preset characters associated with the target plot template; and determining content description information based on the selection of a target character among the multiple preset characters.
[0087] In some embodiments, the target media content is also generated based on prompts entered by the user.
[0088] In some embodiments, the target media content includes multiple parts corresponding to multiple time periods, and presenting the first part of the target media content generated based on the media generation request includes: presenting the first part within the target time period corresponding to the first part.
[0089] In some embodiments, process 500 further includes providing complete target media content after multiple time periods.
[0090] In some embodiments, process 500 further includes: presenting at least one additional media content generated based on a virtual object within a target time period, wherein a second portion of the at least one additional media content corresponds to the target time period.
[0091] In some embodiments, process 500 further includes: in response to receiving an update request from a user for the first image, presenting an updated image corresponding to the first text content.
[0092] In some embodiments, the multiple images include visual content associated with virtual objects.
[0093] In some embodiments, the target media content is generated based on the following process: generating target text based on at least one reference image and / or plot description information; and determining multiple images associated with virtual objects based at least on the target text.
[0094] In some embodiments, determining multiple images associated with a virtual object based at least on target text includes: generating a third set of images associated with the virtual object based on the target text and image data of the virtual object; and determining multiple images based on the third set of images.
[0095] Example devices and equipment
[0096] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 6 shows a schematic structural block diagram of an example apparatus 600 for generating media content according to certain embodiments of this disclosure. Apparatus 600 may be implemented as or included in an electronic device. The various modules / components in apparatus 600 may be implemented by hardware, software, firmware, or any combination thereof.
[0097] As shown in Figure 6, the device 600 includes a first presentation module 610 configured to present a target interface associated with a virtual object, the virtual object being created based on user configuration operations; an acquisition module 620 configured to acquire a media generation request associated with the virtual object via the target interface, the media generation request indicating content description information and / or at least one reference image; and a second presentation module 630 configured to present a first portion of target media content generated based on the media generation request, the target media content including multiple images and target text associated with the virtual object, the first portion including a first image among the multiple images and a first text content among the target text.
[0098] In some embodiments, content description information is used to describe the plot associated with the target media content.
[0099] In some embodiments, at least one reference image includes a first set of images generated based on image data of a virtual object, and multiple images include the first set of images.
[0100] In some embodiments, the target text is generated based on a first set of images, and the order of the first set of images in the target media content corresponds to the content order in the target text.
[0101] In some embodiments, the plurality of images also includes at least one image generated based on the image data of the virtual object and the target text.
[0102] In some embodiments, at least one reference image includes a second set of images input by the user, and multiple images are generated based on the second set of images, image data of a virtual object, and target text.
[0103] In some embodiments, the device 600 further includes a story module, which is configured to: present multiple story entries in a target interface, the multiple story entries corresponding to different story templates; and, in response to receiving a selection of a target story entry among the multiple story entries, determine content description information based on the target story template associated with the target story entry.
[0104] In some embodiments, the story module is further configured to: in response to the selection of a target story entry point, present a story settings interface, the story settings interface displaying multiple preset characters associated with the target story template; and determine content description information based on the selection of a target character among the multiple preset characters.
[0105] In some embodiments, the target media content is also generated based on prompts entered by the user.
[0106] In some embodiments, the target media content includes multiple parts corresponding to multiple time periods, and the second presentation module 630 is further configured to present the first part within the target time period corresponding to the first part.
[0107] In some embodiments, the apparatus 600 further includes a providing module configured to provide complete target media content after a plurality of time periods.
[0108] In some embodiments, the apparatus 600 further includes a third presentation module configured to present at least one additional media content generated based on a virtual object within a target time period, wherein a second portion of the at least one additional media content corresponds to the target time period.
[0109] In some embodiments, the apparatus 600 further includes an update module configured to: in response to receiving an update request from a user for a first image, present an updated image corresponding to the first text content.
[0110] In some embodiments, the multiple images include visual content associated with virtual objects.
[0111] In some embodiments, the target media content is generated based on the following process: generating target text based on at least one reference image and / or content description information; and determining a plurality of images associated with the virtual object based at least on the target text.
[0112] In some embodiments, determining multiple images associated with a virtual object based at least on target text includes: generating a third set of images associated with the virtual object based on the target text and image data of the virtual object; and determining multiple images based on the third set of images.
[0113] Figure 7 illustrates a block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 700 shown in Figure 7 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 700 shown in Figure 7 can be used in electronic devices.
[0114] As shown in Figure 7, the electronic device 700 is in the form of a general-purpose electronic device. Components of the electronic device 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage devices 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processing unit 710 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 720. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 700.
[0115] Electronic device 700 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 730 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 700.
[0116] Electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 720 may include computer program product 725 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0117] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 700 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0118] Input device 750 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 700 can also communicate with one or more external devices (not shown) via communication unit 740 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 700, or with any device that enables electronic device 700 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0119] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0120] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0121] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0122] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0123] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0124] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1.A method of generating media content, comprising: presenting a target interface associated with a virtual object, the virtual object being created based on a configuration operation of a user; obtaining, via the target interface, a media generation request associated with the virtual object, the media generation request indicating content description information and / or at least one reference image; and presenting a first portion of target media content generated based on the media generation request, the target media content including a plurality of images and target text associated with the virtual object, the first portion including a first image of the plurality of images and first text content of the target text. 2.The method of claim 1, wherein the content description information is used to describe a plot associated with the target media content. 3.The method of claim 1, wherein the at least one reference image includes a first set of images generated based on appearance data of the virtual object, and the plurality of images includes the first set of images. 4.The method of claim 3, wherein the target text is generated based on the first set of images, and an order of the first set of images in the target media content corresponds to an order of content in the target text. 5.The method of claim 3, wherein the plurality of images further includes at least one image generated based on the appearance data of the virtual object and the target text. 6.The method of claim 1, wherein the at least one reference image includes a second set of images input by a user, and the plurality of images is generated based on the second set of images, appearance data of the virtual object and the target text. 7.The method of claim 1, further comprising: presenting, in the target interface, a plurality of plot entries, the plurality of plot entries corresponding to different plot templates; and determining the content description information based on a target plot template associated with a target plot entry of the plurality of plot entries in response to receiving a selection of the target plot entry. 8.The method of claim 7, wherein determining the content description information based on a target plot template associated with a target plot entry comprises: presenting, in response to the selection of the target plot entry, a plot setting interface, the plot setting interface displaying a plurality of preset characters associated with the target plot template; and determining the content description information based on a selection of a target character of the plurality of preset characters. 9.The method of claim 1, wherein the target media content is further generated based on a hint item input by a user. 10.The method of claim 1, the target media content including a plurality of portions corresponding to a plurality of time periods, and presenting a first portion of target media content generated based on the media generation request comprises: presenting the first portion within a target time period corresponding to the first portion. 11.The method of claim 10, further comprising: providing the complete target media content after the plurality of time periods. 12.The method of claim 10, further comprising: during the target time period, presenting a second portion of at least one additional media content, the at least one additional media content being generated based on the virtual object, and the second portion of the at least one additional media content corresponding to the target time period. 13.The method of claim 1, further comprising: in response to receiving an update request of a user for the first image, presenting an updated image corresponding to the first text content. 14.The method of claim 1, wherein the plurality of images comprises visual content associated with the virtual object. 15.The method of claim 1, wherein the target media content is generated based on a process comprising: generating the target text based on the at least one reference image and / or the content description information; and determining the plurality of images associated with the virtual object based at least on the target text. 16.The method of claim 15, wherein determining the plurality of images associated with the virtual object based at least on the target text comprises: generating a third set of images associated with the virtual object based on the target text and figurative data of the virtual object; and determining the plurality of images based on the third set of images. 17.An apparatus for generating media content, comprising: a first presentation module configured to present a target interface associated with a virtual object, the virtual object being created based on a configuration operation of a user; an acquisition module configured to acquire, via the target interface, a media generation request associated with a virtual object, the media generation request indicating content description information and / or at least one reference image; and a second presentation module configured to present a first portion of target media content generated based on the media generation request, the target media content comprising a plurality of images and target text associated with the virtual object, the first portion comprising a first image of the plurality of images and first text content of the target text. 18.An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method according to any one of claims 1-16. 19.A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1-16.
Citation Information
Patent Citations
Video generation method and device and storage medium
CN110162667A
Content generation method and device, equipment and storage medium
CN117046094A
Interaction method and device, equipment and storage medium
CN117519528A
Interaction method and device, equipment and storage medium
CN117850937A
Information processing apparatus, method of controlling apparatus, and computer program
EP3032521A1