Method and device for generating media content, equipment and storage medium
By providing an audio content update entry point during the media content generation process, and using the audio content to drive the target object to generate second media content that matches the audio content, the problem of difficulty in controlling the generation process in traditional technologies is solved, thus improving the quality of media content.
Patent Information
- Application Number
- CN202410565322.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-08
- Publication Date
- 2025-11-11
AI Technical Summary
Traditional methods of generating media content make it difficult to control the generation process in detail, which affects the quality of the media content.
By displaying the first media content, responding to preset conditions to provide a target entry point, selecting the target audio content, and using the target audio content to update the target object in the first media content, a second media content matching the audio content is generated.
It improves the quality of the generated media content, making it match the audio content, and enhances the control precision and effectiveness of the generation process.
Smart Images

Figure CN120935387A_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatus, devices, and computer-readable storage media for generating media content. Background Technology
[0002] With the rapid development of smart technology, various forms of electronic devices are greatly enriching people's daily lives. For example, users can generate various types of media content, such as videos and images, using electronic devices. How to efficiently generate media content that meets user needs is a key concern. Summary of the Invention
[0003] In a first aspect of this disclosure, a method for generating media content is provided. The method includes: displaying first media content; providing a target entry associated with the first media content in response to the first media content satisfying preset conditions; determining target audio content in response to selection of the target entry; and displaying second media content matching the target audio content, the second media content being generated by updating a target object in the first media content using the target audio content.
[0004] In a second aspect of this disclosure, an apparatus for generating media content is provided. The apparatus includes: a first display module configured to display first media content; an entry point providing module configured to provide a target entry point associated with the first media content in response to the first media content meeting preset conditions; an audio determination module configured to determine target audio content in response to selection of the target entry point; and a second display module configured to display second media content matching the target audio content, the second media content being generated by updating a target object in the first media content using the target audio content.
[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
[0007] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0009] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;
[0010] Figures 2A to 2E The diagram shows interface schematics according to some embodiments of the present disclosure;
[0011] Figure 3 A flowchart illustrating a process for generating media content according to some embodiments of the present disclosure is shown;
[0012] Figure 4 A schematic structural block diagram of an apparatus for generating media content according to certain embodiments of the present disclosure is shown;
[0013] Figure 5 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation
[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0015] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0016] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0017] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0018] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0019] As discussed above, generative artificial intelligence techniques can be used to generate various types of media content. For example, video content can be generated based on text or images. However, in traditional generation processes, it is difficult to have detailed control over the media content generation process, which affects the quality of the generated media content.
[0020] Embodiments of this disclosure propose a scheme for generating media content. According to this scheme, first media content can be displayed. Further, in response to the first media content satisfying preset conditions, a target entry point associated with the first media content can be provided. Further, in response to the selection of the target entry point, target audio content can be determined. Accordingly, second media content matching the target audio content can be displayed, the second media content being generated by updating a target object in the first media content using the target audio content.
[0021] Embodiments of this disclosure can support updating media content using audio content, thereby generating new media content that matches the audio content. Therefore, embodiments of this disclosure can improve the quality of the generated media content.
[0022] Example Environment
[0023] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, example environment 100 may include electronic device 110.
[0024] In this example environment 100, an application 120 is installed on an electronic device 110. A user 140 can interact with the application 120 via the electronic device 110 and / or its attached devices. The application 120 can be an application that generates media content, or any other suitable application.
[0025] exist Figure 1 In environment 100, if application 120 is active, application 120 can provide a presentation interface 150 for user 140. User 140 can perform operations to generate media content based on interface 150.
[0026] In some embodiments, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).
[0027] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for applications 120 that support content presentation in electronic devices 110.
[0028] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus, and Wi-Fi connections; the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, server 130 and electronic device 110 can achieve signaling interaction through the communication connection between them.
[0029] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0030] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.
[0031] Example Interaction
[0032] The following will be referenced Figures 2A to 2E To describe an example interaction process according to an embodiment of this disclosure. Figures 2A to 2E Example interfaces 200A to 200E according to some embodiments of the present disclosure are shown. Interfaces 200A to 200E may be, for example, provided by... Figure 1 The electronic device 110 shown is provided.
[0033] like Figure 2A As shown, electronic device 110 can present interface 200A based on user requests to generate media content. For example... Figure 2A As shown, in interface 200A, electronic device 110 can provide media generation page 205, which may include generation component 210 for receiving generation information.
[0034] As an example, in the scenario of generating video content, such a generation component can be used to receive input text (e.g., prompts), input images, video control parameters, and other information to generate corresponding video content.
[0035] Accordingly, the electronic device 110 can display media content 215 (also referred to as first media content) generated based on the received information on the interface 200A. In some embodiments, the first media content 215 may be appropriate media content with image information generated based on generation parameters, such as video content, image content, etc.
[0036] In addition to the generated media content, in some embodiments, the first media content 215 may also include, for example, media content uploaded by the user or media content specified by means such as a search.
[0037] Furthermore, the electronic device 110 can determine whether the first media content 215 meets preset conditions. Such preset conditions can be associated with using audio content to drive the images in the first media content 215.
[0038] In some embodiments, the electronic device 110 and / or other suitable device may determine whether the first media content 215 meets preset conditions by detecting whether the first media content 215 includes a target object. As an example, such a target object may include a facial object.
[0039] For example, if it is determined that the first media content 215 does not include a facial object, the electronic device 110 can determine that the first media content 215 does not meet the preset conditions.
[0040] In some embodiments, the electronic device 110 may also determine whether the first media content 215 meets preset conditions based on the screen information of the detected target object in the first media content 215. For example, the target conditions may be related to one or more of the following judgments: whether the target object is fully displayed in the first media content 215, whether the display duration of the target object in the first media content 215 is greater than a threshold, whether the display angle of the target object in the first media content 215 meets the conditions, etc.
[0041] Furthermore, if the first media content 215 is determined to meet preset conditions, the electronic device 110 can provide an entry 220 associated with the first media content 215.
[0042] Conversely, if the first media content 215 is determined to meet preset conditions, the electronic device 110 may, for example, disable or not display the entry point 220. Alternatively, the electronic device 110 may also provide a prompt indicating that the screen using audio content to drive the first media content is not supported.
[0043] Furthermore, upon receiving a selection for entry 220, the electronic device 110 can display, as follows: Figure 2B The interface shown is 200B. (As shown in the image) Figure 2B As shown, the electronic device 110 can provide an audio input component 225 for determining audio content.
[0044] In some embodiments, the audio input component 225 can be configured to determine the target audio content for driving the first media content 215 based on received input information. Specifically, as Figure 2B As shown, the audio input component 225 can support two audio input modes, namely text-to-speech mode 230 and local dubbing upload mode 235.
[0045] like Figure 2B As shown, in text-to-speech mode 230, audio input component 225 can provide text input control 240 to receive text content input by the user. Furthermore, audio input component 225 may also include one or more parameter configuration controls, such as parameter configuration control 245 and parameter configuration control 250.
[0046] Such parameter configuration controls 245 and 250 can be used to set audio parameters for generating audio content. For example, parameter configuration control 245 can be used to set speech rate, and parameter configuration control 250 can be used to set timbre parameters. As an example, the reference matching control 250 can also provide a preview of the audio content with the corresponding timbre parameters to help users perceive the corresponding timbre effect.
[0047] Furthermore, the electronic device 110 can utilize an appropriate speech generation model to generate audio content (also referred to as first audio content) corresponding to the received text content. As an example, such first audio content can be generated based on preset audio parameters, or it can be generated based on audio parameters input by the user via a parameter configuration control.
[0048] Continue to refer to Figure 2C In the local dubbing upload mode 235, the audio input component 225 can provide the upload module 250 to receive the audio content uploaded by the user (also known as the second audio content).
[0049] In some embodiments, the electronic device 110 may directly use the first audio content or the second audio content determined above as the target audio content for driving the first media content 215. For example, when the first media content 215 is image content, the electronic device 110 may trigger the use of the first audio content or the second audio content to drive the target object (e.g., a face) in the first media content 215, thereby generating video content of a corresponding length.
[0050] In some other embodiments, the electronic device 110 may also compare a first duration of the first audio content or the second audio content (collectively referred to as reference audio content) with a second duration associated with the first media content 215. In some embodiments, such a second duration may be the total duration of the second media content 215; or, the second duration may be the corresponding duration of the target object in the second media content 215, for example, the duration of facial appearance.
[0051] In some embodiments, where the first duration is longer than the second duration, the electronic device 110 may, for example, guide the user to edit the reference audio content. For example, the electronic device 110 may provide, for example... Figure 2D The editing window shown is 200D.
[0052] In some embodiments, the electronic device 110 may determine a target segment of the reference audio content 265 via an editing window 220 such that the target segment corresponds to a second duration. For example, if the second duration is 3 seconds, the user may determine a 3-second audio segment in the reference audio content 265 as the target audio content by dragging the time window.
[0053] The above description uses text-to-speech and audio upload as examples to illustrate the process of acquiring target audio content. In some embodiments, such target audio content can also be acquired through other appropriate methods, such as real-time user recording.
[0054] Furthermore, upon receiving, for example, a selection of button 255, the electronic device 110 can generate matching second media content based on the determined target audio content. Specifically, such as Figure 2E As shown, taking text-to-speech mode 230 as an example, electronic device 110 can drive the target object (e.g., face) in first media content 215 to generate second media content 290 based on the target audio content generated by the received text content 270.
[0055] In some embodiments, such as Figure 2E As shown, the electronic device 110 can also display time information (e.g., 3 seconds) corresponding to the text content 270. This time information can indicate the duration of the audio content generated based on the text content 270. It should be understood that such duration can also vary according to speech rate parameters or other audio parameters.
[0056] It should be understood that the electronic device 110 can generate the second media content 290 locally or using any suitable model deployed on a remote device, and this disclosure is not intended to limit the specific generation process of the second media content 290.
[0057] like Figure 2E As shown, the electronic device 110 can display the generated second media 290 on the interface 200E. This second media content 290 can have visuals corresponding to the defined target audio content. For example, the characters in the second media content 290 can exhibit facial movements corresponding to the target audio content. Such facial movements can be implemented, for example, using appropriate facial driving technology.
[0058] In some embodiments, the electronic device 110 may display target elements associated with the second media content 290 to indicate that the second media content 290 was generated using target audio content. As an example, such as Figure 2E As shown, the text “lip-syncing” can be displayed in the upper left corner of the second media content 290 to indicate that it was generated by driving the facial object in the first media content 215 using the target audio content.
[0059] Furthermore, such as Figure 2EAs shown, the electronic device 110 can also provide an entry point 275 associated with the second media content 290. Upon receiving a selection from this entry point 275, the electronic device 110 can switch the interface 200E to display the corresponding first media content 215.
[0060] Additionally, such as Figure 2E As shown, the electronic device 110 may also provide an entry point 280 associated with the second media content 290. Upon receiving a preset operation (e.g., a click or hover) from the entry point 280, the electronic device 110 may display a window 295.
[0061] like Figure 2E As shown, window 295 can be used to play target audio content corresponding to the second media content 290. Alternatively or additionally, window 295 can also be used to display text corresponding to the target audio content, which may, for example, be the same as text content 270.
[0062] Additionally, such as Figure 2E As shown, the electronic device 110 can also provide an entry point 285 associated with the second media content 290. Upon receiving an selection at this entry point 285, the electronic device 110 can provide an audio input component 225 to obtain new target audio content.
[0063] Based on the process described above, embodiments of this disclosure can support secondary generation of the generated media content, enabling the new media content to match the corresponding audio content. Therefore, embodiments of this disclosure can improve the quality of the generated media content.
[0064] Example process
[0065] Figure 3 A flowchart of a process 300 for generating media content according to some embodiments of the present disclosure is shown. Process 300 can be implemented at electronic device 110. Reference is made below. Figure 1 Describe the process 300.
[0066] In frame 310, electronic device 110 displays first media content.
[0067] In frame 320, in response to the first media content meeting preset conditions, electronic device 110 provides a target entry associated with the first media content.
[0068] In box 330, in response to the selection of a target entry, electronic device 110 determines the target audio content.
[0069] In box 340, electronic device 110 displays second media content that matches the target audio content, the second media content being generated by updating the target object in the first media content using the target audio content.
[0070] In some embodiments, process 300 further includes: determining whether the first media content includes a target object; and in response to the first media content including a target object and the screen information of the target object satisfying the target constraint, determining that the first media content satisfies a preset condition.
[0071] In some embodiments, the second media content is generated by using the target audio content to drive a target object in the first media content, such that the target object displays motion that matches the target audio content.
[0072] In some embodiments, determining target audio content in response to the selection of a target entry point includes: displaying an audio input component in response to the selection of a target entry point; determining reference audio content based on input information from the audio input component; and determining target audio content based on the reference audio content.
[0073] In some embodiments, determining the target audio content based on input information from the audio input component includes: obtaining text content via a text input control of the audio input component; and obtaining first audio content generated based on the text content as reference audio content.
[0074] In some embodiments, the first audio content is generated based on text content and at least one audio parameter, the at least one audio parameter including preset audio parameters and / or user-input audio parameters.
[0075] In some embodiments, at least one audio parameter includes a timbre parameter and / or a speech rate parameter.
[0076] In some embodiments, process 300 further includes: displaying time information corresponding to text content, the time information indicating the duration of first audio content generated based on the text content.
[0077] In some embodiments, determining reference audio content based on input information from the audio input component includes: acquiring second audio content uploaded via the audio input component as reference audio content.
[0078] In some embodiments, determining target audio content based on reference audio content includes: displaying an editing window in response to a first duration of reference audio content being greater than a second duration associated with first media content; and determining a target segment of the reference audio content via the editing window as the target audio content, the target segment having the second duration.
[0079] In some embodiments, process 300 further includes: displaying a target element associated with the second media content, the target element indicating that the second media content was generated using target audio content.
[0080] In some embodiments, process 300 further includes: providing a first entry point associated with the second media content; and switching to display the first media content in response to selection of the first entry point.
[0081] In some embodiments, process 300 further includes: providing a second entry point associated with the second media content; and, in response to a preset operation on the second entry point, playing target audio content and / or displaying text corresponding to the target audio content.
[0082] In some embodiments, process 300 further includes: in response to the first media content not meeting preset conditions, displaying a prompt message, the prompt message indicating that the screen using audio content to drive the first media content is not supported.
[0083] Example devices and equipment
[0084] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 4 A schematic structural block diagram of an apparatus 400 for generating media content according to certain embodiments of the present disclosure is shown. The apparatus 400 may be implemented as or included in the electronic device 110 discussed above. The various modules / components in the apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0085] like Figure 4 As shown, the device 400 includes a first display module 410 configured to display first media content; an entry providing module 420 configured to provide a target entry associated with the first media content in response to the first media content meeting preset conditions; an audio determination module 430 configured to determine target audio content in response to the selection of the target entry; and a second display module 440 configured to display second media content matching the target audio content, the second media content being generated by updating a target object in the first media content using the target audio content.
[0086] In some embodiments, the apparatus 400 further includes an object detection module configured to: determine whether the first media content includes a target object; and determine that the first media content meets a preset condition in response to the first media content including a target object and the screen information of the target object satisfying the target constraint.
[0087] In some embodiments, the second media content is generated by using the target audio content to drive a target object in the first media content, such that the target object displays motion that matches the target audio content.
[0088] In some embodiments, the audio determination module 430 is further configured to: display an audio input component in response to the selection of a target entry; determine reference audio content based on input information from the audio input component; and determine target audio content based on the reference audio content.
[0089] In some embodiments, the audio determination module 430 is further configured to: acquire text content via a text input control of the audio input component; and acquire first audio content generated based on the text content as reference audio content.
[0090] In some embodiments, the first audio content is generated based on text content and at least one audio parameter, the at least one audio parameter including preset audio parameters and / or user-input audio parameters.
[0091] In some embodiments, at least one audio parameter includes a timbre parameter and / or a speech rate parameter.
[0092] In some embodiments, the first display module 410 is further configured to display time information corresponding to the text content, the time information indicating the duration of the first audio content generated based on the text content.
[0093] In some embodiments, the audio determination module 430 is further configured to: acquire second audio content uploaded via the audio input component as reference audio content.
[0094] In some embodiments, the audio determination module 430 is further configured to: display an editing window in response to a first duration of reference audio content being greater than a second duration associated with the first media content; and determine a target segment of the reference audio content via the editing window as the target audio content, the target segment having the second duration.
[0095] In some embodiments, the second display module 440 is further configured to display a target element associated with the second media content, the target element indicating that the second media content was generated using target audio content.
[0096] In some embodiments, the second display module 440 is further configured to: provide a first entry point associated with the second media content; and switch to displaying the first media content in response to the selection of the first entry point.
[0097] In some embodiments, the second display module 440 is further configured to: provide a second entry point associated with the second media content; and, in response to a preset operation on the second entry point, play target audio content and / or display text corresponding to the target audio content.
[0098] In some embodiments, the first display module 410 is further configured to: display a prompt message in response to the first media content not meeting a preset condition, the prompt message indicating that the screen does not support updating the first media content using audio content.
[0099] The units included in device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units may be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 400 may be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.
[0100] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 5 The electronic device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown can be used to achieve Figure 1 The electronic device 110 shown.
[0101] like Figure 5 As shown, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.
[0102] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within electronic device 500.
[0103] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0104] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0105] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0106] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0107] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0108] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0109] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0110] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0111] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for generating media content, comprising: Display primary media content; In response to the first media content meeting preset conditions, a target entry point associated with the first media content is provided; In response to the selection of the target entry point, the target audio content is determined; as well as Display second media content that matches the target audio content, the second media content being generated by updating the target object in the first media content using the target audio content.
2. The method according to claim 1, further comprising: Determine whether the target object is included in the first media content; as well as In response to the first media content including the target object and the screen information of the target object satisfying the target constraint, it is determined that the first media content satisfies the preset condition.
3. The method of claim 2, wherein the second media content is generated through the following process: The target audio content is used to drive the target object in the first media content, so that the target object displays a motion that matches the target audio content.
4. The method of claim 1, wherein determining the target audio content in response to selecting the target entry point comprises: In response to the selection of the target entry, an audio input component is displayed; Based on the input information from the audio input component, the reference audio content is determined; as well as The target audio content is determined based on the reference audio content.
5. The method according to claim 4, wherein determining the reference audio content based on the input information of the audio input component includes: Text content is obtained via the text input control of the audio input component; Obtain first audio content generated based on the text content, as the reference audio content.
6. The method according to claim 5, wherein the first audio content is generated based on the text content and at least one audio parameter, wherein the at least one audio parameter includes preset audio parameters and / or user-input audio parameters, wherein the at least one audio parameter includes timbre parameters and / or speech rate parameters.
7. The method according to claim 5, further comprising: Display time information corresponding to the text content, the time information indicating the duration of the first audio content generated based on the text content.
8. The method of claim 4, wherein determining the reference audio content based on the input information of the audio input component comprises: Obtain the second audio content uploaded via the audio input component as the reference audio content.
9. The method of claim 4, wherein determining the target audio content based on the reference audio content comprises: In response to a first duration of the reference audio content being greater than a second duration associated with the first media content, an editing window is displayed; as well as The target segment of the reference audio content is determined via the editing window as the target audio content, and the target segment has the second duration.
10. The method according to claim 1, further comprising: Display a target element associated with the second media content, the target element indicating that the second media content was generated using the target audio content.
11. The method according to claim 1, further comprising: Provide a first entry point associated with the second media content; as well as In response to the selection of the first entry point, the system switches to displaying the first media content.
12. The method according to claim 1, further comprising: Provide a second entry point associated with the second media content; as well as In response to a preset operation on the second entry point, the target audio content is played and / or text corresponding to the target audio content is displayed.
13. The method according to claim 1, further comprising: In response to the first media content not meeting the preset conditions, a prompt message is displayed, which indicates that the use of audio content to drive the picture of the first media content is not supported.
14. An apparatus for generating media content, comprising: The first display module is configured to display the first media content; The entry point providing module is configured to provide a target entry point associated with the first media content in response to the first media content meeting preset conditions; The audio determination module is configured to determine the target audio content in response to the selection of the target entry point; as well as The second display module is configured to display second media content that matches the target audio content, the second media content being generated by updating the target object in the first media content using the target audio content.
15. An electronic device comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 13.
16. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method according to any one of claims 1 to 13.