A page interaction method and device for multimedia creation, electronic equipment and storage medium

By displaying video frames and audio features from multiple perspectives on the media creation page, the target video is generated, solving the problem of inconsistent features of the main object in AI-generated videos and ensuring the professionalism and consistency of the videos.

CN122131949APending Publication Date: 2026-06-02BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
Filing Date
2026-01-30
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing AI-generated videos, the features of core elements such as people, products, and scenes drift, deform, or are lost between frames, affecting the credibility of the content and its dissemination effect.

Method used

This paper provides a multimedia creation page interaction method. By displaying preset images on the media creation page, adding materials of the target subject object, including video frames and sound features from multiple perspectives, the target video is generated, ensuring the consistency of the appearance and features of the subject object.

Benefits of technology

It achieves a high degree of consistency in the appearance and features of the main object across different shots, actions, and scenes, avoiding feature distortion, deviation, and style drift, thus enhancing the professionalism of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122131949A_ABST
    Figure CN122131949A_ABST
Patent Text Reader

Abstract

This disclosure relates to a multimedia creation page interaction method, apparatus, electronic device, and storage medium. The method includes: displaying a preset image containing a target subject object on a media creation page; displaying material of the target subject object on the media creation page in response to an instruction to add material for the target subject object; the material of the target subject object includes a preset video containing at least two video frames and preset sound features; the target subject object in the two video frames has at least different perspectives; generating a target video associated with the preset image in response to a video creation instruction for the preset image and the material of the target subject object; the target video includes the target subject object; the sound of the target subject object in the target video is generated based on preset sound features. By providing video and audio from multiple perspectives for the subject object, embodiments of this application can ensure a high degree of consistency in the appearance and features of the subject object when implementing different shots, actions, and scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of Internet technology, and in particular to a method, apparatus, electronic device, and storage medium for multimedia creation page interaction. Background Technology

[0002] With the deep penetration of AI-generated models in video creation, AI-generated videos have been widely used in e-commerce advertising, short animated dramas, content marketing, and other scenarios. However, the current technology still faces a core bottleneck: the issue of theme consistency. The theme consistency problem refers to the drift, distortion, or loss of features (such as appearance, style, and logical relationships) of core elements such as people, products, and scenes in the generated video between frames, which seriously affects the credibility and dissemination effect of the content. Summary of the Invention

[0003] This disclosure provides a multimedia creation page interaction method, apparatus, electronic device, and storage medium. The technical solution of this disclosure is as follows: According to a first aspect of the present disclosure, a multimedia creation page interaction method is provided, comprising: Display a preset image containing the target subject on the media creation page; In response to the instruction to add material for the target subject, the material for the target subject is displayed on the media creation page; the material for the target subject includes preset video and preset sound features containing at least two video frames; the target subject in the two video frames has at least a different perspective; In response to a video creation instruction for a preset image and a target subject, a target video associated with the preset image is generated; the target video includes the target subject; the sound of the target subject in the target video is generated based on preset sound features.

[0004] In some possible embodiments, the media creation page includes a material addition control; in response to a material addition instruction for a target subject object, the media creation page displays the material for the target subject object, including: In response to a material addition command triggered by a material addition control, a material display area is displayed on the media creation page; the material display area displays materials with multiple primary subject objects; the materials with multiple primary subject objects originate from the subject object library; In response to a confirmation command to add material for a target subject object, an identifier of the target subject object is displayed on the material addition control; displaying the identifier of the target subject object on the material addition control indicates that the material for the target subject object has been selected; the material for the target subject object is at least one of multiple materials for a first subject object; The material for each target subject is a preset video containing at least two video frames and preset sound features.

[0005] In some possible embodiments, the material display area displays materials of multiple first subject objects, multiple first sound features, and an add confirmation control; the multiple first sound features are derived from a sound library; In response to a confirmation command to add material targeting a specific subject, the identifier of the target subject is displayed on the material addition control, including: In response to a first add command for material of a second subject object, material of the selected second subject object is displayed; the material of the second subject object is at least one of a plurality of materials of the first subject objects. In response to a second add instruction for a preset sound feature, the selected preset sound feature is displayed; the preset sound feature is at least one of a plurality of first sound features; In response to an add confirmation command triggered by the add confirmation control, the identifier of the target subject object is displayed on the material add control; the identifier of the target subject object includes the identifier of the second subject object and the identifier of the preset sound characteristics; The target subject's material includes the video frames and preset sound features corresponding to the material of the second subject.

[0006] In some possible embodiments, when the material of the second subject object is a preset video including at least two video frames and a second sound feature, and a selected first sound feature exists among the multiple first sound features, the material of the target subject object includes the video frame corresponding to the material of the second subject object, the first type of sound feature among the second sound features, and the second type of sound feature among the selected first sound features.

[0007] In some possible embodiments, the preset sound features include at least one of pitch, timbre, and dialect.

[0008] In some possible embodiments, the material display area further includes a main object creation control; the method also includes: In response to a subject object creation command triggered by a subject object creation control, a subject creation page is displayed; the subject creation page includes a subject image collection area and subject collection prompts. When the image of the subject being collected is displayed in the subject image collection area, the subject recording instruction is triggered in response to the subject being collected performing activities according to the subject collection prompt information, thereby generating the material of the subject being collected; the material of the subject being collected is in video format. The collected material of the main object is stored in the main object library as material of a first main object.

[0009] In some possible embodiments, when the image of the subject being collected is displayed in the subject image collection area, a subject recording instruction is triggered in response to the subject being collected performing activities according to the subject collection prompt information, thereby generating the material of the collected subject being collected, including: When the image of the subject being collected is displayed in the subject image collection area, the subject recording instruction is triggered in response to the subject being collected performing activities according to the subject collection prompt information, thereby generating the video of the subject being collected. Display voice prompts on the main page; In response to the collected subject object reciting according to the voice reading prompt, the voice recording instruction is triggered to generate the audio of the collected subject object; The collected video and audio of the subject matter constitute the material of the collected subject matter.

[0010] In some possible embodiments, the preset image is an image containing the target subject object; the preset image is the first frame image, the last frame image, or an intermediate keyframe image in the target video; The preset images are two images, each containing the target subject; the two preset images are the first frame and the last frame of the target video, respectively.

[0011] In some possible embodiments, the preset images are a first preset image containing a first target subject and a second preset image containing a second target subject; the first preset image and the second preset image are the first frame and the last frame of the target video, respectively; The materials for the target subject include the materials corresponding to the first target subject and the materials corresponding to the second target subject.

[0012] In some possible embodiments, the main object library is a sub-library of the comprehensive material library; the comprehensive material library also includes at least a scene material library, a prop material library, and a costume material library; The target scenes in the scene material library, the target props in the prop material library, and the target clothing in the clothing material library are used to combine preset images and target subject materials to generate target videos associated with preset images.

[0013] In some possible embodiments, the method further includes: Display video generation prompts in the video prompt area on the media creation page; In response to video creation instructions for preset images and target subjects, generate a target video associated with the preset images, including: The system generates prompts using video, responding to video creation instructions for preset images and target subjects, and generates target videos associated with the preset images. The video generation prompts include video generation prompts, specified shot information, transition information, and video duration.

[0014] According to a second aspect of the present disclosure, a multimedia authoring page interaction device is provided, comprising: The first display module is configured to display a preset image containing the target subject on the media creation page; The second display module is configured to execute a response to an instruction to add material for a target subject object, and display the material of the target subject object on the media creation page; the material of the target subject object includes preset video and preset sound features containing at least two video frames; the target subject object in the two video frames has at least a different perspective; The third display module is configured to execute video creation instructions in response to preset images and target subject materials, and generate a target video associated with the preset images; the target video includes the target subject; the sound of the target subject in the target video is generated based on preset sound features.

[0015] In some possible embodiments, the media creation page includes material addition controls; a second display module is configured to perform: In response to a material addition command triggered by a material addition control, a material display area is displayed on the media creation page; the material display area displays materials with multiple primary subject objects; the materials with multiple primary subject objects originate from the subject object library; In response to a confirmation command to add material for a target subject object, an identifier of the target subject object is displayed on the material addition control; displaying the identifier of the target subject object on the material addition control indicates that the material for the target subject object has been selected; the material for the target subject object is at least one of multiple materials for a first subject object; The material for each target subject is a preset video containing at least two video frames and preset sound features.

[0016] In some possible embodiments, the material display area displays materials of multiple first subject objects, multiple first sound features, and an add confirmation control; the multiple first sound features are derived from a sound library; The second display module is configured to execute: In response to a first add command for material of a second subject object, material of the selected second subject object is displayed; the material of the second subject object is at least one of a plurality of materials of the first subject objects. In response to a second add instruction for a preset sound feature, the selected preset sound feature is displayed; the preset sound feature is at least one of a plurality of first sound features; In response to an add confirmation command triggered by the add confirmation control, the identifier of the target subject object is displayed on the material add control; the identifier of the target subject object includes the identifier of the second subject object and the identifier of the preset sound characteristics; The target subject's material includes the video frames and preset sound features corresponding to the material of the second subject.

[0017] In some possible embodiments, when the material of the second subject object is a preset video including at least two video frames and a second sound feature, and a selected first sound feature exists among the multiple first sound features, the material of the target subject object includes the video frame corresponding to the material of the second subject object, the first type of sound feature among the second sound features, and the second type of sound feature among the selected first sound features.

[0018] In some possible embodiments, the preset sound features include at least one of pitch, timbre, and dialect.

[0019] In some possible embodiments, the material display area further includes a subject object creation control; the apparatus also includes a subject creation module configured to perform: In response to a subject object creation command triggered by a subject object creation control, a subject creation page is displayed; the subject creation page includes a subject image collection area and subject collection prompts. When the image of the subject being collected is displayed in the subject image collection area, the subject recording instruction is triggered in response to the subject being collected performing activities according to the subject collection prompt information, thereby generating the material of the subject being collected; the material of the subject being collected is in video format. The collected material of the main object is stored in the main object library as material of a first main object.

[0020] In some possible embodiments, the subject creation module is configured to perform: When the image of the subject being collected is displayed in the subject image collection area, the subject recording instruction is triggered in response to the subject being collected performing activities according to the subject collection prompt information, thereby generating the video of the subject being collected. Display voice prompts on the main page; In response to the collected subject object reciting according to the voice reading prompt, the voice recording instruction is triggered to generate the audio of the collected subject object; The collected video and audio of the subject matter constitute the material of the collected subject matter.

[0021] In some possible embodiments, the preset image is an image containing the target subject object; the preset image is the first frame image, the last frame image, or an intermediate keyframe image in the target video; The preset images are two images, each containing the target subject; the two preset images are the first frame and the last frame of the target video, respectively.

[0022] In some possible embodiments, the preset images are a first preset image containing a first target subject and a second preset image containing a second target subject; the first preset image and the second preset image are the first frame and the last frame of the target video, respectively; The materials for the target subject include the materials corresponding to the first target subject and the materials corresponding to the second target subject.

[0023] In some possible embodiments, the main object library is a sub-library of the comprehensive material library; the comprehensive material library also includes at least a scene material library, a prop material library, and a costume material library; The target scenes in the scene material library, the target props in the prop material library, and the target clothing in the clothing material library are used to combine preset images and target subject materials to generate target videos associated with preset images.

[0024] In some possible embodiments, the display device further includes a fourth display module configured to perform: Display video generation prompts in the video prompt area on the media creation page; The third display module is configured to execute: The system generates prompts using video, responding to video creation instructions for preset images and target subjects, and generates target videos associated with the preset images. The video generation prompts include video generation prompts, specified shot information, transition information, and video duration.

[0025] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute instructions to implement the method as described in either the first or second aspect above.

[0026] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method of any one of the first or second aspects of the present disclosure.

[0027] According to a fifth aspect of the present disclosure, a computer program product is provided, the computer program product including a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from the readable storage medium and executes the computer program, causing the computer device to perform the method of any one of the first or second aspects of the present disclosure.

[0028] The technical solutions provided by the embodiments of this disclosure bring at least the following beneficial effects: The media creation page displays a preset image containing the target subject object; in response to an instruction to add material for the target subject object, the media creation page displays the material for the target subject object; the material for the target subject object includes a preset video containing at least two video frames and preset sound features; the target subject object in the two video frames has at least different perspectives; in response to a video creation instruction for the preset image and the material for the target subject object, a target video associated with the preset image is generated; the target video includes the target subject object; the sound of the target subject object in the target video is generated based on preset sound features. This embodiment of the application can ensure a high degree of consistency in the appearance and features of the subject object when implementing different shots, actions, and scenes by providing video and audio from multiple perspectives for the subject object, avoiding problems such as deformation, deviation, and style drift of the subject object's features due to insufficient reference, thereby making the video more professional.

[0029] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 This is a schematic diagram illustrating an application environment for a multimedia creation page interaction method according to an exemplary embodiment; Figure 2 This is a flowchart illustrating a multimedia creation page interaction method according to an exemplary embodiment; Figure 3 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 1 ; Figure 4 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 2 ; Figure 5 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 3 ; Figure 6 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 4 ; Figure 7 This is a schematic diagram illustrating a material display area according to an exemplary embodiment; Figure 8 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 5 ; Figure 9 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 6 ; Figure 10 This is a block diagram illustrating a multimedia creation page interaction device according to an exemplary embodiment; Figure 11 This is a block diagram illustrating an electronic device for interactive page creation for multimedia authoring, according to an exemplary embodiment. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar first objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0034] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0035] Please see Figure 1 , Figure 1This is a schematic diagram illustrating an application environment for a multimedia creation page interaction method according to an exemplary embodiment, such as... Figure 1 As shown, the application environment may include server 011 and client 012.

[0036] In some possible embodiments, server 011 may be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The operating system running on the server may include, but is not limited to, Android, iOS, Linux, Windows, Unix, etc.

[0037] In some possible embodiments, the client 012 described above may include, but is not limited to, image-type clients such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices. It may also be software running on the client, such as applications or mini-programs. Optionally, the operating system running on the client may include, but is not limited to, Android, iOS, Linux, Windows, and Unix systems.

[0038] In some possible embodiments, client 012 displays a preset image containing the target subject on the media creation page; in response to an instruction to add material for the target subject, the client displays material for the target subject on the media creation page; the material for the target subject includes a preset video and preset sound features comprising at least two video frames; the target subject in the two video frames has at least different perspectives; in response to a video creation instruction for the preset image and the material for the target subject, a target video associated with the preset image is generated; the target video includes the target subject; the sound of the target subject in the target video is generated based on preset sound features. This embodiment of the application can ensure a high degree of consistency in the appearance and features of the subject by providing video and audio from multiple perspectives for the subject, ensuring that different shots, actions, and scenes are achieved. This avoids problems such as deformation, deviation, and style drift of the subject's features due to insufficient reference, thereby making the video more professional.

[0039] In one exemplary implementation, both the client and server databases can be node devices in the blockchain system, capable of sharing acquired and generated information with other node devices within the blockchain system, thus enabling information sharing among multiple node devices. Multiple node devices in the blockchain system can be configured with the same blockchain, which consists of multiple blocks. Adjacent blocks are related, ensuring that any data tampering in any block can be detected by the next block, thereby preventing data tampering and guaranteeing the security and reliability of the data in the blockchain.

[0040] Figure 2 This is a flowchart illustrating a multimedia creation page interaction method according to an exemplary embodiment. It should be noted that this specification provides the operational steps of the method described in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operational steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many steps and does not represent the only execution order. In actual system or product execution, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment). Specifically, as shown... Figure 2 As shown, this flowchart includes at least the following steps S201-S203: In step S201, a preset image containing the target subject is displayed on the media creation page.

[0041] In this embodiment, the media creation page can be a page provided in an application for creating media content. This target page displays various interactive areas, which serve as entry points for users to input, edit, and view media content. This media content can be the media materials required by the application for generating media content. These media materials can cover single or multiple types of media content. This embodiment does not limit the type of media content, including text, audio, images, video, or other formats.

[0042] In this embodiment, the preset image serves as reference media content for generating the target video. This embodiment does not limit the number of preset images; one or more preset images can be used depending on the actual needs of generating the target video.

[0043] In some possible embodiments, the number of preset images is one, and the preset image contains the target subject object. Optionally, after generating the target video based on the preset image, it can be seen that the target video contains the preset image, that is, the preset image is a frame of the target video. Optionally, the preset image can be the first frame of the target video. Optionally, the preset image can be the last frame of the target video. Optionally, the preset image can be a keyframe image in the middle of the target video.

[0044] In some other possible embodiments, the number of preset images is multiple. Taking two preset images as an example, the preset images are two images that both contain the target subject object, such as a first preset image and a second preset image.

[0045] Optionally, after generating the target video based on the first preset image and the second preset image, it can be seen that the target video contains the first preset image and the second preset image. Optionally, the first preset image and the second preset image can be the first frame image and the last frame image of the target video, respectively. Optionally, the first preset image and the second preset image can be keyframe images in the middle of the target video, respectively.

[0046] In this embodiment, the target subject in the preset image is the main object in the image, which is the most core and attractive object in the image, and can carry the main information and express the purpose. Other objects in the image exist to complement it.

[0047] Optionally, this application embodiment does not limit the number or type of target objects in the preset image. Optionally, the target objects in the preset image can be one or more, and the category of the target objects in the preset image can include at least one of people, animals, items, and scenes.

[0048] In some possible embodiments, when the number of preset images is one, the preset image may include one or more target subject objects. Optionally, the preset image may include one target subject object, person A. Optionally, the preset image may include two target subject objects, person A and person B. Optionally, the preset image may include two target subject objects, person A and animal C. Optionally, the preset image may include two target subject objects, person A and item D.

[0049] In another possible embodiment, when there are multiple preset images, different preset images may include different target objects, and a preset image may include one or more target objects.

[0050] Optionally, the preset images include a first preset image containing a first target subject and a second preset image containing a second target subject. Optionally, the number and type of the first target subject and the number and type of the second target subject are not limited.

[0051] Optionally, the first preset image and the second preset image can be the first frame and the last frame of the target video, respectively. Optionally, the first preset image and the second preset image can be keyframe images in the middle of the target video, respectively. Optionally, one of the first preset image and the second preset image can be the first frame image or the last frame image, and the other preset image can be a keyframe image in the middle.

[0052] Figure 3 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 1 ,like Figure 3 As shown, it includes a media creation page 300, and a content reference area 301 and a video display area 302 located on the media creation page.

[0053] like Figure 3 As shown, the content reference area 301 is used to input the main reference content, that is, to input a preset image containing the target subject. Optionally, you can... Figure 3 The content reference area shown allows input of one or more preset images, and allows definition of the order and positional relationship of the input preset images in the final generated target video (e.g., whether it is the first frame image or the last frame image). The video display area 302 is used to present the final generated target video.

[0054] Figure 4 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 2 ,like Figure 4 As shown, it includes a media creation page 300, and a content reference area 301 and a video display area 302 located on the media creation page.

[0055] like Figure 4 As shown, the content reference area 301 may include a first frame input area and a last frame input area. The first frame input area and the last frame input area are used to input the main reference content, i.e., to input a preset image containing the target subject. Optionally, a preset image may only be input in the first frame input area; in this case, the preset image input in the first frame input area is the first frame image of the final generated target video. Optionally, a preset image may only be input in the last frame input area; in this case, the preset image input in the last frame input area is the last frame image of the final generated target video. The video display area 302 is used to display the final generated target video.

[0056] In step S203, in response to the instruction to add material for the target subject, the material of the target subject is displayed on the media creation page; the material of the target subject includes a preset video and preset sound features containing at least two video frames; the target subject in the two video frames has at least a different perspective.

[0057] In this embodiment, in response to an instruction to add material for a target subject, the material for the target subject can be displayed on the media creation page. Optionally, the material for the target subject includes a preset video and preset sound features. The preset video contains at least two video frames, each of which contains the target subject, and the target subject has at least a different viewpoint.

[0058] In some possible embodiments, each of the two video frames contains a target subject object, and the target subject object has different viewpoints and poses.

[0059] Figure 5 This is an interactive illustration of a media creation page according to an exemplary embodiment. Figure 3 ,like Figure 5 As shown, it includes a media creation page 300, and a content reference area 301, a video display area 302, a material addition control 303, and a material display area 400 located on the media creation page, as well as materials 401 and sound identifiers 402 of multiple first subject objects located in the material display area 400.

[0060] Optional, such as Figure 5 As shown, the entire content reference area can share a single material addition control. Optionally, the content reference area can be configured with a corresponding number of material addition controls based on the number of its sub-areas. For example, Figure 5 The content reference area includes both the first frame input area and the last frame input area, each of which can be configured with a material addition control.

[0061] Optionally, the material addition control can be placed anywhere on the media creation page or near the content reference area for easier and faster user access.

[0062] In the embodiments of this application, such as Figure 5 As shown in the first sub-image, the material addition control can be configured to appear below the content reference area. In response to a material addition command triggered by the material addition control, the material display area can be displayed on the media creation page. Specifically, when a preset operation (such as click, double-click, swipe, etc.) is detected in the material addition control, it can be displayed as follows: Figure 5 The final second sub-image shows the area where the material is displayed.

[0063] Optionally, the material display area 400 can always be displayed on the media creation page. When the material addition control is triggered, it means that the materials in the material display area can be used to match the target subject object or other materials for the preset image.

[0064] Optionally, the material display area 400 is not always displayed on the media creation page. Instead, it is displayed only when the material addition control is triggered. Subsequently, the materials in the material display area can be used to match target subjects or other materials with preset images. This allows the interface elements corresponding to the functions to be used to be displayed on the media creation page through interaction, and the interface elements can be hidden when there is no interaction, thereby reducing visual interference, improving space utilization, and enhancing the user experience.

[0065] In this embodiment of the application, the material display area can display materials 401 containing multiple first subject objects. In response to the confirmation instruction for adding material of a target subject object among the materials of multiple first subject objects, the identifier of the target subject object can be displayed on the material adding control.

[0066] like Figure 5 As shown, the materials of multiple first subject objects include two types of materials: one is video material including audio, and the other is video material without audio. The sound identifier 402 on the material of a first subject object indicates that the material of that first subject object is a video material including audio.

[0067] Figure 6 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 4 ,like Figure 6 As shown, it includes a media creation page 300, and a content reference area 301, a video display area 302, a material addition control 303, and a material display area 400 located on the media creation page, a material 401 and a sound identifier 402 of multiple first subject objects located in the material display area 400, and an identifier 410 of the target subject object displayed on the material addition control 303.

[0068] Optionally, when the material display area shows multiple materials for the primary subject object, the material for the primary subject object corresponding to the target subject object in the preset image can be selected. The material for the primary subject object is then used as the material for the target subject object. When a selection operation for the material of the target subject object is detected, i.e., in response to a confirmation instruction to add the material for the target subject object, the identifier of the target subject object can be displayed on the material addition control.

[0069] like Figure 6As shown, in the material display area, the material of the first main object located in the third position of the first row is selected as the material of the target main object. In this way, the identifier of the target main object can be displayed on the material adding control 303.

[0070] In this embodiment, the material of the first subject object corresponding to the target subject object in the preset image refers to the fact that the target subject object in the preset image and the subject object in the material of the corresponding first subject object are the same subject object. For example, if the target subject object in the preset image is person A, then the subject object in the material of the corresponding first subject object is also person A. In this way, the consistency of the subject object (such as person A) in the target video generated based on the preset image and the material of the corresponding first subject object can be guaranteed, and the situation of feature error of the subject object can be minimized or eliminated, such as the inconsistency between the features of the side face and the features of the front face.

[0071] In this embodiment of the application, the identifier 410 of the target subject object is displayed on the material addition control to indicate that the material of the target subject object has been selected.

[0072] In this embodiment of the application, the material of the target subject object can be at least one of multiple materials of the first subject object. That is, the number of materials of the target subject object can be determined based on the number of target subject objects contained in the preset image. If the preset image includes one target subject object, the number of materials of the target subject object is one. If one or more preset images include multiple target subject objects, the number of materials of the target subject object is the same as the number of target subject objects, and the materials of the target subject object must match the target subject objects, thereby preventing the problem that a preset image containing target subject object A cannot be matched with materials containing target subject object A, thus causing inconsistency of subjects.

[0073] In this way, when there are multiple target objects, the consistency of each target object can be guaranteed by providing materials for multiple target objects.

[0074] In some possible embodiments, the materials displayed in the material display area for the multiple first subject objects are all video materials of the first subject objects. Optionally, the materials of the multiple first subject objects may include video materials with audio (e.g., Figure 6 The material display area includes the materials of the first and third main objects in the first row, the materials of the second and fourth main objects in the second row, and video materials without audio (such as...). Figure 6The material display area shows the material of the second and fourth main objects in the first row, and the material of the first and third main objects in the second row.

[0075] Optionally, regardless of whether the video footage has audio or not, the video must contain at least two video frames, and the viewing angles of the target subject in these two video frames must be different. For example, a video may contain three video frames: the target subject is shown from the front in the first video frame, from the left side of the target subject in the second video frame, and from the right side of the target subject in the third video frame.

[0076] In this embodiment of the application, the material for each target subject includes a preset video containing at least two video frames and preset audio features. That is, the material for each target subject must not only include a video containing at least two video frames, but also audio features.

[0077] In some possible embodiments, the material of the target subject is a preset video comprising at least two video frames and preset sound features. That is, the material of the target subject is video material with its own audio (e.g.,...). Figure 6 The material display area shows the materials for the first and third main subjects in the first row, and the materials for the second and fourth main subjects in the second row. Because this type of material contains preset sound features, it can be directly used as material for the target subject.

[0078] In other possible embodiments, the material of the target subject may be a combination of silent video containing at least two video frames and other preset sound features.

[0079] In this embodiment, the material display area displays materials of multiple first subject objects, multiple first sound features, and an add confirmation control. Optionally, in response to a first add instruction for materials of a second subject object, the material of the selected second subject object can be displayed. The material of the second subject object can be silent video material (e.g., ...). Figure 6 The material display area shows the material of the second and fourth main objects in the first row, and the material of the first and third main objects in the second row.

[0080] Optionally, the number of materials for the second subject object can be determined based on the number of target subject objects in the preset image. Therefore, the number of materials for the second subject object is at least one of the materials for multiple first subject objects.

[0081] Optionally, when the material of the second subject object is selected, in response to the second add instruction for the preset sound features, the selected preset sound features can be displayed.

[0082] Similarly, the number of preset sound features can be determined based on the number of target objects in the preset image. Therefore, the preset sound features are at least one of a plurality of first sound features.

[0083] In this embodiment, when there are multiple target subjects in a preset image, multiple corresponding second subject materials can be selected for the multiple target subjects first, and then multiple corresponding preset sound features can be selected for the multiple target subjects. Alternatively, the corresponding second subject material and preset sound features can be selected for the first target subject, and then the corresponding second subject material and preset sound features can be selected for the next target subject.

[0084] Optionally, after selecting the corresponding second subject object's material and preset sound features for the target subject object, in response to the add confirmation command triggered by the add confirmation control, the target subject object's identifier can be displayed on the material add control. This identifier includes both the second subject object's identifier and the preset sound feature's identifier. Thus, the target subject object's material is obtained by combining the video frames corresponding to the second subject object's material with the preset sound features.

[0085] In other possible embodiments, the material of the target subject can be obtained by combining video frames from a video containing at least two video frames with other preset sound features. That is, in the process of generating the target video, only the video frames of the target subject containing multiple perspectives from the video containing sound are used, and the audio of the video itself is not used.

[0086] In this embodiment of the application, the preset sound features (including the audio that comes with the video and multiple first sound features) include at least one of pitch, timbre and dialect.

[0087] In other possible embodiments, the material of the target subject object may be obtained by combining video frames from a video containing at least two video frames, some sound features, and other preset sound features.

[0088] Optionally, when the material of the second subject object is a preset video including at least two video frames and a second sound feature, and a selected first sound feature exists among the multiple first sound features, the material of the target subject object may include the video frame corresponding to the material of the second subject object, the first type of sound feature among the second sound features, and the second type of sound feature among the selected first sound features.

[0089] For example, when the material of the second subject object is a video with built-in audio, and the first sound feature is selected from multiple first sound features, the timbre and pitch of the video frames and audio in the material of the second subject object can be preserved and combined with the Cantonese dialect in the first sound feature to obtain the material of the target subject object.

[0090] Thus, through the above embodiments, a flexible combination of video frames and sound features can be achieved, thereby generating target videos that better suit the user's preferences and are more flexible and versatile.

[0091] In this embodiment of the application, the materials of multiple first subject objects are sourced from a subject object library, and the materials of multiple first sound features are sourced from a sound library. Both the subject object library and the sound library can be sub-libraries in a comprehensive material library.

[0092] Optionally, the comprehensive material library may also include at least a scene material library, a prop material library, and a costume material library. The target scene in the scene material library, the target prop in the prop material library, and the target costume in the costume material library are used to combine the materials of the preset image and the target subject to generate a target video associated with the preset image.

[0093] Figure 7 This is a schematic diagram illustrating a material display area according to an exemplary embodiment, such as... Figure 7 As shown, it includes multiple material controls located on the material display area 400, such as the main material control 420, the sound material control 430, the scene material control, the prop material control, and the costume material control, etc.

[0094] In this embodiment, for the sake of page aesthetics and order, different material controls can be placed in the material display area to trigger the display of corresponding materials. For example, when the main material control is triggered, multiple materials of the first main object can be displayed; when the sound material control is triggered, multiple first sound features can be displayed.

[0095] This allows for flexible configuration of various materials in the material library, enabling the generation of videos in various styles and improving user convenience.

[0096] Figure 8 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 5 ,like Figure 8 As shown, it includes a media creation page 300, a material display area 400 and a subject creation page 600 located on the media creation page, and a subject object creation control 500 located in the material display area 400.

[0097] like Figure 8As shown in the first sub-image, the material display area also includes a main object creation control, which can be placed in front of multiple first main object materials.

[0098] Optional, such as Figure 8 As shown in the second sub-figure, the main object creation page is displayed in response to the main object creation command triggered by the main object creation control.

[0099] Optionally, the subject creation page 600 may include a subject image collection area and subject collection prompts. The subject image collection area instructs the user to place an image of their face within that area, and the subject collection prompts guide the user on how to perform the activity.

[0100] Optionally, when the image of the subject being collected is displayed in the subject image collection area, a subject recording instruction is triggered in response to the subject being collected moving according to the subject collection prompt information, thereby generating the collected subject material. The collected subject material includes video frames from different perspectives of the subject.

[0101] In this embodiment of the application, the collected subject object material is video material. The collected subject object material can be displayed as a first subject object material in the material display area and stored in the subject object library.

[0102] In this way, when there is a lack of material for the target subject, the material for the target subject can be quickly collected through the subject object creation control.

[0103] Furthermore, in order to collect sound features while collecting images, voice reading prompts can be displayed on the main body creation page.

[0104] In this embodiment, when the image of the collected subject object is displayed in the subject image collection area, a subject object recording instruction is triggered in response to the collected subject object performing activities according to the subject collection prompts, thereby generating a video of the collected subject object. Next, voice reading prompts can be displayed on the subject creation page. In response to the collected subject object reading according to the voice reading prompts, a voice recording instruction is triggered, thereby generating audio of the collected subject object. The video and audio of the collected subject object constitute the material of the collected subject object.

[0105] Optionally, the collected subject matter material can be displayed as a first subject matter material in the material display area and stored in the subject matter library.

[0106] Similarly, sound features from the sound library, scenes from the scene material library, props from the prop material library, and clothing from the clothing material library can also be collected in the same way, which will not be elaborated here.

[0107] In step S205, in response to a video creation instruction for materials containing a preset image and a target subject, a target video associated with the preset image is generated; the target video includes the target subject; the sound of the target subject in the target video is generated based on preset sound features.

[0108] In this embodiment of the application, in response to a video creation instruction for a preset image and a target subject object, a target video associated with the preset image is generated. The target video includes the target subject object, and the sound of the target subject object in the target video is generated based on preset sound features.

[0109] Optionally, the relationship between the target video and the preset image can mean that the preset image is a video frame in the target video, such as the first frame, a keyframe in the middle, or the last frame.

[0110] In this embodiment, video generation prompts can be displayed in the video prompt area of ​​the media creation page. Using the video generation prompts, in response to video creation instructions for materials based on preset images and target subjects, a target video associated with the preset images can be generated. The video generation prompts include video generation prompts, specified shot information, transition information, and video duration.

[0111] Figure 9 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 6 ,like Figure 9 As shown, it includes: a media creation page 300, and a content reference area 301, a video display area 302, and a video prompt area 700 located on the media creation page.

[0112] Optionally, in response to an input operation, a video generation prompt message corresponding to the target video can be displayed in the video prompt area.

[0113] In this way, in response to video creation instructions for preset images, video generation prompts, specified shot information, transition information, video duration, and target subject material, a target video associated with the preset image can be generated.

[0114] In this embodiment of the application, before generating the target video based on preset images, video generation prompts, specified shot information, transition information, video duration, and target subject material, an image sequence that conforms to video logic can be generated to ensure the logical accuracy of the final target video.

[0115] In this embodiment, in response to a video logic creation instruction for preset images, video generation prompts, specified shot information, transition information, video duration, and target subject material, a sequence of key node images including the subject object is displayed, wherein the temporal sequence of the key node images is consistent with the temporal sequence of the target video. In response to the video creation instruction for preset images, video generation prompts, specified shot information, transition information, video duration, target subject material, and key node image sequence, the target video is generated and played.

[0116] Taking the prompt "A girl is walking on the street when it suddenly starts raining. She takes an umbrella out of her bag and walks in the rain with it" as an example, after preparing preset images, video generation prompts, specified shot information, transition information, video duration, and target subject material on the video creation page, in response to the video logic creation instructions, multiple images can be displayed. These multiple images are a sequence of images ordered according to key nodes, and each image includes the target subject. Based on the prompt "A girl is walking on the street when it suddenly starts raining. She takes an umbrella out of her bag and walks in the rain with it," the system generates the first image corresponding to "A girl is walking on the street," the second image corresponding to "A girl takes an umbrella out of her bag," and the third image corresponding to "Walking in the rain with an umbrella." This demonstrates that the logic of the AI ​​in generating the target video is accurate.

[0117] If the generated keyframe image sequence is, in order, the first image is "a girl taking an umbrella out of her bag," the second image is "a girl walking on the street," and the third image is "walking in the rain with an umbrella," the image logic is clearly flawed, and therefore the logic of the final generated target video will also be incorrect. Therefore, if the displayed keyframe image sequence, including the main object, does not conform to logic, you can either re-enter the prompt information or regenerate the keyframe image sequence based on the prompt information.

[0118] Optionally, provided that the temporal sequence of key node images is determined and the logic is correct, the target video is generated and played in response to video creation instructions for preset images, video generation prompts, specified shot information, transition information, video duration, and target subject material.

[0119] In this way, with the provision of information from multiple sources, a target video with a consistent subject and logical consistency can be generated.

[0120] In some possible embodiments, the media creation page includes a subject association area. This subject association area is used to display association information between multiple subject objects in response to input operations when multiple target subject objects are preset to be included. Optionally, if the multiple target subject objects are two human figures, the association information may include their positional relationship, height / weight ratio, or closeness (e.g., lovers, friends). The association information between multiple target subject objects is used to instruct the generation of a target video containing the target subject objects, making the relationships between the multiple subject objects presented in the target video more consistent with the association information.

[0121] In summary, users only need to upload one main reference object. By binding the provided main object, the barrier to entry for users is lowered. By providing multiple perspectives and state diagrams for the main object, a high degree of consistency in the appearance and characteristics of the main object can be ensured when implementing different shots, actions, and scenes. This avoids problems such as deformation, deviation, and style drift of the main object's characteristics due to insufficient reference, thus making the video more professional.

[0122] Figure 10 A block diagram of a multimedia creation page interaction device is shown according to an exemplary embodiment. It has the function of implementing the data processing method in the above-described method embodiments; the function can be implemented in hardware or by hardware executing corresponding software. (Refer to...) Figure 10 The device includes a first display module 1001, a second display module 1002, and a third display module 1003. The first display module 1001 is configured to display a preset image containing the target subject on the media creation page; The second display module 1002 is configured to execute a response to an instruction to add material to a target subject object, and display the material of the target subject object on the media creation page; the material of the target subject object includes preset video and preset sound features containing at least two video frames; the target subject object in the two video frames has at least a different perspective; The third display module 1003 is configured to execute a video creation instruction in response to a preset image and a target subject object, and generate a target video associated with the preset image; the target video includes the target subject object; the sound of the target subject object in the target video is generated based on preset sound features.

[0123] In some possible embodiments, the media creation page includes material addition controls; a second display module is configured to perform: In response to a material addition command triggered by a material addition control, a material display area is displayed on the media creation page; the material display area displays materials with multiple primary subject objects; the materials with multiple primary subject objects originate from the subject object library; In response to a confirmation command to add material for a target subject object, an identifier of the target subject object is displayed on the material addition control; displaying the identifier of the target subject object on the material addition control indicates that the material for the target subject object has been selected; the material for the target subject object is at least one of multiple materials for a first subject object; The material for each target subject is a preset video containing at least two video frames and preset sound features.

[0124] In some possible embodiments, the material display area displays materials of multiple first subject objects, multiple first sound features, and an add confirmation control; the multiple first sound features are derived from a sound library; The second display module is configured to execute: In response to a first add command for material of a second subject object, material of the selected second subject object is displayed; the material of the second subject object is at least one of a plurality of materials of the first subject objects. In response to a second add instruction for a preset sound feature, the selected preset sound feature is displayed; the preset sound feature is at least one of a plurality of first sound features; In response to an add confirmation command triggered by the add confirmation control, the identifier of the target subject object is displayed on the material add control; the identifier of the target subject object includes the identifier of the second subject object and the identifier of the preset sound characteristics; The target subject's material includes the video frames and preset sound features corresponding to the material of the second subject.

[0125] In some possible embodiments, when the material of the second subject object is a preset video including at least two video frames and a second sound feature, and a selected first sound feature exists among the multiple first sound features, the material of the target subject object includes the video frame corresponding to the material of the second subject object, the first type of sound feature among the second sound features, and the second type of sound feature among the selected first sound features.

[0126] In some possible embodiments, the preset sound features include at least one of pitch, timbre, and dialect.

[0127] In some possible embodiments, the material display area further includes a subject object creation control; the apparatus also includes a subject creation module configured to perform: In response to a subject object creation command triggered by a subject object creation control, a subject creation page is displayed; the subject creation page includes a subject image collection area and subject collection prompts. When the image of the subject being collected is displayed in the subject image collection area, the subject recording instruction is triggered in response to the subject being collected performing activities according to the subject collection prompt information, thereby generating the material of the subject being collected; the material of the subject being collected is in video format. The collected material of the main object is stored in the main object library as material of a first main object.

[0128] In some possible embodiments, the subject creation module is configured to perform: When the image of the subject being collected is displayed in the subject image collection area, the subject recording instruction is triggered in response to the subject being collected performing activities according to the subject collection prompt information, thereby generating the video of the subject being collected. Display voice prompts on the main page; In response to the collected subject object reciting according to the voice reading prompt, the voice recording instruction is triggered to generate the audio of the collected subject object; The collected video and audio of the subject matter constitute the material of the collected subject matter.

[0129] In some possible embodiments, the preset image is an image containing the target subject object; the preset image is the first frame image, the last frame image, or an intermediate keyframe image in the target video; The preset images are two images, each containing the target subject; the two preset images are the first frame and the last frame of the target video, respectively.

[0130] In some possible embodiments, the preset images are a first preset image containing a first target subject and a second preset image containing a second target subject; the first preset image and the second preset image are the first frame and the last frame of the target video, respectively; The materials for the target subject include the materials corresponding to the first target subject and the materials corresponding to the second target subject.

[0131] In some possible embodiments, the main object library is a sub-library of the comprehensive material library; the comprehensive material library also includes at least a scene material library, a prop material library, and a costume material library; The target scenes in the scene material library, the target props in the prop material library, and the target clothing in the clothing material library are used to combine preset images and target subject materials to generate target videos associated with preset images.

[0132] In some possible embodiments, the display device further includes a fourth display module configured to perform: Display video generation prompts in the video prompt area on the media creation page; The third display module is configured to execute: The system generates prompts using video, responding to video creation instructions for preset images and target subjects, and generates target videos associated with the preset images. The video generation prompts include video generation prompts, specified shot information, transition information, and video duration.

[0133] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0134] Figure 11 This is a block diagram illustrating a page interaction device 3000 for multimedia authoring according to an exemplary embodiment. For example, device 3000 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0135] Reference Figure 11 The device 3000 may include one or more of the following components: a processing component 3002, a memory 3004, a power component 3006, a multimedia component 3008, an audio component 3010, an input / output (I / O) interface 3012, a sensor component 3014, and a communication component 3016.

[0136] Processing component 3002 typically controls the overall operation of device 3000, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 3002 may include one or more processors 3020 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 3002 may include one or more modules to facilitate interaction between processing component 3002 and other components. For example, processing component 3002 may include a multimedia module to facilitate interaction between multimedia component 3008 and processing component 3002.

[0137] Memory 3004 is configured to store various image types of data to support operation of device 3000. Examples of this data include instructions for any application or method operating on device 3000, contact data, phonebook data, messages, pictures, videos, etc. Memory 3004 can be implemented by any image type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0138] Power supply component 3006 provides power to various components of device 3000. Power supply component 3006 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 3000.

[0139] Multimedia component 3008 includes a screen that provides an output interface between the device 3000 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 3008 includes a front-facing camera and / or a rear-facing camera. When the device 3000 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0140] Audio component 3010 is configured to output and / or input audio signals. For example, audio component 3010 includes a microphone (MIC) configured to receive external audio signals when device 3000 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 3004 or transmitted via communication component 3016. In some embodiments, audio component 3010 also includes a speaker for outputting audio signals.

[0141] I / O interface 3012 provides an interface between processing component 3002 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0142] Sensor assembly 3014 includes one or more sensors for providing status assessments of various aspects of device 3000. For example, sensor assembly 3014 may detect the on / off state of device 3000, the relative positioning of components such as the display and keypad of device 3000, changes in the position of device 3000 or a component of device 3000, the presence or absence of user contact with device 3000, the orientation or acceleration / deceleration of device 3000, and temperature changes of device 3000. Sensor assembly 3014 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 3014 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 3014 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0143] Communication component 3016 is configured to facilitate wired or wireless communication between device 3000 and other devices. Device 3000 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 3G, 4G, or 5G), or combinations thereof. In one exemplary embodiment, communication component 3016 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 3016 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0144] In an exemplary embodiment, the apparatus 3000 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0145] Embodiments of the present invention also provide a computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a multimedia creation page interaction method. The at least one instruction or the at least one program is loaded and executed by the processor to implement the multimedia creation page interaction method provided in the above method embodiments.

[0146] In an exemplary embodiment, a storage medium including instructions is also provided, such as a memory 3004 including instructions, which can be executed by a processor 3020 of the device 3000 to perform the above-described method. Optionally, the storage medium may be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device.

[0147] Embodiments of the present invention also provide a computer-readable storage medium that, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform the method of any one of the first or second aspects of the embodiments of the present disclosure.

[0148] Embodiments of the present invention also provide a computer program product comprising a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from the readable storage medium and executes the computer program, causing the computer device to perform the method of any one of the first or second aspects of the embodiments of the present disclosure.

[0149] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, interactive page layout and parallel processing for multimedia authoring are possible or may be advantageous.

[0150] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0151] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0152] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A multimedia creation page interaction method, characterized in that, Applied to the first client, including: Display a preset image containing the target subject on the media creation page; In response to an instruction to add material for the target subject, the material for the target subject is displayed on the media creation page; the material for the target subject includes a preset video and preset sound features containing at least two video frames; the target subject in the two video frames has at least a different perspective; In response to a video creation instruction for the preset image and the target subject object, a target video associated with the preset image is generated; the target video includes the target subject object; the sound of the target subject object in the target video is generated based on the preset sound features.

2. The multimedia creation page interaction method according to claim 1, characterized in that, The media creation page includes a material addition control; the step of displaying the material of the target subject object on the media creation page in response to the material addition instruction for the target subject object includes: In response to a material addition command triggered by the material addition control, a material display area is displayed on the media creation page; the material display area displays materials of multiple first subject objects; the materials of the multiple first subject objects are sourced from the subject object library; In response to a confirmation instruction to add material for the target subject object, an identifier of the target subject object is displayed on the material adding control; the display of the identifier of the target subject object on the material adding control indicates that the material of the target subject object has been selected; the material of the target subject object is at least one of the materials of the plurality of first subject objects; The material for each target subject is a preset video containing at least two video frames and the preset sound features.

3. The multimedia creation page interaction method according to claim 2, characterized in that, The material display area displays materials of multiple first main objects, multiple first sound features, and an add confirmation control; the multiple first sound features are derived from a sound library; The step of displaying the identifier of the target subject object on the material adding control in response to a confirmation instruction for adding material for the target subject object includes: In response to a first add instruction for material of the second subject object, material of the selected second subject object is displayed; the material of the second subject object is at least one of the materials of the plurality of first subject objects; In response to a second add instruction for the preset sound feature, the selected preset sound feature is displayed; the preset sound feature is at least one of the plurality of first sound features; In response to an add confirmation command triggered by the add confirmation control, the identifier of the target subject object is displayed on the material add control; the identifier of the target subject object includes the identifier of the second subject object and the identifier of the preset sound feature; The material of the target subject includes the video frames corresponding to the material of the second subject and the preset sound features.

4. The multimedia creation page interaction method according to claim 3, characterized in that, When the material of the second subject object is a preset video containing at least two video frames and a second sound feature, and a selected first sound feature exists among multiple first sound features, the material of the target subject object includes the video frame corresponding to the material of the second subject object, the first type of sound feature among the second sound features, and the second type of sound feature among the selected first sound features.

5. The multimedia creation page interaction method according to any one of claims 1-4, characterized in that, The preset sound features include at least one of pitch, timbre, and dialect.

6. The multimedia creation page interaction method according to any one of claims 2-4, characterized in that, The material display area also includes a main object creation control; the method further includes: In response to a subject object creation instruction triggered by the subject object creation control, a subject creation page is displayed; the subject creation page includes a subject image collection area and subject collection prompts. When the image of the subject being collected is displayed in the subject image collection area, in response to the subject being collected performing activities according to the subject collection prompt information, a subject recording instruction is triggered to generate the material of the subject being collected; the material of the subject being collected is in video format. The collected material of the main object is stored in the main object library as material of a first main object.

7. The multimedia creation page interaction method according to claim 6, characterized in that, When the image of the subject being collected is displayed in the subject image collection area, in response to the subject being collected performing activities according to the subject collection prompt information, a subject recording instruction is triggered to generate the material of the collected subject, including: When the image of the subject being collected is displayed in the subject image collection area, in response to the subject being collected performing activities according to the subject collection prompt information, a subject recording instruction is triggered to generate a video of the subject being collected. The main body creation page displays voice reading prompts; In response to the collected subject object reciting according to the voice recitation prompt information, thereby triggering a voice recording instruction, audio of the collected subject object is generated; The collected video and audio of the subject object constitute the material of the collected subject object.

8. The multimedia creation page interaction method according to any one of claims 1-4, characterized in that, The preset image is an image containing the target subject; the preset image is the first frame, the last frame, or a keyframe in the target video. The preset images are two images, each containing the target subject; the two preset images are the first frame and the last frame of the target video, respectively.

9. The multimedia creation page interaction method according to any one of claims 1-4, characterized in that, The preset images are a first preset image containing a first target subject and a second preset image containing a second target subject; the first preset image and the second preset image are respectively the first frame image and the last frame image in the target video; The materials for the target subject include the materials corresponding to the first target subject and the materials corresponding to the second target subject.

10. The multimedia creation page interaction method according to any one of claims 2-4, characterized in that, The main object library is a sub-library of the comprehensive material library; the comprehensive material library also includes at least a scene material library, a prop material library, and a costume material library; The target scene in the scene material library, the target prop in the prop material library, and the target clothing in the clothing material library are used to combine the preset image and the material of the target subject to generate a target video associated with the preset image.

11. The multimedia creation page interaction method according to any one of claims 1-4, characterized in that, The method further includes: Video generation prompts are displayed in the video prompt area of ​​the media creation page. The step of generating a target video associated with the preset image in response to a video creation instruction for the preset image and the target subject object includes: Using the video to generate prompt information, in response to a video creation instruction for the preset image and the target subject object, a target video associated with the preset image is generated; The video generation prompt information includes video generation prompt words, specified shot information, transition information, and video duration.

12. A multimedia creation page interactive device, characterized in that, include: The first display module is configured to display a preset image containing the target subject on the media creation page; The second display module is configured to execute a material addition instruction for the target subject object and display the material of the target subject object on the media creation page; the material of the target subject object includes a preset video and preset sound features containing at least two video frames; the target subject object in the two video frames has at least a different perspective; The third display module is configured to execute a video creation instruction in response to the preset image and the material of the target subject object, and generate a target video associated with the preset image; the target video includes the target subject object; the sound of the target subject object in the target video is generated based on the preset sound features.

13. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the multimedia creation page interaction method as described in any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the multimedia authoring page interaction method as described in any one of claims 1 to 11.

15. A computer program product, characterized in that, The computer program product includes a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from and executes the computer program, causing the device to perform a multimedia authoring page interaction method as described in any one of claims 1 to 11.