Page interaction method and device for multimedia creation, electronic equipment and storage medium

By displaying media content from multiple perspectives on the media creation page and utilizing feature decomposition and identification functions, the problem of inconsistent features of the main object in AI-generated videos has been solved, achieving a high degree of consistency in the appearance and features of the main object in the video and generating more professional videos.

CN121832811APending Publication Date: 2026-04-10BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
Filing Date
2025-11-24
Publication Date
2026-04-10

Smart Images

  • Figure CN121832811A_ABST
    Figure CN121832811A_ABST
Patent Text Reader

Abstract

The invention relates to a multimedia creation page interaction method and device, electronic equipment and a storage medium. The method comprises the steps that first media content containing a main body object is displayed on a media creation page; displaying at least one second media content in response to a content generation instruction for the first media content; the second media content comprises a main body object; the main body object in the first media content and the main body object in the second media content at least have different viewing angles; generating a target video in response to a video creation instruction for the first media content and the second media content; the target video comprises the main body object. According to the embodiment of the invention, a plurality of view angles and state diagrams can be provided for the main body object, so that the appearance and the feature of the main body object are highly consistent when different lenses, actions and scenes are realized, the problems of feature deformation, deviation, style drift and the like of the main body object caused by insufficient reference are avoided, and the accuracy of the image processing is improved. Therefore, the video is more professional.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of Internet, and particularly relates to a page interaction method and device for multimedia creation, an electronic device and a storage medium. BACKGROUND

[0002] With the deep penetration of artificial intelligence (AI) generation models in the field of video creation, AI-generated videos have been widely used in e-commerce advertising, short drama animation, content marketing and other scenarios. However, the current technology still faces a core bottleneck, namely the problem of theme consistency. The problem of theme consistency refers to the drift, deformation or loss of the features (such as shape, style, logical relationship) of core elements (such as characters, goods and scenes) in the generated video between frames, which seriously affects the content credibility and dissemination effect. SUMMARY

[0003] The present disclosure provides a page interaction method and device for multimedia creation, an electronic device and a storage medium. The technical solutions of the present disclosure are as follows: According to a first aspect of an embodiment of the present disclosure, a page interaction method for multimedia creation is provided, comprising: displaying a first media content containing a subject object on a media creation page; in response to a content generation instruction for the first media content, displaying at least one second media content; the second media content includes the subject object; the subject object in the first media content and the subject object in the second media content at least have different perspectives; in response to a video creation instruction for the first media content and the second media content, generating a target video; the target video includes the subject object.

[0004] In some possible embodiments, in response to the content generation instruction for the first media content, displaying at least one second media content comprises: in response to the content generation instruction for the first media content, generating a plurality of candidate media contents; in response to a selection instruction for the plurality of candidate media contents, displaying at least one second media content; the at least one second media content is a part of the plurality of candidate media contents.

[0005] In some possible embodiments, the media creation page includes a feature display area; in response to the content generation instruction for the first media content, generating a plurality of candidate media contents comprises: in response to a feature decomposition instruction for the subject object in the first media content, displaying, in the feature display area, a plurality of features corresponding to the subject object and description information corresponding to each feature in the plurality of features; In response to the content generation instruction, display a plurality of candidate media contents generated based on the first media content, the plurality of features, and description information corresponding to each feature in the plurality of features.

[0006] In some possible embodiments, the media creation page includes a deletion control of each feature and a feature addition control; In response to the content generation instruction, before displaying, on the media creation page, a plurality of candidate media contents generated based on the first media content, the plurality of features, and description information corresponding to each feature in the plurality of features, the method further includes: In response to a feature addition instruction triggered based on the feature addition control, display a feature addition area on the media creation page; the feature addition area includes a picture addition area or a text addition area; When the picture addition area adds a picture of a feature to be added or the text addition area adds text of the feature to be added, in response to a feature positioning operation on the first media content, display, in the feature display area, the feature to be added and description information corresponding to the feature to be added.

[0007] In some possible embodiments, in response to a video creation instruction on the first media content and the second media content, generating a target video includes: In response to a video creation instruction on the first media content, the second media content, the plurality of features, and description information corresponding to each feature in the plurality of features, generating the target video.

[0008] In some possible embodiments, after displaying, in the feature display area, the plurality of features corresponding to the subject object and the description information corresponding to each feature in the plurality of features, the method further includes: In response to an identification addition operation on a target feature in the plurality of features, display a preset identification in a preset area corresponding to the target feature; the preset identification includes a fixed feature identification. The target feature carrying the fixed feature identification indicates that the subject object in the target video always includes the target feature.

[0009] In some possible embodiments, the media creation page includes a prompt information area corresponding to prompt information; the method further includes: In response to an input operation on the input area for the prompt information, display, in the prompt information area, prompt information corresponding to the target video; the prompt information corresponding to the target video is used to guide a generation logic of the target video.

[0010] In some possible embodiments, in response to a video creation instruction on the first media content and the second media content, generating a target video includes: In response to video logic creation instructions for primary and secondary media content, a sequence of key node images, including the main object, is displayed; the temporal sequence of the key node images is consistent with the temporal sequence of the target video. In response to video creation instructions for primary media content, secondary media content, and key node image sequences, generate the target video.

[0011] In some possible embodiments, the content creation page includes a subject association area; the subject association area is used to display association information between multiple subject objects in response to an input operation in the subject association area when the first media content includes multiple subject objects; the association information between multiple subject objects is used to indicate the generation of a target video containing subject objects.

[0012] In some possible embodiments, the subject object includes at least one of a person, an animal, and an object; The first media content is media content that has undergone background processing; the first media content includes a first video or a first image, and the second media content is a second video or a second image; the second video is a video clip from the first video; The perspectives or states of any two subjects in the primary media content and the secondary media content are different; states include actions and expressions.

[0013] According to a second aspect of the present disclosure, a multimedia authoring page interaction device is provided, comprising: The first display module is configured to display the first media content containing the main object on the media creation page; The second display module is configured to execute a content generation instruction in response to the first media content and display at least one second media content; the second media content includes a subject object; the subject object in the first media content and the subject object in the second media content have at least different perspectives; The third display module is configured to execute video creation instructions in response to the first media content and the second media content, and generate a target video; the target video includes a main object.

[0014] In some possible embodiments, the second display module is configured to perform: In response to a content generation instruction for the primary media content, generate multiple candidate media content; In response to a selection instruction for multiple candidate media content, at least one second media content is displayed; the at least one second media content is a part of the multiple candidate media content.

[0015] In some possible embodiments, the media creation page includes a feature display area; a second display module is configured to perform: In response to a feature decomposition instruction for the main object in the first media content, multiple features corresponding to the main object and descriptive information for each feature are displayed in the feature display area. In response to a content generation instruction, multiple candidate media contents are displayed based on a first media content, multiple features, and descriptive information corresponding to each of the multiple features.

[0016] In some possible embodiments, the media creation page includes a delete control and a feature addition control for each feature; The device also includes an information adding module, configured to perform: In response to a feature addition command triggered by a feature addition control, a feature addition area is displayed on the media creation page; the feature addition area may include an image addition area or a text addition area. When an image with features to be added is added to the image addition area, or text with features to be added is added to the text addition area, in response to the feature positioning operation for the first media content, the feature to be added and its corresponding descriptive information are displayed in the feature display area.

[0017] In some possible embodiments, the third display module is configured to perform: In response to video creation instructions for first media content, second media content, multiple features, and descriptive information corresponding to each of the multiple features, a target video is generated.

[0018] In some possible embodiments, the apparatus further includes an identifier adding module configured to perform: In response to an operation that adds an identifier to a target feature among multiple features, a preset identifier is displayed in a preset area corresponding to the target feature; the preset identifier includes a fixed feature identifier; Among them, the target feature carrying a fixed feature identifier indicates that the main object in the target video always includes the target feature.

[0019] In some possible embodiments, the media creation page includes a prompt information area corresponding to the prompt information; the device also includes a prompt information input module configured to perform: In response to input operations on the prompts in the input area, the prompts corresponding to the target video are displayed in the prompt area; the prompts corresponding to the target video are used to guide the generation logic of the target video.

[0020] In some possible embodiments, the third display module is configured to perform: In response to video logic creation instructions for primary and secondary media content, a sequence of key node images, including the main object, is displayed; the temporal sequence of the key node images is consistent with the temporal sequence of the target video. In response to video creation instructions for primary media content, secondary media content, and key node image sequences, generate the target video.

[0021] In some possible embodiments, the content creation page includes a subject association area; the subject association area is used to display association information between multiple subject objects in response to an input operation in the subject association area when the first media content includes multiple subject objects; the association information between multiple subject objects is used to indicate the generation of a target video containing subject objects.

[0022] In some possible embodiments, the subject object includes at least one of a person, an animal, and an object; The first media content is media content that has undergone background processing; the first media content includes a first video or a first image, and the second media content is a second video or a second image; the second video is a video clip from the first video; The perspectives or states of any two subjects in the primary media content and the secondary media content are different; states include actions and expressions.

[0023] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute instructions to implement the method as described in either the first or second aspect above.

[0024] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method of any one of the first or second aspects of the present disclosure.

[0025] According to a fifth aspect of the present disclosure, a computer program product is provided, the computer program product including a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from the readable storage medium and executes the computer program, causing the computer device to perform the method of any one of the first or second aspects of the present disclosure.

[0026] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: The media creation page displays first media content containing the main object; in response to a content generation instruction for the first media content, at least one second media content is displayed; the second media content includes the main object; the main object in the first media content and the main object in the second media content have at least different perspectives; in response to a video creation instruction for the first media content and the second media content, a target video is generated; the target video includes the main object. This embodiment of the application can ensure a high degree of consistency in the appearance and characteristics of the main object when implementing different shots, actions, and scenes by providing multiple perspectives and state diagrams for the main object, avoiding problems such as deformation, deviation, and style drift of the main object's characteristics due to insufficient reference, thereby making the video more professional.

[0027] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 This is a schematic diagram illustrating an application environment for a multimedia creation page interaction method according to an exemplary embodiment; Figure 2 This is a flowchart illustrating a multimedia creation page interaction method according to an exemplary embodiment; Figure 3 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 1 ; Figure 4 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 2 ; Figure 5 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 3 ; Figure 6 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 4 ; Figure 7 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 5 ; Figure 8 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 6; Figure 9 This is a block diagram illustrating a multimedia creation page interaction device according to an exemplary embodiment; Figure 10 This is a block diagram illustrating an electronic device for interactive page creation for multimedia authoring, according to an exemplary embodiment. Detailed Implementation

[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0031] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar first objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0032] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0033] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application environment for a multimedia creation page interaction method according to an exemplary embodiment, such as... Figure 1 As shown, the application environment may include server 011 and client 012.

[0034] In some possible embodiments, server 011 may be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The operating system running on the server may include, but is not limited to, Android, iOS, Linux, Windows, Unix, etc.

[0035] In some possible embodiments, the client 012 described above may include, but is not limited to, image-type clients such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices. It may also be software running on the client, such as applications or mini-programs. Optionally, the operating system running on the client may include, but is not limited to, Android, iOS, Linux, Windows, and Unix systems.

[0036] In some possible embodiments, client 012 displays first media content containing the main object on the media creation page; in response to a content generation instruction for the first media content, it displays at least one second media content; the second media content includes the main object; the main object in the first media content and the main object in the second media content have at least different perspectives; in response to a video creation instruction for the first media content and the second media content, it generates a target video; the target video includes the main object. This embodiment of the application can ensure a high degree of consistency in the appearance and features of the main object when implementing different shots, actions, and scenes by providing multiple perspectives and state diagrams for the main object, avoiding problems such as deformation, deviation, and style drift of the main object's features due to insufficient reference, thereby making the video more professional.

[0037] In one exemplary implementation, both the client and server databases can be node devices in the blockchain system, capable of sharing acquired and generated information with other node devices within the blockchain system, thus enabling information sharing among multiple node devices. Multiple node devices in the blockchain system can be configured with the same blockchain, which consists of multiple blocks. Adjacent blocks are related, ensuring that any data tampering in any block can be detected by the next block, thereby preventing data tampering and guaranteeing the security and reliability of the data in the blockchain.

[0038] Figure 2 This is a flowchart illustrating a multimedia creation page interaction method according to an exemplary embodiment. It should be noted that this specification provides the operational steps of the method described in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operational steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many steps and does not represent the only execution order. In actual system or product execution, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment). Specifically, as shown... Figure 2 As shown, this flowchart includes at least the following steps S201-S203: In step S201, the first media content containing the main object is displayed on the media creation page.

[0039] In this embodiment, the media creation page can be a page provided in an application for creating media content. This target page displays various interactive areas, which serve as entry points for users to input, edit, and view media content. This media content can be the media materials required by the application for media content generation. These media materials can cover single or multiple types of media content. This embodiment does not limit the type of media content, including text, audio, images, video, or other formats.

[0040] In this embodiment of the application, the first media content is the reference media content used to generate the target video. The first media content may include images or videos. This embodiment of the application does not limit the number of first media contents, which may be one or more.

[0041] In some possible embodiments, the number of first media contents is one, such as including a first image, or including a first video.

[0042] In this embodiment, the main object refers to the most central and attractive object in the frame of the first media content, which carries the main information and expresses the purpose of the content. Other objects in the media content exist to complement it.

[0043] Optionally, this application embodiment does not limit the number or type of main objects in the first media content. Optionally, the main objects in the first media content can be one or more, and the categories of the main objects in the first media content can include at least one of people, animals, items, and scenes.

[0044] For example, the first media content includes one main object: Person A. For example, the first media content includes two main objects: Person A and Person B. For example, the first media content includes two main objects: Person A and Animal A. For example, the first media content includes two main objects: Person A and Item A.

[0045] Figure 3 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 1 ,like Figure 3 As shown, it includes a media creation page 300, and a first area 301, a second area 302, a third area 303 and a creation control 304 located on the media creation page.

[0046] like Figure 3 As shown, the first area 301 is used to input the main reference content, that is, to input the first media content containing the main object. The second area 302 is used to input reference content from other perspectives or to display second media content generated based on the first media content. The third area 303 is used to input descriptive information of the main object within the first media. The authoring control 304 is used to trigger the generation of the final target video.

[0047] Optionally, when multiple reference contents exist corresponding to multiple perspectives, a primary reference content can be selected as the first media content and added to the first area, while the other reference contents can be input as the second media content into the second area. The perspectives of the main objects in the first and second media content are different.

[0048] Optionally, when a reference content exists corresponding to a viewpoint, the reference content can be added to the first area as the first media content, and a second media content can be generated based on the first media content and displayed in the second area. The viewpoints of the main objects in the first and second media content are different.

[0049] Optionally, the main subject in the first media content is shown from the front, while the main subject in the second media content is shown from the side or back. The main subject in the second media content can present various side views depending on the angle from which it is viewed.

[0050] In step S203, in response to the content generation instruction for the first media content, at least one second media content is displayed; the second media content includes a subject object; the subject object in the first media content and the subject object in the second media content have at least different perspectives.

[0051] In some possible embodiments, the second media content may be media content generated based on the primary reference content (i.e., the first media content) in the presence of a reference content.

[0052] In other possible embodiments, the second media content may be media content generated based on the multiple main reference contents (i.e., multiple first media contents) when multiple first media contents exist. Optionally, when multiple first media contents exist, they may contain the same subject but have different perspectives.

[0053] In this embodiment of the application, the first media content may include images or videos. This embodiment of the application does not limit the number of second media contents, which may be one or more.

[0054] Optionally, the number of second media content items may be one, such as including a second image or a second video. Optionally, the number of second media content items may be multiple, such as including multiple second images or multiple second videos.

[0055] In some possible embodiments, when there are one or more second videos, the second video may be a video segment extracted from the first video.

[0056] In this embodiment, the subject objects in the first media content and the subject objects in the second media content have different perspectives, including different perspectives and states. Optionally, any two subject objects in the first media content and the second media content may have different perspectives or different states. Optionally, states include actions and expressions; for example, states include standing, running, jumping, raising an arm, etc. Optionally, expressions include smiling, being angry, crying, etc.

[0057] For example, the perspective of the main subject in the first media content and the second media content can be the same, but the main subject in the first media content is standing, while the main subject in the second media content is running. The state of the main subject in both the first and second media content can be the same, both smiling and standing still, but the main subject in the first media content is facing forward, while the main subject in the second media content is from the side. The state and perspective of the main subject in both the first and second media content can be different.

[0058] Figure 4 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 2 ,like Figure 4 As shown, it includes a media creation page 300, and a first area 301, a second area 302, a third area 303 and a creation control 304 located on the media creation page, and also includes a content generation control 3021 located in the second area 302.

[0059] In some possible embodiments, when the first area displays input first media content and the content generation control 3021 is triggered, a content generation instruction for the first media content can be generated, and at least one second media content can be generated based on the content generation instruction and displayed in the second area.

[0060] In some other possible embodiments, when the first area displays the input first media content, the third area displays the description information of the main object in the first media content, and the content generation control 3021 is triggered, a content generation instruction for the first media content can be generated, and at least one second media content can be generated based on the content generation instruction and displayed in the second area.

[0061] Therefore, when generating second media content, not only can the first media content be referenced, but the description of the main object in the third area can also be used to make the identification of the main object more precise. Especially when there are multiple objects in the first media content (such as girl A and girl B), but the main object is girl A, that is, when the main object in the generated second media content is girl A, the description of the main object in the third area can quickly identify the main object in the generated second media content, avoiding interference from girl B.

[0062] In order to reduce the interference of the background corresponding to the main object in the first media content on the main object, and to reduce the deviation of the features of the main object in the generated second media content, such as treating the features in the background as part of the main object, in this embodiment of the application, the first media content can be processed, such as blurring the background or cutting out the main object in the first media content.

[0063] Figure 5 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 3 ,like Figure 5 As shown, it includes a media creation page 300, and a first area 301, a second area 302 and a creation control 304 located on the media creation page.

[0064] like Figure 5 As shown, in some possible embodiments, when the first area displays the input first media content and the content generation control is triggered, a content generation instruction for the first media content can be generated, and multiple candidate media contents can be generated based on the content generation instruction. These multiple candidate media contents can be displayed in the second area, and the user can view each candidate media content by swiping up and down.

[0065] Optionally, in response to a selection instruction for multiple candidate media content, the client can select a portion of the candidate media content as the second media content and display it in the second area, while the other candidate media content can disappear from the second area.

[0066] Optionally, after generating multiple candidate media content, a "Regenerate" control can also be displayed in the second area. This "Regenerate" control is used to regenerate candidate media content based at least on the first media content.

[0067] It is evident that the client can use AI to intelligently complete at least the first media content to generate multiple candidate media content, and give technical personnel a certain degree of choice to select the more suitable second media content with more diverse perspectives or states, thereby providing a foundation for the subsequent generation of vivid, consistent target videos containing the main subject.

[0068] In step S205, in response to video creation instructions for the first media content and the second media content, a target video is generated; the target video includes a main object.

[0069] In some possible embodiments, when the media creation page displays first media content and multiple identified second media, when the creation control is detected to be triggered, a video creation instruction for the first media content and the second media content can be generated, and a target video can be generated based on the video creation instruction, which can then be played on the media creation page.

[0070] In other possible embodiments, the target video can also be generated based on the prompt information corresponding to the target video, the first media content, and a plurality of determined second media. Figure 6 A schematic diagram of a media creation page according to an exemplary embodiment. Figure 4 ,like Figure 6 As shown, it includes a media creation page 300, and a first area 301, a second area 302, a prompt information area 400, and a creation control 304 located on the media creation page.

[0071] like Figure 6 As shown, when the media creation page displays the first media content and multiple identified second media, and the prompt message "A girl is walking on the street when it suddenly starts to rain. The girl takes out an umbrella from her bag and walks in the rain with the umbrella open," when the creation control is detected to be triggered, a video creation instruction for the first and second media content can be generated, and a target video can be generated based on the video creation instruction, which can then be played on the media creation page.

[0072] The prompts for the target video guide the generation logic of the target video, such as the order of the videos, the style of the videos, and what details to pay attention to.

[0073] In this application embodiment, the generation of the second media content or target video may lack a certain feature of the subject object or have an extra feature, resulting in inconsistency of the subject object. Based on this, this application can perform targeted feature processing on the first media content to ensure the consistency of the subject object's features in the future.

[0074] Figure 7 A schematic diagram of a media creation page according to an exemplary embodiment. Figure 5 ,like Figure 7 As shown, it includes a media creation page 300, and a first area 301, a second area 302, a feature display area 500, and a creation control 304 located on the media creation page.

[0075] like Figure 7 As shown, after displaying the first media content in the first area, in response to the feature decomposition instruction for the main object in the first media content, the client can display multiple features corresponding to the main object and the descriptive information corresponding to each of the multiple features in the feature display area 500. Then, in response to the content generation instruction, the client can display multiple candidate media contents generated based on the first media content, multiple features, and the descriptive information corresponding to each of the multiple features.

[0076] Optionally, the video creation page includes a feature decomposition control. Specifically, after the first media content is displayed in the first area, in response to a feature decomposition instruction triggered by the feature decomposition control for the main object in the first media content, the client can perform feature decomposition on the first media content to obtain the various features of the main object and the descriptive information of each feature.

[0077] In this embodiment, various features can be listed and displayed in the feature display area. Each feature may include a feature image corresponding to the feature and descriptive information for that feature. Optionally, the features of the subject object may include features of the subject object itself and features of the subject object's clothing. Features of the subject object itself may include hairstyle, face shape, body shape, etc. Features of the subject object's clothing may include clothing, gestures, backpack, etc.

[0078] Taking hairstyle as an example, the feature display area can show the name of the feature, such as "hairstyle," and display an illustration of the hairstyle corresponding to the subject object, as well as the corresponding descriptive information, such as "long hair, slightly wavy, black." Taking earrings as an example, the feature display area can show the name of the feature, such as "earrings," and display an illustration of the hairstyle corresponding to the subject object, as well as the corresponding descriptive information, such as "pearls, flowers."

[0079] Optionally, the information in the feature display area can be modified. For example, a user can change the diagram or description of a particular feature. This allows for optimization based on user changes when the client's feature decomposition and summary is incorrect, resulting in more accurate feature information in the display area.

[0080] In some possible embodiments, the media creation page includes delete controls for each feature and feature addition controls. The delete control for each feature can remove all information about the corresponding feature from the feature display area. This can eliminate errors that may occur during AI feature decomposition, ensuring the accuracy of feature decomposition. Alternatively, an unimportant feature can be deleted so that it is not considered when generating secondary media content.

[0081] Optionally, the feature addition control is used to add features that exist in the subject object but have not been decomposed into the feature display area. This provides a remedial measure in case the client does not decompose the feature completely correctly. Optionally, in response to a feature addition command triggered by the feature addition control, the client can display the feature addition area on the media creation page. Optionally, this feature addition area can be located in the feature display area, and may include an image addition area or a text addition area.

[0082] When an image with features to be added is added to the image addition area, or text with features to be added to the text addition area, in response to the feature localization operation for the first media content, the client can display the feature to be added and its corresponding descriptive information in the feature display area. For example, if the client does not display the glasses worn by the main object in the feature display area during feature decomposition, the user can add an image of glasses to the image addition area and / or add a text description of eyes to the text addition area. Then, in response to the feature localization operation for the first media content, the client can display the descriptive information for the eyes and glasses in the feature display area.

[0083] In this way, by breaking down and listing the features of the primary media content, the process of generating multiple candidate media content on the client side can focus more on the features of the main object, thus avoiding feature bias.

[0084] In some possible embodiments, the client can generate multiple candidate media contents based on a first media content, multiple features, and descriptive information corresponding to each of the multiple features. Then, after determining a second media content from the multiple candidate media contents, the client can generate a target video based on the first media content and the second media content.

[0085] In other possible embodiments, the client can generate multiple candidate media contents based on a first media content, multiple features and description information corresponding to each of the multiple features. Then, after determining a second media content from the multiple candidate media contents, in response to a video creation instruction for the first media content, the second media content, multiple features and description information corresponding to each of the multiple features, the client can generate and play a target video.

[0086] In some possible embodiments, the features of the main object can also be marked with preset identifiers so that the content in the target video can be controlled based on the functions corresponding to the preset identifiers during the process of generating the target video using artificial intelligence.

[0087] In this embodiment, the preset identifier includes a fixed feature identifier. In response to an identifier addition operation for a target feature among multiple features, the fixed feature identifier is displayed in a preset area corresponding to the target feature. The target feature carrying the fixed feature identifier indicates that the main object in the target video always includes the target feature. For example, when the hairstyle feature is marked with a fixed feature identifier, it indicates that the hairstyle of the main object in the generated target video cannot be changed later.

[0088] If a feature is not marked with a fixed feature identifier, it means that the feature may or may not appear on the main object in the target video, or it may appear in some segments of the target video but not in others. This application does not restrict whether features not marked with fixed feature identifiers appear in the target video; rather, it is based on video logic, allowing for reasonable appearance.

[0089] In this embodiment, the media creation page also includes an input area for prompting information, which guides the generation logic of the target video. Figure 8 This is a schematic diagram of a media creation page according to an exemplary embodiment. Figure 6 ,like Figure 8 As shown, it includes a media creation page 300, and a first area 301, a second area 302, a feature display area 500, a prompt information area 400, and a creation control 304 located on the media creation page.

[0090] Optionally, in response to input operations on prompts in the input area, the client can display prompts corresponding to the target video in the prompt area.

[0091] In some possible embodiments, the client can generate multiple candidate media contents based on the first media content, multiple features and description information corresponding to each of the multiple features. Then, after determining the second media content from the multiple candidate media contents, the client can input the prompt information corresponding to the target video in the prompt information area, and generate the target video based on the prompt information, the first media content and the second media content.

[0092] In some other possible embodiments, the client can generate multiple candidate media contents based on the first media content, multiple features and description information corresponding to each of the multiple features. Then, after determining the second media content from the multiple candidate media contents, the client can input the prompt information corresponding to the target video in the prompt information area, and generate the target video based on the prompt information, multiple features and description information corresponding to each of the multiple features, the first media content and the second media content.

[0093] In this embodiment of the application, before generating the target video based at least on the first media content and the second media content, a sequence of images that conforms to the video logic can be generated to ensure the logical accuracy of the final target video.

[0094] In this embodiment, in response to a video logic creation instruction for the first media content and the second media content, a sequence of key node images, including the main object, is displayed, wherein the temporal sequence of the key node images is consistent with the temporal sequence of the target video. In response to the video creation instruction for the first media content, the second media content, and the key node image sequence, the target video is generated and played.

[0095] Taking the aforementioned prompt "A girl is walking on the street when it suddenly starts raining. She takes an umbrella out of her bag and walks in the rain with it" as an example, after preparing the primary media content, multiple secondary media content, multiple features and their descriptions, and the target video prompt on the video creation page, in response to the video logic creation instructions, multiple images can be displayed. These images are a sequence of images ordered according to key nodes, and each image includes the main subject. Based on the prompt "A girl is walking on the street when it suddenly starts raining. She takes an umbrella out of her bag and walks in the rain with it," the system generates the first image corresponding to "A girl is walking on the street," the second image corresponding to "A girl takes an umbrella out of her bag," and the third image corresponding to "Walking in the rain with an umbrella." This demonstrates that the logic of the AI ​​in generating the target video is accurate.

[0096] If the generated keyframe image sequence is, in order, the first image is "a girl taking an umbrella out of her bag," the second image is "a girl walking on the street," and the third image is "walking in the rain with an umbrella," the image logic is clearly flawed, and therefore the logic of the final generated target video will also be incorrect. Therefore, if the displayed keyframe image sequence, including the main object, does not conform to logic, you can either re-enter the prompt information or regenerate the keyframe image sequence based on the prompt information.

[0097] Optionally, provided that the temporal sequence of key node images is determined and the logic is correct, the target video is generated and played in response to video creation instructions for the first media content, the second media content, the key node image sequence, prompts, multiple features, and descriptions of each feature.

[0098] In this way, with the provision of information from multiple sources, a target video with a consistent subject and logical consistency can be generated.

[0099] In some possible embodiments, the content creation page includes a subject association area. This subject association area is used to display association information between multiple subject objects in response to input operations within the area when the first media content includes multiple subject objects. Optionally, if the multiple subject objects are two human figures, the association information can include their positional relationship, height / weight ratio, and intimacy (e.g., lovers, friends). This association information between multiple subject objects is used to instruct the generation of a target video containing the subject objects, making the relationships between the multiple subject objects presented in the target video more consistent with the association information.

[0100] In summary, users only need to upload one main reference object, and the provided supplementary viewpoints lower the barrier to entry. By providing multiple viewpoints and state images for the main object, it ensures a high degree of consistency in the appearance and features of the main object when implementing different shots, actions, and scenes. This avoids problems such as deformation, deviation, and style drift of the main object's features due to insufficient reference, thus making the video more professional.

[0101] Secondly, by breaking down and controlling the various features of the main object during the creative process, artificial intelligence can focus more on the features when generating images and target videos from other perspectives, reducing the possibility of flaws in the features of the final subject and thus preventing continuity errors.

[0102] Furthermore, by using prompts corresponding to the second media content and the target video, the generated media content can be made more logical, thereby reducing the user's entry barrier and ensuring that the generated content better meets the user's needs.

[0103] Figure 9 A block diagram of a multimedia creation page interaction device is shown according to an exemplary embodiment. It has the function of implementing the data processing method in the above-described method embodiments; the function can be implemented in hardware or by hardware executing corresponding software. (Refer to...) Figure 9 The device includes a first display module 901, a second display module 902, and a third display module 903. The first display module 901 is configured to display first media content containing the main object on the media creation page; The second display module 902 is configured to execute a content generation instruction in response to the first media content and display at least one second media content; the second media content includes a subject object; the subject object in the first media content and the subject object in the second media content have at least different perspectives; The third display module 903 is configured to execute video creation instructions in response to the first media content and the second media content, and generate a target video; the target video includes a main object.

[0104] In some possible embodiments, the second display module is configured to perform: In response to a content generation instruction for the primary media content, generate multiple candidate media content; In response to a selection instruction for multiple candidate media content, at least one second media content is displayed; the at least one second media content is a part of the multiple candidate media content.

[0105] In some possible embodiments, the media creation page includes a feature display area; a second display module is configured to perform: In response to a feature decomposition instruction for the main object in the first media content, multiple features corresponding to the main object and descriptive information for each feature are displayed in the feature display area. In response to a content generation instruction, multiple candidate media contents are displayed based on a first media content, multiple features, and descriptive information corresponding to each of the multiple features.

[0106] In some possible embodiments, the media creation page includes a delete control and a feature addition control for each feature; The device also includes an information adding module, configured to perform: In response to a feature addition command triggered by a feature addition control, a feature addition area is displayed on the media creation page; the feature addition area may include an image addition area or a text addition area. When an image with features to be added is added to the image addition area, or text with features to be added is added to the text addition area, in response to the feature positioning operation for the first media content, the feature to be added and its corresponding descriptive information are displayed in the feature display area.

[0107] In some possible embodiments, the third display module is configured to perform: In response to video creation instructions for first media content, second media content, multiple features, and descriptive information corresponding to each of the multiple features, a target video is generated.

[0108] In some possible embodiments, the apparatus further includes an identifier adding module configured to perform: In response to an operation that adds an identifier to a target feature among multiple features, a preset identifier is displayed in a preset area corresponding to the target feature; the preset identifier includes a fixed feature identifier; Among them, the target feature carrying a fixed feature identifier indicates that the main object in the target video always includes the target feature.

[0109] In some possible embodiments, the media creation page includes a prompt information area corresponding to the prompt information; the device also includes a prompt information input module configured to perform: In response to input operations on the prompts in the input area, the prompts corresponding to the target video are displayed in the prompt area; the prompts corresponding to the target video are used to guide the generation logic of the target video.

[0110] In some possible embodiments, the third display module is configured to perform: In response to video logic creation instructions for primary and secondary media content, a sequence of key node images, including the main object, is displayed; the temporal sequence of the key node images is consistent with the temporal sequence of the target video. In response to video creation instructions for primary media content, secondary media content, and key node image sequences, generate the target video.

[0111] In some possible embodiments, the content creation page includes a subject association area; the subject association area is used to display association information between multiple subject objects in response to an input operation in the subject association area when the first media content includes multiple subject objects; the association information between multiple subject objects is used to indicate the generation of a target video containing subject objects.

[0112] In some possible embodiments, the subject object includes at least one of a person, an animal, and an object; The first media content is media content that has undergone background processing; the first media content includes a first video or a first image, and the second media content is a second video or a second image; the second video is a video clip from the first video; The perspectives or states of any two subjects in the primary media content and the secondary media content are different; states include actions and expressions.

[0113] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0114] Figure 10 This is a block diagram illustrating a page interaction device 3000 for multimedia authoring according to an exemplary embodiment. For example, device 3000 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0115] Reference Figure 10 The device 3000 may include one or more of the following components: a processing component 3002, a memory 3004, a power component 3006, a multimedia component 3008, an audio component 3010, an input / output (I / O) interface 3012, a sensor component 3014, and a communication component 3016.

[0116] Processing component 3002 typically controls the overall operation of device 3000, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 3002 may include one or more processors 3020 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 3002 may include one or more modules to facilitate interaction between processing component 3002 and other components. For example, processing component 3002 may include a multimedia module to facilitate interaction between multimedia component 3008 and processing component 3002.

[0117] Memory 3004 is configured to store various image types of data to support operation of device 3000. Examples of this data include instructions for any application or method operating on device 3000, contact data, phonebook data, messages, pictures, videos, etc. Memory 3004 can be implemented by any image type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0118] Power supply component 3006 provides power to various components of device 3000. Power supply component 3006 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 3000.

[0119] Multimedia component 3008 includes a screen that provides an output interface between the device 3000 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 3008 includes a front-facing camera and / or a rear-facing camera. When the device 3000 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0120] Audio component 3010 is configured to output and / or input audio signals. For example, audio component 3010 includes a microphone (MIC) configured to receive external audio signals when device 3000 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 3004 or transmitted via communication component 3016. In some embodiments, audio component 3010 also includes a speaker for outputting audio signals.

[0121] I / O interface 3012 provides an interface between processing component 3002 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, start buttons, and lock buttons.

[0122] Sensor assembly 3014 includes one or more sensors for providing status assessments of various aspects of device 3000. For example, sensor assembly 3014 may detect the on / off state of device 3000, the relative positioning of components such as the display and keypad of device 3000, changes in the position of device 3000 or a component of device 3000, the presence or absence of user contact with device 3000, the orientation or acceleration / deceleration of device 3000, and temperature changes of device 3000. Sensor assembly 3014 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 3014 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 3014 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0123] Communication component 3016 is configured to facilitate wired or wireless communication between device 3000 and other devices. Device 3000 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 3G, 4G, or 5G), or combinations thereof. In one exemplary embodiment, communication component 3016 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 3016 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0124] In an exemplary embodiment, the apparatus 3000 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0125] Embodiments of the present invention also provide a computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a multimedia creation page interaction method. The at least one instruction or the at least one program is loaded and executed by the processor to implement the multimedia creation page interaction method provided in the above method embodiments.

[0126] In an exemplary embodiment, a storage medium including instructions is also provided, such as a memory 3004 including instructions, which can be executed by a processor 3020 of the device 3000 to perform the above-described method. Optionally, the storage medium may be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device.

[0127] Embodiments of the present invention also provide a computer-readable storage medium that, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform the method of any one of the first or second aspects of the embodiments of the present disclosure.

[0128] Embodiments of the present invention also provide a computer program product comprising a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from the readable storage medium and executes the computer program, causing the computer device to perform the method of any one of the first or second aspects of the embodiments of the present disclosure.

[0129] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, interactive page layout and parallel processing for multimedia authoring are possible or may be advantageous.

[0130] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0131] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0132] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A multimedia creation page interaction method, characterized in that, Applied to the first client, including: The media creation page displays the first media content containing the main subject. In response to a content generation instruction for the first media content, at least one second media content is displayed; the second media content includes the subject object; the subject object in the first media content and the subject object in the second media content have at least different perspectives; In response to video creation instructions for the first media content and the second media content, a target video is generated; the target video includes the main object.

2. The multimedia creation page interaction method according to claim 1, characterized in that, The step of displaying at least one second media content in response to a content generation instruction for the first media content includes: In response to the content generation instruction for the first media content, multiple candidate media contents are generated; In response to a selection instruction for the plurality of candidate media content, at least one second media content is displayed; the at least one second media content is a portion of the plurality of candidate media content.

3. The multimedia creation page interaction method according to claim 2, characterized in that, The media creation page includes a feature display area; In response to the content generation instruction for the first media content, a plurality of candidate media contents are generated, including: In response to a feature decomposition instruction for a main object in the first media content, multiple features corresponding to the main object and descriptive information corresponding to each of the multiple features are displayed in the feature display area; In response to the content generation instruction, multiple candidate media contents generated based on the first media content, the plurality of features, and the description information corresponding to each of the plurality of features are displayed.

4. The multimedia creation page interaction method according to claim 3, characterized in that, The media creation page includes a delete control and a feature addition control for each feature; Before displaying multiple candidate media contents generated based on the first media content, the plurality of features, and description information corresponding to each of the plurality of features on the media creation page in response to the content generation instruction, the method further includes: In response to a feature addition instruction triggered by the feature addition control, a feature addition area is displayed on the media creation page; the feature addition area includes an image addition area or a text addition area. When an image with the feature to be added is added to the image addition area, or text with the feature to be added is added to the text addition area, in response to the feature positioning operation for the first media content, the feature to be added and the corresponding descriptive information are displayed in the feature display area.

5. The multimedia creation page interaction method according to claim 3 or 4, characterized in that, The step of generating a target video in response to video creation instructions for the first media content and the second media content includes: The target video is generated in response to a video creation instruction for the first media content, the second media content, the plurality of features, and the descriptive information corresponding to each of the plurality of features.

6. The multimedia creation page interaction method according to claim 5, characterized in that, After displaying multiple features corresponding to the main object and descriptive information corresponding to each of the multiple features in the feature display area, the method further includes: In response to an operation of adding an identifier to a target feature among the plurality of features, a preset identifier is displayed in a preset area corresponding to the target feature; the preset identifier includes a fixed feature identifier; The target feature carrying the fixed feature identifier indicates that the main object in the target video always includes the target feature.

7. The multimedia creation page interaction method according to any one of claims 1-4 and 6, characterized in that, The media creation page includes a prompt information area corresponding to the prompt information; The method further includes: In response to an input operation for the prompt information in the input area, the prompt information corresponding to the target video is displayed in the prompt information area; the prompt information corresponding to the target video is used to guide the generation logic of the target video.

8. The multimedia creation page interaction method according to claim 7, characterized in that, The step of generating a target video in response to video creation instructions for the first media content and the second media content includes: In response to the video logic creation instructions for the first media content and the second media content, a sequence of key node images including the main object is displayed; the temporal sequence of the key node images is consistent with the temporal sequence of the target video. The target video is generated in response to a video creation instruction for the first media content, the second media content, and the key node image sequence.

9. The multimedia creation page interaction method according to claim 1, characterized in that, The content creation page includes a subject association area; the subject association area is used to display association information between the multiple subject objects in response to an input operation in the subject association area when the first media content includes multiple subject objects; the association information between the multiple subject objects is used to indicate the generation of a target video containing the subject objects.

10. The multimedia creation page interaction method according to any one of claims 1-4 and 6, characterized in that, The main object includes at least one of a person, an animal, or an item; The first media content is media content after background processing; the first media content includes a first video or a first image, and the second media content is a second video or a second image; the second video is a video segment from the first video. The perspectives or states of any two main objects in the first media content and the second media content are different; the states include actions and expressions.

11. A multimedia creation page interactive device, characterized in that, include: The first display module is configured to display the first media content containing the main object on the media creation page; The second display module is configured to execute a content generation instruction in response to the first media content and display at least one second media content; the second media content includes the subject object; the subject object in the first media content and the subject object in the second media content have at least different perspectives; The third display module is configured to execute a target video in response to video creation instructions for the first media content and the second media content; The target video includes the main object.

12. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the multimedia creation page interaction method as described in any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the multimedia authoring page interaction method as described in any one of claims 1 to 10.

14. A computer program product, characterized in that, The computer program product includes a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from and executes the computer program, causing the device to perform a multimedia authoring page interaction method as described in any one of claims 1 to 10.