Multimedia resource generation method and device, electronic equipment and storage medium
By displaying the description components of video features and logos in the multimedia generation interface, the problem of insufficient coordination between text information and video is solved, and the efficiency of multimedia resource generation is improved.
Patent Information
- Application Number
- CN202510473806.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, when users generate multimedia resources, the matching degree between text information and video is insufficient, resulting in the generated multimedia resources that cannot meet user creativity, and require multiple inputs and regeneration, which reduces the generation efficiency.
The description component is displayed in the multimedia generation interface, including video features and video identifiers. The use of the prompt information input box and the description component is used to generate target prompt information so that the user can clearly and accurately describe the creativity, thereby improving the generation efficiency.
By matching the description components of video features and logos, users can express their creativity more accurately, reduce multiple inputs and regenerate tasks, and improve the efficiency of multimedia resource generation.
Smart Images

Figure CN120264087A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of multimedia technologies, and particularly to a method, apparatus, electronic device, and storage medium for generating multimedia resources. Background Art
[0002] With the development of multimedia technologies, users can use artificial intelligence models to process videos to generate new multimedia resources, thereby completing secondary creation of the videos.
[0003] To facilitate users' secondary creation, a terminal can provide a multimedia generation interface, where users can input videos to be processed. Users can also input text information in the multimedia generation interface to describe their creative intentions (hereinafter referred to as creativity) for the input videos. The terminal provides the videos and text information input by the users to an artificial intelligence model, and the artificial intelligence model uses the text information as prompt information for processing the videos, and processes the input videos according to the prompt information to generate new multimedia resources.
[0004] However, since the text information input by users does not match well with the videos to be processed input by users, the generated multimedia resources may not meet users' creativity, and users need to re-enter the text information and provide it to the artificial intelligence model, and the newly generated multimedia resources may not necessarily meet users' creativity either, thus reducing the generation efficiency of multimedia resources. Summary of the Invention
[0005] The present disclosure provides a method, apparatus, electronic device, and storage medium for generating multimedia resources to at least solve the problem of low generation efficiency of multimedia resources in related technologies. The technical solutions of the present disclosure are as follows:
[0006] According to one aspect of the embodiments of the present disclosure, there is provided a method for generating multimedia resources, including:
[0007] Display a multimedia generation interface, where the multimedia generation interface includes a prompt information input box for inputting prompt information, and the prompt information indicates an editing method for a video;
[0008] Based on a first video obtained on the multimedia generation interface, display a first description component in the prompt information input box, where the first description component includes video features and a video identifier of the first video;
[0009] Display an obtained input statement in the prompt information input box, and the input statement and the first description component in the prompt information input box form target prompt information;
[0010] In response to a multimedia resource generation operation, display a target multimedia resource generated based on the target prompt information.
[0011] Optionally, after displaying the multimedia generation interface, the method further includes:
[0012] In response to a multimedia resource input operation in the prompt information input box, display the input first video in a first display window in the multimedia generation interface.
[0013] Optionally, the prompt information input box includes a first entry component for inputting a video into the first display window.
[0014] The step of displaying the input first video in the first display window in the multimedia generation interface in response to a multimedia resource input operation in the prompt information input box includes:
[0015] In response to a selection operation on the first entry component, display the input first video in the first display window in the multimedia generation interface.
[0016] Optionally, the step of displaying the input first video in the first display window in the multimedia generation interface in response to a multimedia resource input operation in the prompt information input box includes:
[0017] In response to an invocation operation for the first entry component in the prompt information input box, display the first entry component in the prompt information input box, where the first entry component is used to input a video into the first display window.
[0018] In response to a selection operation on the first entry component, display the input first video in the first display window in the multimedia generation interface.
[0019] Optionally, the multimedia generation interface further includes a second entry component for inputting a video into the first display window, and the second entry component is located outside the prompt information input box.
[0020] Optionally, the multimedia generation interface further includes a candidate multimedia area, and a third entry component corresponding to a candidate video is displayed in the candidate multimedia area, where the third entry component is used to input the corresponding candidate video into the first display window.
[0021] After displaying the multimedia generation interface, the method further includes:
[0022] In response to a selection operation on the third entry component, display the candidate video in the first display window in the multimedia generation interface, and the candidate video in the first display window is the first video.
[0023] Optionally, the multimedia generation interface further includes a fourth entry component and / or a fifth entry component. The fourth entry component is located in the prompt information input box, and the fifth entry component is located outside the prompt information input box. Both the fourth entry component and the fifth entry component are used to input pictures into the second display window in the multimedia generation interface, and the pictures in the second display window are used to generate the target prompt information.
[0024] Optionally, after displaying the obtained input statement in the prompt information input box, the method further includes:
[0025] In response to a moving operation on the first description component, move the first description component from the first position in the prompt information input box to the second position in the input statement.
[0026] Optionally, after displaying the first description component in the prompt information input box, the method further includes:
[0027] In response to obtaining a second video on the multimedia generation interface, update the video feature of the first video in the first description component to the video feature of the second video.
[0028] Optionally, after displaying the first description component in the prompt information input box, the method further includes:
[0029] In response to a deletion operation on the first description component in the prompt information input box, display the deleted first description component in the candidate component area of the multimedia generation interface;
[0030] In response to a selection operation on the first description component in the candidate component area, add the first description component back to the prompt information input box.
[0031] Optionally, after displaying the first description component in the prompt information input box, the method further includes:
[0032] In response to a component addition operation in the prompt information input box, display a candidate component panel, and the candidate component panel includes the first description component;
[0033] In response to a selection operation on the first description component in the candidate component panel, add the first description component to the prompt information input box.
[0034] Optionally, before the first description component is displayed in the prompt information input box based on the first video obtained on the multimedia generation interface, the method further includes:
[0035] Displaying a prompt template in the multimedia generation interface, where the prompt template includes a reference statement for generating the prompt information, and the input statement matches the syntax structure of the reference statement.
[0036] Optionally, the multimedia generation interface includes a plurality of editing components, and different editing components indicate different editing methods for the video; the displaying of the prompt template in the multimedia generation interface includes:
[0037] In response to a selection operation on any one of the editing components, displaying in the multimedia generation interface the prompt template related to the editing method indicated by the editing component.
[0038] Optionally, after the prompt template is displayed in the multimedia generation interface, the method further includes:
[0039] Updating the reference statement in the prompt template based on the first object in the first video obtained on the multimedia generation interface.
[0040] Optionally, after the prompt template is displayed in the multimedia generation interface, the method further includes:
[0041] Updating the reference statement in the prompt template based on the first object in the first video and the second object in the target picture, where the target picture can be obtained on the multimedia generation interface.
[0042] Optionally, the displaying of the prompt template in the multimedia generation interface includes:
[0043] Displaying the prompt template in the prompt information input box in the multimedia generation interface.
[0044] Optionally, the prompt template includes a first entry component, and the first entry component is used to input a video into a first display window in the multimedia generation interface.
[0045] According to another aspect of the embodiments of the present disclosure, there is provided a multimedia resource device, including:
[0046] A first display unit configured to execute displaying a multimedia generation interface, where the multimedia generation interface includes a prompt information input box for inputting prompt information, and the prompt information indicates an editing method for a video;
[0047] A second display unit, configured to display a first description component in the prompt information input box based on a first video obtained on the multimedia generation interface, where the first description component includes video features and a video identifier of the first video;
[0048] A third display unit, configured to display an obtained input statement in the prompt information input box, and the input statement in the prompt information input box and the first description component form target prompt information;
[0049] A fourth display unit, configured to display a target multimedia resource generated based on the target prompt information in response to a multimedia resource generation operation.
[0050] Optionally, the device further includes:
[0051] A fifth display unit, configured to display an input first video in a first display window on the multimedia generation interface in response to a multimedia resource input operation in the prompt information input box.
[0052] Optionally, the prompt information input box includes a first entry component for inputting a video into the first display window; the fifth display unit is configured to display an input first video in the first display window on the multimedia generation interface in response to a selection operation on the first entry component.
[0053] Optionally, the fifth display unit is configured to perform:
[0054] In response to an activation operation on the first entry component in the prompt information input box, display the first entry component in the prompt information input box, where the first entry component is used to input a video into the first display window;
[0055] In response to a selection operation on the first entry component, display the input first video in the first display window on the multimedia generation interface.
[0056] Optionally, the multimedia generation interface further includes a second entry component for inputting a video into the first display window, and the second entry component is located outside the prompt information input box.
[0057] Optionally, the multimedia generation interface further includes a candidate multimedia area in which a third entry component corresponding to a candidate video is displayed, and the third entry component is used to input the corresponding candidate video into the first display window; the device further includes:
[0058] A fifth display unit, configured to execute a selection operation in response to the third entry component, and display the candidate video in the first display window in the multimedia generation interface, where the candidate video in the first display window is the first video.
[0059] Optionally, the multimedia generation interface further includes a fourth entry component and / or a fifth entry component. The fourth entry component is located in the prompt information input box, and the fifth entry component is located outside the prompt information input box. Both the fourth entry component and the fifth entry component are used to input pictures into the second display window in the multimedia generation interface, and the pictures in the second display window are used to generate the target prompt information.
[0060] Optionally, the device further includes:
[0061] A moving unit, configured to execute a moving operation in response to the first description component, and move the first description component from the first position in the prompt information input box to the second position in the input statement.
[0062] Optionally, the device further includes:
[0063] A first update unit, configured to execute an update operation in response to obtaining a second video on the multimedia generation interface, and update the video feature of the first video in the first description component to the video feature of the second video.
[0064] Optionally, the device further includes:
[0065] A deletion unit, configured to execute a deletion operation in response to the first description component in the prompt information input box, and display the deleted first description component in the candidate component area of the multimedia generation interface;
[0066] A first addition unit, configured to execute an addition operation in response to the selection operation of the first description component in the candidate component area, and add the first description component back to the prompt information input box.
[0067] Optionally, the device further includes:
[0068] A sixth display unit, configured to execute a component addition operation in response to the prompt information input box, and display a candidate component panel, where the candidate component panel includes the first description component;
[0069] A second addition unit, configured to execute an addition operation in response to the selection operation of the first description component in the candidate component panel, and add the first description component to the prompt information input box.
[0070] Optionally, the device further includes:
[0071] A seventh display unit configured to execute displaying a prompt template in the multimedia generation interface, the prompt template including a reference statement for generating the prompt information, and the input statement matching the syntax structure of the reference statement.
[0072] Optionally, the multimedia generation interface includes a plurality of editing components, and different editing components indicate different editing methods for the video;
[0073] The seventh display unit is configured to execute, in response to a selection operation on any one of the editing components, displaying the prompt template related to the editing method indicated by the editing component in the multimedia generation interface.
[0074] Optionally, the apparatus further includes:
[0075] A second update unit configured to execute updating the reference statement in the prompt template based on a first object in a first video obtained on the multimedia generation interface.
[0076] Optionally, the apparatus further includes:
[0077] A third update unit configured to execute updating the reference statement in the prompt template based on a first object in a first video and a second object in a target picture, and the target picture can be obtained on the multimedia generation interface.
[0078] Optionally, the seventh display unit is configured to execute displaying the prompt template in a prompt information input box in the multimedia generation interface.
[0079] Optionally, the prompt template includes a first entry component for inputting a video into a first display window in the multimedia generation interface.
[0080] According to another aspect of an embodiment of the present disclosure, there is provided an electronic device, including:
[0081] One or more processors;
[0082] One or more memories for storing executable instructions of the one or more processors;
[0083] Wherein, the one or more processors are configured to execute the multimedia resource generation method in any possible implementation manner of the above-mentioned one aspect.
[0084] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium. When at least one instruction in the computer-readable storage medium is executed by one or more processors of an electronic device, the electronic device is enabled to execute the multimedia resource generation method in any possible implementation manner of the above aspect.
[0085] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product including one or more instructions, which can be executed by one or more processors of an electronic device, enabling the electronic device to execute the multimedia resource generation method in any possible implementation manner of the above aspect.
[0086] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0087] By displaying description components in the prompt information input box based on the video input by the user on the multimedia generation interface, since the description components include the video features and video identifier of the video, the description components can match the video, so that the user can use the description components when inputting prompt information, mix the input statement and the description components to discharge the target prompt information, making the target prompt information more clearly and accurately describe the user's creativity, and further enabling the multimedia resources generated based on the target prompt information to meet the user's creativity, thus eliminating the need for the user to input prompt information multiple times on the multimedia generation interface and submit the task of regenerating multimedia resources multiple times, and further improving the generation efficiency of multimedia resources.
[0088] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.
[0090] Figure 1 is a schematic diagram of an implementation environment of a multimedia resource generation method shown according to an exemplary embodiment;
[0091] Figure 2 is a flowchart of a multimedia resource generation method shown according to an exemplary embodiment;
[0092] Figure 3 is a schematic diagram of a multimedia generation interface in a modified manner shown according to an exemplary embodiment;
[0093] Figure 4Schematic diagram of a multimedia generation interface in an addition mode shown according to an exemplary embodiment;
[0094] Figure 5 Schematic diagram of a multimedia generation interface in a deletion mode shown according to an exemplary embodiment;
[0095] Figure 6 Schematic diagram of a multimedia generation interface in another modification mode shown according to an exemplary embodiment;
[0096] Figure 7 Flowchart of another multimedia resource generation method shown according to an exemplary embodiment;
[0097] Figure 8 Schematic diagram of input multimedia resources in an addition mode shown according to an exemplary embodiment;
[0098] Figure 9 Schematic diagram of input multimedia resources in a modification mode shown according to an exemplary embodiment;
[0099] Figure 10 Schematic diagram of input multimedia resources in a deletion mode shown according to an exemplary embodiment;
[0100] Figure 11 Schematic diagram of input of target prompt information in a deletion mode shown according to an exemplary embodiment;
[0101] Figure 12 Schematic diagram of input of target prompt information in an addition mode shown according to an exemplary embodiment;
[0102] Figure 13 Schematic diagram of input of target prompt information in a modification mode shown according to an exemplary embodiment;
[0103] Figure 14 Flowchart of another multimedia resource generation method shown according to an exemplary embodiment;
[0104] Figure 15 Schematic diagram of display of a description component shown according to an exemplary embodiment;
[0105] Figure 16 Schematic diagram of mixed arrangement of target prompt information shown according to an exemplary embodiment;
[0106] Figure 17 Schematic diagram of another mixed arrangement of target prompt information shown according to an exemplary embodiment;
[0107] Figure 18It is a schematic diagram showing a component deletion description according to an exemplary embodiment;
[0108] Figure 19 It is a schematic diagram showing an update of a reference statement according to an exemplary embodiment;
[0109] Figure 20 It is a flowchart of another method for generating multimedia resources according to an exemplary embodiment;
[0110] Figure 21 It is a structural block diagram of a multimedia resource generation device according to an exemplary embodiment;
[0111] Figure 22 It is a structural block diagram of another multimedia resource generation device according to an exemplary embodiment;
[0112] Figure 23 It is a structural block diagram of an electronic device according to an exemplary embodiment. Detailed implementation
[0113] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0114] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0115] The user information involved in the present disclosure may be information authorized by the user or fully authorized by all parties.
[0116] In some embodiments, the meaning of A and / or B includes: three cases: A and B, A, and B.
[0117] Next, some terms related to the present disclosure will be introduced.
[0118] Multimedia application: It refers to an application program that supports the generation of multimedia resources. Multimedia applications include, but are not limited to: at least one of an application specifically for generating multimedia resources, a short video application, an audio and video application, a shopping application, a food delivery application, a travel application, a game application, or a social application.
[0119] Multimedia resource: It is composed of at least one multimedia element. The multimedia resource includes but is not limited to at least one of video, image, or text. A multimedia element is an element that constitutes a multimedia resource. For example, a video frame in a video, an element in a video frame, an element in an image, or an element in text, etc. The text involved in the present disclosure can be plain text that only contains text information, or rich text that contains various multimedia elements such as graphic and text information, components, etc. Here, the embodiments of the present disclosure do not limit the type of text.
[0120] Figure 1 It is a schematic diagram of an implementation environment of a multimedia resource generation method shown according to an exemplary embodiment. Refer to Figure 1 This implementation environment includes at least one terminal 101 and a server 102. The terminal 101 is directly or indirectly connected to the server 102 through a wired network or a wireless network. The embodiments of the present disclosure do not limit the connection method between the terminal 101 and the server 102.
[0121] The terminal 101 is a user device. The terminal 101 can be at least one of a smart phone, a tablet computer, an e-book reader, a music player, a personal computer, a laptop portable computer, a desktop computer, and a wearable device. Here, the embodiments of the present disclosure do not limit the device type of the terminal 101. Figure 1 3 terminals 101 are shown, but the number of terminals 101 is not limited to 3, and can also be less than 3 or more than 3. Here, the embodiments of the present disclosure do not limit the number of terminals 101 in this implementation environment.
[0122] A multimedia application is installed and run on the terminal 101. The user logs in to the multimedia application on the terminal 101 through a certain user account. The user inputs the multimedia resource to be processed and the prompt information in the multimedia generation interface provided by the multimedia application, triggering the terminal 101 to send the input multimedia resource and the prompt information to the server 102. The server 102 uses an artificial intelligence model to process the multimedia resource and the prompt information input by the user to generate a new multimedia resource, and returns the generated multimedia resource to the terminal 101. The terminal 101 displays the generated multimedia resource, thereby completing the secondary creation of the input multimedia resource. Among them, the prompt information is used to describe the user's creativity for the input multimedia resource to prompt the artificial intelligence model to perform what processing on the input multimedia resource, or it can also prompt what forms of expression the multimedia elements in the multimedia resource to be generated have, so that the artificial intelligence model processes the input multimedia resource based on the prompt information to generate a new multimedia resource.
[0123] Server 102 is used to provide background services for the multimedia applications running on terminal 101. Server 102 includes at least one of a single server, multiple servers, a cloud computing platform, or a virtualization center. Optionally, server 102 undertakes the main computing work and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work and terminal 101 undertakes the main computing work; or, a distributed computing architecture is adopted between server 102 and terminal 101 for collaborative computing.
[0124] The above is described by taking the deployment of the artificial intelligence model on server 102 as an example. In some other embodiments, the artificial intelligence model is deployed on terminal 101. Terminal 101 uses the artificial intelligence model to process the multimedia resources and prompt information input by the user to generate new multimedia resources. In this case, server 102 may not be included in this implementation environment.
[0125] Next, based on Figure 1 The shown is the implementation environment. In combination with Figure 2 The process of the multimedia resource generation method provided by the present disclosure will be introduced. Among them, Figure 2 is a flowchart of a multimedia resource generation method shown according to an exemplary embodiment. Refer to Figure 2 , this multimedia resource generation method is applied to a terminal, and this terminal may be terminal 101 in the above implementation environment. This multimedia resource generation method includes the following steps.
[0126] In step 201, the terminal displays a multimedia generation interface, and a prompt template is displayed in the multimedia generation interface. The prompt template includes reference statements for generating prompt information, and the prompt information indicates the editing method of the multimedia resources.
[0127] Among them, the multimedia generation interface is a human-computer interaction interface provided by the multimedia application, which is used to provide the function of secondary generation of multimedia resources. The function of secondary generation of multimedia resources refers to the function of generating new multimedia resources based on multimedia resources.
[0128] The function of secondary generation of multimedia resources includes multimodal editing. Multimodal editing refers to the way of generating multimedia resources based on multimedia resources of multiple resource types. For example, the way of generating multimedia resources based on multimedia resources of at least two resource types among video, image, and text.
[0129] Exemplarily, such as Figures 3 to 5Each of the multimedia generation interfaces 300 shown, the multimedia generation interface 300 includes a multimodal editing sub-interface 310, and the multimodal editing sub-interface 310 provides a multimodal editing function. The multimodal editing sub-interface 310 includes a multimodal editing component 31, and the multimodal editing component 31 indicates multimodal editing. The multimodal editing component 31 is in a selected state to indicate the selection of using multimodal editing to generate multimedia resources.
[0130] Multimodal editing supports multiple editing methods for multimedia resources. The multiple editing methods include a modification method, an addition method, and a deletion method. Among them, the modification method includes modifying and / or replacing multimedia elements in the multimedia resource, the addition method includes adding multimedia elements to the multimedia resource, and the deletion method includes deleting multimedia elements in the multimedia resource.
[0131] The multimedia generation interface includes multiple editing components, and different editing components indicate different editing methods for multimedia resources. Exemplarily, as Figures 3 to 5 shown in each of the multimedia generation interfaces 300, the multimodal editing sub-interface 310 includes a modification editing component 311, an addition editing component 312, and a deletion editing component 313. Among them, the modification editing component 311 indicates the modification method, the addition editing component 312 indicates the addition method, and the deletion editing component 313 indicates the deletion method.
[0132] The editing method currently selected and used among the multiple editing methods is called the current editing method. The editing component that indicates the current editing method among the multiple editing components is in a selected state. That is, in the case where any editing component is in a selected state, the editing method indicated by this editing component is the current editing method. As Figure 3 shown, the modification editing component 311 is in a selected state, and the modification method is the current editing method. As Figure 4 shown, the addition editing component 312 is in a selected state, and the addition method is the current editing method. As Figure 5 shown, the deletion editing component 313 is in a selected state, and the deletion method is the current editing method. Among them, the selected state indicates performing multimodal editing using the editing method indicated by the editing component (i.e., the current editing method).
[0133] The current editing method is the default editing method or the editing method selected by the user. The default editing method is any one of the multiple editing methods. For example, when the multimodal editing sub-interface 310 is initially displayed, the current editing method is the default editing method, and the default editing method is any one of the multiple editing methods. The user can perform a selection operation on any editing component to select the editing method indicated by this editing component as the current editing method.
[0134] In the present disclosure, the selection operation on any component includes, but is not limited to, a click operation on the component, an operation of selecting the component by voice instruction or other types of instructions.
[0135] In some embodiments, the multimedia resource secondary generation function further includes unimodal editing. Unimodal editing refers to a way of generating multimedia resources based on multimedia resources of one resource type. For example, a way of generating multimedia resources based on one of text, images, or videos. Exemplarily, as Figures 3 to 5 shown in each multimedia generation interface 300, the multimodal editing sub-interface 310 further includes a unimodal editing component 32 and a unimodal editing component 33. Among them, the unimodal editing component 32 indicates unimodal editing for generating a video based on text, and the unimodal editing component 33 indicates unimodal editing for generating a video based on pictures. The unimodal editing components 32 and 33 are respectively associated with unimodal editing sub-interfaces, and the modal editing sub-interfaces are used to implement the unimodal editing functions indicated by the associated unimodal editing components. The user triggers the terminal to jump to the associated unimodal editing sub-interface by selecting any unimodal editing component.
[0136] The prompt information indicates the editing method of the multimedia resource. For example, it indicates which multimedia elements in the multimedia resource to modify and how to modify the multimedia elements, or indicates which multimedia elements to add to the multimedia resource, or which elements to delete from the multimedia resource, etc. The reference statement is a statement provided for the user to refer to in the prompt information. For example, Figure 3 the reference statement S1 shown is "Use the reference diagram to replace the specified content in the original video", Figure 4 the reference statement S2 shown "Based on the original video content, in a natural and attractive way, integrate the content of the reference diagram into the video", Figure 5 the reference statement S3 shown "Delete the specified picture content in the original video (it is better to specify the subject)" and so on. These reference statements can clearly describe the editing method of the multimedia resource (original video or reference diagram), and how the reference diagram and the original video are substituted and interacted with each other. Subsequently, if the user inputs a statement whose syntax structure matches the reference statement when inputting the prompt information, it can make the input prompt information (such as the target prompt information below) clearly describe the editing method of the multimedia resource and the interaction method between different multimedia resources, and thus can clearly describe the user's creativity, so that the subsequent multimedia resources generated based on this prompt information can better conform to the user's creativity. There is no need for the user to input the prompt information multiple times on the multimedia generation interface and submit the task of regenerating the multimedia resource multiple times, thereby reducing the number of times the user inputs the prompt information and the number of times of regenerating the multimedia resource, and thus improving the generation efficiency of the multimedia resource.
[0137] The hint template is a template for hint messages. There are multiple hint templates, and each hint template is related to an editing method supported by multimodal editing. The relationship between the hint template and the editing method can be manifested in that the reference statement in the hint template can reflect the editing method. Exemplarily, as Figure 3 shown, the hint template 1 including the reference statement S1 can reflect the modification method, so the hint template 1 is related to the modification method; as Figure 4 shown, the hint template 2 including the reference statement S2 can reflect the addition method, so the hint template 2 is related to the modification method; as Figure 5 shown, the hint template 3 including the reference statement S3 can reflect the deletion method, so the hint template 3 is related to the modification method.
[0138] The above introduced the content of the multimedia generation interface. Next, taking Figures 3 to 5 the illustration of each multimedia generation interface 300 as an example, the display process of the multimedia generation interface will be introduced.
[0139] The user issues an instruction to the terminal to open the multimedia generation interface in the multimedia application. In response to this instruction, the terminal displays Figures 3 to 5 the multimedia generation interface 300 shown in any of the attached drawings. The multimedia generation interface 300 includes a multimodal editing sub-interface 310, and the multimodal editing component 31 in the multimodal editing sub-interface 310 is in a selected state to indicate that the multimodal editing function is currently selected for use. If the user performs a selection operation on any one of the single-modal editing components 32 or 33 in the multimodal editing sub-interface 310, the terminal adjusts to the one associated with the single-modal editing component in response to this selection operation. Among them, the multimodal editing component 31, the single-modal editing component 32, and the single-modal editing component 33 are optional components. In some embodiments, the multimodal editing sub-interface 310 does not include the multimodal editing component 31, the single-modal editing component 32, and the single-modal editing component 33.
[0140] When the terminal initially displays the multimedia generation interface 300 in response to this instruction, the current editing method is the default editing method, and the editing component in the multimodal editing sub-interface 310 indicating the current editing method is in a selected state. For example, assuming the default editing method is the modification method, the terminal displays Figure 3 the multimedia generation interface 300 shown, and the modification editing component 311 is in a selected state. Assuming the default editing method is the addition method, the terminal displays Figure 4 the multimedia generation interface 300 shown, and the addition editing component 312 is in a selected state. Assuming the default editing method is the deletion method, the terminal displays Figure 5 the multimedia generation interface 300 shown, and the deletion editing component 313 is in a selected state.
[0141] The editing component indicating the current editing mode is called the current editing component. Assume that the current editing mode is the first editing mode among multiple editing modes, and the editing component indicating the first editing mode among multiple editing components is the first editing component. The user can perform a selection operation on the second editing component among multiple editing components. In response to this selection operation, the terminal switches the first editing mode to the second editing mode to meet the different creative needs of the user. Among them, the first editing component and the second editing component are different, the second editing mode is the editing mode indicated by the second editing component, and the second editing mode and the first editing mode are different. Changing the first editing component mode to the second editing mode can be expressed as: switching the display state of the first editing component from the selected state to the unselected state, and switching the display state of the second editing component from the unselected state to the selected state, so as to switch the current editing component from the first editing component to the second editing component.
[0142] Take Figure 3 the multimodal editing sub-interface 310 shown in the figure as an example. The current editing component is the modification editing component 311 (i.e., the first editing component), and the current editing mode is the modification mode (i.e., the first editing mode). The user performs a selection operation on the addition editing component 312. As Figure 4 shown in the figure, in response to this selection operation, the terminal switches the display state of the modification editing component 311 from the selected state to the unselected state, and switches the display state of the addition editing component 312 from the unselected state to the selected state. At this time, the addition editing component 312 is the current editing component, and the addition mode (i.e., the second editing mode) is the current editing mode, so as to realize switching the current editing mode from the modification mode to the addition mode. When the current editing mode switches between any two of the modification mode, the addition mode, and the deletion and modification mode, the process is the same as this switching process. It should be noted that when switching from the addition mode or the modification mode to the deletion and modification mode, the terminal also deletes the second display window R2 and the fifth entry component 80 in the multimodal editing sub-interface 310. The second display window R2 and the fifth entry component 80 are introduced in the following text and will not be elaborated here.
[0143] To facilitate the user to better use the multimodal editing function provided by the multimodal editing sub-interface 310, the terminal displays a prompt template in the multimodal editing sub-interface 310. For example, a prompt template related to the current editing mode is displayed in the multimodal editing sub-interface 310. For example, Figures 3 to 5 in the modal editing sub-interface 310 shown in the figure, prompt template 1, prompt template 2, and prompt template 3 are respectively displayed.
[0144] When the user switches the editing mode, the terminal updates the prompt template in the multimedia generation interface. For example, in response to a selection operation on any editing component, the terminal displays, in the multimedia generation interface, a prompt template related to the editing mode indicated by the editing component. Exemplarily, assume that the current editing mode is the first editing mode, and the user performs a selection operation on the second editing component in the multimedia generation interface. In response to the selection operation, the terminal displays, in the multimedia generation interface, a prompt template related to the second editing mode.
[0145] Take Figure 3 as an example. The current editing component is the modification editing component 311, the first editing mode is the modification mode, and the prompt template 1 is displayed in the multimodal editing sub-interface 310. The user performs a selection operation on the addition editing component 312. As Figure 4 shown, in response to this selection operation, the terminal replaces the prompt template 1 in the multimodal editing sub-interface 310 with the prompt template 2.
[0146] Different editing components can trigger prompt templates related to different editing modes, so that the user can trigger prompt templates related to different editing modes by selecting different editing components, to meet the different creative needs of the user and improve the generation efficiency of multimedia resources.
[0147] The multimedia generation interface includes a prompt information input box, which is used to input prompt information (such as the target prompt information involved below). The terminal can display the prompt template in the prompt information input box in the multimedia generation interface. Exemplarily, as Figures 3 to 5 shown in each multimedia generation interface 300, the multimedia generation interface 300 includes a prompt information input box 40. As Figure 3 shown, the prompt template 1 is displayed in the prompt information input box 40. As Figure 4 shown, the prompt template 2 is displayed in the prompt information input box 40. As Figure 5 shown, the prompt template 3 is displayed in the prompt information input box 40.
[0148] Since the prompt information input box is used to input prompt information, and the prompt template is displayed in the prompt information input box, while providing the prompt template for the user, it can also prompt the user to input prompt information in the prompt information input box, so that the user can quickly locate the prompt information input box when inputting prompt information, thereby improving the generation efficiency of multimedia resources.
[0149] The above is an example of displaying the prompt template in the prompt information input box. In some other embodiments, the terminal can also display the prompt information template in an area outside the prompt information input box in the multimedia generation interface. Here, the present disclosure embodiment does not limit the display position of the prompt template.
[0150] The multimedia resources used when generating multimedia resources are referred to as reference multimedia resources. Under multimodal editing, reference multimedia resources of at least one resource type are used, for example, at least one of video and picture. When the reference multimedia resource is a video, the reference multimedia resource can be referred to as the original video, and when the reference multimedia resource is a picture, the reference multimedia resource can be referred to as the reference picture.
[0151] Under the modification method or the addition method, reference multimedia resources of multiple resource types are used. For example, the reference multimedia resources used under these two editing methods include video and picture. Under the deletion method, reference multimedia resources of one resource type are used, such as video or picture.
[0152] The reference multimedia resources used under multimodal editing are respectively the basic multimedia resource and the additional multimedia resource. Among them, the basic multimedia resource is the basis of the multimedia resource to be generated, and a new multimedia resource can be generated based on the basic multimedia resource. The additional multimedia resource is used to provide multimedia elements for the basic multimedia resource, and the multimedia elements in the additional multimedia resource can be added to the basic multimedia resource to generate a new multimedia resource. The resource types of the basic multimedia resource and the additional multimedia resource are different. For example, the basic multimedia resource is a video and the additional multimedia resource is a picture.
[0153] When performing multimodal editing, the reference multimedia resources required in the current editing mode are input by the user to the multimedia generation interface so that the terminal can obtain the reference multimedia resources input by the user from the multimedia generation interface. Taking the current editing mode as the modification method or the addition method as an example, the user inputs the basic multimedia resource and the additional multimedia resource to the multimedia generation interface, and the terminal obtains the basic multimedia resource and the additional multimedia resource input by the user from the multimedia generation interface. Or, taking the current editing mode as the deletion method as an example, the user inputs the basic multimedia resource to the multimedia generation interface, and the terminal obtains the basic multimedia resource input by the user from the multimedia generation interface.
[0154] In step 202, the terminal displays target prompt information in the multimedia generation interface based on the first multimedia resource and the input statement obtained on the multimedia generation interface, and the syntax structure of the input statement matches that of the reference statement.
[0155] Among them, the first multimedia resource is the basic multimedia resource input by the user on the multimedia generation interface. Exemplarily, when the current editing mode is any one of the modification method, the addition method, or the deletion method, the terminal obtains the basic multimedia resource input by the user and uses the obtained basic multimedia resource as the first multimedia resource.
[0156] The input statement is the statement entered by the user in the prompt message input box. The input statement includes first description information, which is the description information of the first multimedia resource and is used to describe the first multimedia resource. The first description information includes the multimedia features and resource identifier of the first multimedia resource. In the embodiments of the present disclosure, the multimedia features of any multimedia resource can be various features that can describe the multimedia resource. For example, the thumbnail of the multimedia resource. In the case where the multimedia resource is a video, the thumbnail of the multimedia resource is the thumbnail of a certain video frame in the video. This video frame can be the first video frame in the video or the highlight frame of the video. The highlight frame refers to a specific video frame in the video that has prominent features or important information and can attract the special attention of the audience. Or, the multimedia features of the multimedia resource can also be other features that can describe the multimedia resource other than the thumbnail. The embodiments of the present disclosure do not limit these other features. In the embodiments of the present disclosure, the resource identifier of any multimedia resource is used to indicate the multimedia resource. Taking the multimedia resource as a video as an example, the resource identifier of the multimedia resource is "video", or the resource identifier of the multimedia resource is the number of the video, such as video 1 or 2. Taking the multimedia resource as a picture as an example, the resource identifier of the multimedia resource is "picture", or the resource identifier of the multimedia resource is the number of the picture, such as Figure 1 or 2.
[0157] The first description information can be carried by a certain interactive component. The interactive component carrying the first description information is the first description component, that is, the first description component includes the first description information and is used to describe the first multimedia resource.
[0158] The matching of the syntax structure between the input statement and the reference statement can be manifested as: the editing method reflected by the input statement is the same as the editing method reflected by the reference statement. For example, the input statement "Delete the kittens that appear in xxx" and the reference statement S3 can both reflect the deletion method, so the syntax structure of this input statement matches the reference statement S3. Among them, "xxx" is the first description information or the description component carrying the first description information.
[0159] The target prompt message is the prompt message entered by the user in the prompt message input box, and it is the prompt message to be used when generating the multimedia resource this time. The target prompt message includes this input statement. Since the reference statement can clearly describe the editing method of the reference multimedia resource and the interaction method between different reference multimedia resources, if the input statement in the target prompt message can match the syntax structure of the reference statement, the target prompt message can clearly describe the editing method of the reference multimedia resource output by the user and the interaction method between different reference multimedia resources, and thus can clearly describe the user's creativity. Subsequently, the multimedia resource generated based on the target prompt message can better conform to the user's creativity and improve the generation efficiency of the multimedia resource.
[0160] After obtaining the first multimedia resource, the terminal displays the first multimedia resource input by the user in the multimedia generation interface. The terminal uploads the first multimedia resource to the server, and the server generates first description information based on the first multimedia resource, and uses the first description component carrying the first description information as a candidate description component and adds it to the candidate component panel associated with the user, so that the candidate component panel includes the first description component.
[0161] After the user finishes inputting the first multimedia resource, the user inputs the target prompt message in the prompt message input box. During the process of inputting the target prompt message, the user can input the first description component in the candidate component panel into the prompt message input box, and input text information in the prompt message input box, so that the text information and the first description component are mixed into an input statement. The user can imitate the way the reference statement expresses the current editing method and input text information that can express the current editing method in the prompt message input box, so that the input statement where the text information is located matches the syntax structure of the reference statement.
[0162] After the user finishes inputting in the prompt message input box, a multimedia resource generation operation is performed. The multimedia resource generation operation instructs to generate a multimedia resource. The terminal responds to this multimedia resource generation operation and uses the input statement in the information prompt message box as the target prompt message.
[0163] Exemplarily, Figures 3 to 5 For the generation component K in the shown multimedia generation interface 300, the user performs a selection operation on the generation component K to implement the multimedia resource generation operation. The terminal responds to this selection operation and uses the input statement in the information prompt message box as the target prompt message. Among them, the generation component K is used to indicate the generation of a multimedia resource.
[0164] In step 203, the terminal responds to the multimedia resource generation operation and displays the second multimedia resource generated based on the target prompt message.
[0165] Among them, the second multimedia resource is the newly generated multimedia resource in this time. The second multimedia resource can be a video or a picture. Here, the embodiments of the present disclosure do not limit the resource type of the multimedia resource. The second multimedia resource is generated based on the target prompt information and the multimedia resources described by each description component in the target prompt information. The description components in the target prompt information are the description components input by the user in the input statement.
[0166] Exemplarily, the user performs a selection operation on the generation component in the multimedia generation interface to implement the multimedia resource generation operation. The terminal responds to the multimedia resource generation operation and sends the target prompt information input by the user in the multimedia generation interface to the server. The server inputs the first multimedia resource and the target prompt information into the artificial intelligence model. The artificial intelligence model processes the input first multimedia resource based on the target prompt information, obtains and outputs the second multimedia resource. The server returns the second multimedia resource to the terminal, and the terminal displays the returned second multimedia resource.
[0167] In some embodiments, the multimedia generation interface further includes at least one configuration component, and this configuration component is used to configure the multimedia resource to be generated or the reference multimedia resource. For example, Figures 3 to 5 As shown in the multimedia generation interface 300, the multimodal editing sub-interface 310 includes multiple configuration components H1 to H4. The configuration components H1 to H3 are respectively used to configure the mode of generating the multimedia resource, the duration and the number of the multimedia resources to be generated. The configuration component H4 is used to configure whether to retain the basic multimedia resource. In some embodiments, the multimodal editing sub-interface 310 includes at least one of the multiple configuration components H1 to H4, or the multimodal editing sub-interface 310 includes other configuration components. Here, the embodiments of the present disclosure do not limit the number and configuration content of the configuration components in the multimodal editing sub-interface 310. The terminal sends the content configured by these configuration components to the server so that the server can use the AI model to generate the second multimedia resource that meets the configuration. The user can configure the multimedia resource to be generated or the reference multimedia resource through these configuration components so as to generate the multimedia resource that meets the configuration, thereby meeting the creative needs of the user and improving the generation efficiency of the multimedia resource.
[0168] The method provided by the embodiments of the present disclosure displays a prompt template in the multimedia generation interface before the user inputs the prompt information. Since the prompt template includes reference statements for generating the prompt information, so that the user can input an input statement that matches the syntax structure of the reference statement in the multimedia generation interface, the target prompt information obtained based on the multimedia resources and the input statement obtained on the multimedia interface can more clearly and accurately describe the user's creativity. Furthermore, the multimedia resources generated based on the target prompt information can better conform to the user's creativity, without the user having to input the prompt information multiple times on the multimedia generation interface and submit the task of regenerating the multimedia resources multiple times, thereby reducing the number of times the user inputs the prompt information and the number of times of regenerating the multimedia resources, and thus improving the generation efficiency of the multimedia resources.
[0169] The multimedia generation interface includes a first display window for displaying the basic multimedia resources input by the user. The first display window is located in the multimodal editing sub-interface. For example, Figures 3 to 5 as shown in the multimedia generation interface 300, the multimodal editing sub-interface 310 includes a first display window R1.
[0170] An input entry for the basic multimedia resources is provided in the multimedia generation interface for inputting the basic multimedia resources. For example, the multimedia generation interface includes a first entry component and / or a second entry component, and both the first entry component and the second entry component are used to input multimedia resources into the first display window in the multimedia generation interface.
[0171] The input entry for the basic multimedia resources is located in the multimodal editing sub-interface, so that when the user performs multimodal editing in the multimodal editing sub-interface, the basic multimedia resources can be input through this input entry. Exemplarily, Figures 3 to 5 as shown in the multimedia generation interface 300, the multimodal editing sub-interface 310 includes a first entry component 50 and a second entry component 60.
[0172] The first entry component is located in the prompt template, that is, the prompt template includes the first entry component. Figures 3 to 5Taking the multimedia generation interface 300 shown as an example, prompt templates 1-3 all include a first entry component 50. By setting the first entry component in the prompt template, on the one hand, it can prompt the user to input basic multimedia resources through the first entry component. During the process of browsing the prompt template, the user can input basic multimedia resources into the first display window through the first entry component in the prompt template, which is convenient for the user to input multimedia resources. On the other hand, the first entry component is set in the reference statement of the prompt template, and the first entry component and the text in the reference statement are described integrally, which can help the user better understand the reference statement in the prompt template. So that when inputting the target prompt information subsequently, the user can input an input statement whose syntax structure matches that of the reference statement, thereby being able to give more accurate prompt information to the artificial intelligence model.
[0173] The first entry component is also used to prompt the input of basic multimedia resources and the resource type of the basic multimedia resources (referred to as the first resource type), so as to Figures 3 to 5 Taking the multimedia generation interface 300 shown as an example, the first entry component 50 includes an icon 51 for representing the first resource type and information 52 for prompting the input of basic multimedia resources. So that the user can input multimedia resources of the first resource type through the first entry component 50.
[0174] The second entry component is located in the area outside the prompt information input box in the multimedia generation interface. The second entry component can be the first display window, that is, the first display window has the function of the second entry component. Exemplarily, as Figures 3 to 5 shown in the multimedia generation interface 300, the first display window R1 is the second entry component 60. In some other embodiments, the second entry component 60 and the first display window R1 can also be different interactive components.
[0175] The first input prompt information is displayed in the second entry component. The first input prompt information is used to prompt the input of basic multimedia resources, the source of the basic multimedia resources, and the maximum specification. For example Figures 3 to 5 the first input prompt information 61 in the second entry component 60 shown. So that the user can, according to the prompt of the first input prompt information 61, input basic multimedia resources into the first display window R1 through the second entry component 60.
[0176] The above Figures 3 to 5 is illustrated by taking the multimedia generation interface 300 including the first entry component 50 and the second entry component 60 as an example. In some other embodiments, the multimedia generation interface includes one of the first entry component 50 and the second entry component 60.
[0177] In some embodiments, the input entry for the basic multimedia resource further includes a third entry component. Exemplarily, the multimedia generation interface further includes a candidate multimedia area, in which a third entry component corresponding to the candidate multimedia resource is displayed, and the third entry component is used to input the corresponding candidate multimedia resource into the first display window. Among them, there is at least one multimedia resource in the candidate multimedia area, and the candidate multimedia resource is a candidate basic multimedia resource. Each candidate multimedia resource corresponds to a third entry component, so that the user can input the corresponding candidate multimedia resource into the first display window by triggering the third entry component. The candidate multimedia resources in the candidate multimedia area can be historically generated multimedia resources (i.e., historical multimedia resources), or they can not be historical multimedia resources.
[0178] Take Figure 6 the multimedia generation interface 400 shown as an example. The difference between the multimedia generation interface 400 and Figure 3 the multimedia generation interface 300 shown is that the multimedia generation interface 400 further includes a candidate multimedia area Z, in which resource identifiers Z2 of at least one candidate multimedia resource Z1 are displayed, and a third entry Z3 is displayed on the cover of each candidate multimedia resource Z1, so that each candidate multimedia resource Z1 corresponds to a third entry Z3. In some embodiments, if the user performs a selection operation on the resource identifier Z2 of any candidate multimedia resource Z1, the terminal responds to the selection operation, displays the cover of the candidate multimedia resource Z1 in the candidate multimedia area Z, and the third entry Z3 is displayed on the cover. Figure 6 Taking the candidate multimedia resources in the candidate multimedia area Z as historical multimedia resources or multimedia resources collected by the user as an example, in some other embodiments, the candidate multimedia resources are neither historical multimedia resources nor multimedia resources collected by the user. Historical multimedia resources are multimedia resources generated by the user historically.
[0179] In some embodiments, the multimedia generation interface further includes a second display window, and there is at least one second display window, and the second display window is located in the multimodal editing sub-interface. As Figure 3 and Figure 4 the multimedia generation interface 300 shown, the multimodal editing sub-interface 310 includes a second display window R2. Since additional multimedia resources are not used in the deletion mode, when the current editing mode is the deletion mode, the multimodal editing sub-interface 310 does not include the second display window R2, such as Figure 5 the multimodal editing sub-interface 310 shown.
[0180] The multimedia generation interface is provided with an input entry for additional multimedia resources, which is used to input additional multimedia resources. For example, the multimedia generation interface includes a fourth entry component and / or a fifth entry component, and both the fourth entry component and the fifth entry component are used to input multimedia resources into the second display window in the multimedia generation interface. As Figure 3 and Figure 4 shown in the multimedia generation interface 300, the multimodal editing sub-interface 310 includes a fourth entry component 70 and at least one fifth entry component 80. Since additional multimedia resources are not used in the deletion mode, when the current editing mode is the deletion mode, the multimodal editing sub-interface 310 includes a fourth entry component 70 and a fifth entry component 80, such as Figure 5 shown in the multimodal editing sub-interface 310.
[0181] The fourth entry component is located in the prompt template, that is, the prompt template includes the fourth entry component. Taking Figure 3 and the figures shown as an example, the prompt templates 1-2 both include the fourth entry component 70. Setting the fourth entry component in the prompt template, on the one hand, can prompt the user to input additional multimedia resources through the fourth entry component. During the process of browsing the prompt template, the user can input additional multimedia resources into the second display window through the fourth entry component in the prompt template, which is convenient for the user to input multimedia resources. On the other hand, the fourth entry component is set in the reference statement of the prompt template, and the fourth entry component and the text in the reference statement are described integrally, which can help the user better understand the reference statement in the prompt template, so that when inputting the target prompt information subsequently, the user can input an input statement with a syntax structure matching that of the reference statement, thereby being able to provide more accurate prompt information to the artificial intelligence model.
[0182] The fourth entry component is also used to prompt the resource type of the input additional multimedia resources (referred to as the second resource type) and to prompt the input of additional multimedia resources. As Figure 3 and Figure 4 shown, the fourth entry component 70 includes an icon 71 representing the second resource type and information 72 for prompting the input of additional multimedia resources. So that the user can input multimedia resources of the second resource type through the fourth entry component 71.
[0183] There is at least one fifth entry component, and the fifth entry component is located in the area outside the prompt information input box in the multimedia generation interface. The fifth entry component can be the second display window, that is, the second display window has the function of the fifth entry component. Exemplarily, as Figure 3 and Figure 4 shown, the second display window R2 is the fifth entry component 80. In some other embodiments, the fifth entry component and the second display window can also be different components.
[0184] A second input prompt message is displayed in the fifth input component. The second input prompt message is used to prompt the input of additional multimedia resources and the sources of the additional multimedia resources. For example Figure 3 and Figure 4 the second input prompt message 81 therein. So that the user, according to the prompt of the second input prompt message 81, can input additional multimedia resources to the second display window R2 through the fifth input component 80.
[0185] In some embodiments, such as Figure 3 and Figure 4 shown, the multimedia generation interface 300 further includes an adding component 90. The adding component 90 is used to add the fifth input component 80 and the second display window R2. The user can perform a selection operation on the adding component 90. In response to this selection operation, the terminal adds the fifth input component 80 and the second display window R2 on the multimedia generation interface 300. So that the user can use the added fifth input component 80 to transmit additional multimedia resources to the added second display window R2, enabling the user to input more additional multimedia resources into the multimedia generation interface for generating new multimedia resources. In some embodiments, in a similar manner, the user can also add a first input component and a first display window on the multimedia generation interface, so that the user can input more basic multimedia resources into the multimedia generation interface for generating new multimedia resources.
[0186] The above Figure 3 and Figure 4 are all illustrated by taking the multimedia generation interface 300 including a fourth input component 70 and a fifth input component 80 as an example. In other embodiments, the multimedia generation interface includes one of the fourth input component and the fifth input component.
[0187] When performing multimodal editing on the multimedia generation interface, the user can use the input entrances introduced above to input reference multimedia resources (such as basic multimedia resources, or basic multimedia resources and additional multimedia resources) to the multimedia generation interface. So that the terminal generates new multimedia resources based on the reference multimedia resources input by the user. Next, a description of this implementation manner will be given in combination with Figure 5 this.
[0188] Figure 7 is a flowchart of another multimedia resource generation method shown according to an exemplary embodiment. Refer to Figure 7 This multimedia resource generation method is applied to a terminal. The terminal can be the terminal 101 in the above-described implementation environment. This multimedia resource generation method includes the following steps.
[0189] In step 701, the terminal displays a multimedia generation interface, and a prompt template is displayed in the multimedia generation interface. The prompt template includes reference statements for generating prompt information, and the prompt information indicates an editing method for multimedia resources.
[0190] Among them, this step 701 is the same as the above step 201 and will not be elaborated here.
[0191] In step 702, the terminal obtains the input first multimedia resource on the multimedia generation interface, and displays the input first multimedia resource in the first display window on the multimedia generation interface.
[0192] Among them, the first multimedia resource is the basic multimedia resource input by the user.
[0193] The user can use any one of the first entry component, the second entry component, and the third entry component to input the first multimedia resource into the first display window, and the terminal displays the first multimedia resource input by the user in the first display window. Next, the methods of inputting the first multimedia resource into the first display window by using the first entry component, the second entry component, and the third entry component will be introduced respectively in combination with the following (1)-(3).
[0194] (1) Input the first multimedia resource by using the first entry component
[0195] In some embodiments, the user can input the first multimedia resource from the prompt information input box on the multimedia generation interface, and the terminal displays the input first multimedia resource in the first display window.
[0196] For example, the user performs a multimedia resource input operation in the prompt information input box, and the multimedia resource input operation indicates inputting the basic multimedia resource. The terminal responds to the multimedia resource input operation in the prompt information input box and displays the input first multimedia resource in the first display window on the multimedia generation interface.
[0197] Exemplarily, the prompt template is located in the prompt information input box, and the prompt template includes the first entry component. The user performs a selection operation on the first entry component in the prompt template to implement the multimedia resource input operation. The terminal responds to the selection operation on the first entry component and displays the input first multimedia resource in the first display window on the multimedia generation interface.
[0198] Taking the current editing method as the addition method as an example, such as Figure 8The multimedia generation interface 300 shown has a prompt template 2 displayed in the prompt information input box 40. The user performs a selection operation on the first entry component 50 in the prompt template 2. The terminal, in response to this selection operation, displays a candidate multimedia panel, and the candidate multimedia panel includes at least one candidate multimedia resource. The user performs a selection operation on any candidate multimedia resource A1 in the candidate multimedia panel. The terminal, in response to this selection operation, inputs the candidate multimedia resource A1 into the first display window R1, and the candidate multimedia resource A1 is displayed in the first display window R1. The candidate multimedia resource A1 is the first input multimedia resource.
[0199] Among them, the candidate multimedia resources in the candidate multimedia panel are candidate basic multimedia resources. This candidate multimedia resource can be a historical multimedia resource or not a historical multimedia resource.
[0200] In the case where the current editing mode is the modification mode or the deletion mode, the user can, in a similar manner, use the first entry component 50 to input the first multimedia resource into the first display window. For example, Figure 9 shows the effect of using the first entry component 50 to input the first multimedia resource into the first display window R1 in the modification mode. Figure 10 shows the effect of using the first entry component 50 to input the first multimedia resource into the first display window R1 in the deletion mode, which will not be elaborated here. In the case where the above-mentioned template is outside the prompt information input box, the user can, in a similar manner, use the first entry component 50 in the prompt template to input the first multimedia resource into the first display window R1, which will not be elaborated here.
[0201] In the case where the first entry component is not displayed in the prompt information input box (such as when the prompt template is outside the prompt information input box, or the prompt template in the prompt information input box has been deleted), the user can perform a specific operation in the prompt information input box to call out the first entry component, and then use the called-out first entry component to input the first multimedia resource into the first display window.
[0202] Exemplarily, the user performs a call-out operation for the first entry component in the prompt information input box. This call-out operation instructs to call out the first entry component. The terminal, in response to the call-out operation for the first entry component in the prompt information input box, displays the first entry component in the prompt information input box. The user performs a selection operation on the called-out component. The terminal, in response to this selection operation, displays the input first multimedia resource in the first display window in the multimedia generation interface (this process is as introduced above), thereby inputting the first multimedia resource into the first display window. In this case, the multimedia resource input operation in the prompt information input box includes the call-out operation for the first entry component and the selection operation for the first entry component.
[0203] In the case of a need for multimedia resource input, the user first calls out the first entry component in the prompt message input box, and then inputs the first multimedia resource into the multimedia generation interface through the called-out first entry component. This can reduce the number of entry components on the multimedia generation interface, simplify the multimedia generation interface, enable the user to quickly understand the functions and operation methods of the multimedia function interface, reduce visual interference, and thus improve the user experience.
[0204] (2) Input the first multimedia resource using the second entry component
[0205] The user performs a selection operation on the second entry component in the multimedia generation interface. In response to this selection operation, the terminal displays the input first multimedia resource in the first display window in the multimedia generation interface.
[0206] Take Figure 8 as an example. Assume that the current editing mode is the addition mode. According to the prompt of the first input prompt message 61 in the second entry component 60, the user clicks on the second entry component 60 to perform a selection operation on the second entry component 60. In response to this click operation, the terminal displays the above-mentioned candidate multimedia panel. The user performs a selection operation on any candidate multimedia resource A1 in the candidate multimedia panel. In response to this selection operation, the terminal inputs the candidate multimedia resource A1 into the first display window R1 and displays the candidate multimedia resource A1 in the first display window R1. Alternatively, the user can also trigger the terminal to display the candidate multimedia panel in other ways. For example, the user performs a drag-and-drop operation of dragging the candidate multimedia resource A1 to the second entry component 60. In response to this drag-and-drop operation, the terminal inputs the candidate multimedia resource A1 into the first display window R1 and displays the candidate multimedia resource A1 in the first display window R1. In the case where the current editing mode is the modification mode or the deletion mode, the user can input multimedia resources into the first display window R1 using the second entry component 60 in a similar manner, which will not be elaborated here.
[0207] (3) Input the first multimedia resource using the third entry component
[0208] In the case where the multimedia generation interface includes a candidate multimedia area, a third entry component corresponding to the candidate multimedia resource is displayed in the candidate multimedia area. The user performs a selection operation on this third entry component. In response to the selection operation on this third entry component, the terminal displays the candidate multimedia resource corresponding to the third entry component in the first display window in the multimedia generation interface.
[0209] Take Figure 6Taking the multimedia generation interface 400 shown as an example, the user selects the third entry component Z3 on the cover of a candidate multimedia resource Z1 in the candidate multimedia area Z, and the terminal responds to the selection operation by inputting the candidate multimedia resource Z1 to which the third entry component Z3 belongs into the first display window R1, and displays the input candidate multimedia resource Z1 in the first display window R1.
[0210] The user can trigger the terminal to input the candidate multimedia resources to the first display window by selecting the third entry component, without the terminal popping up a new panel for the user to select the multimedia resources to be input, which simplifies the process of inputting multimedia resources and improves the efficiency of generating multimedia resources. Figure 6 It is illustrated by taking the candidate multimedia resources as historical multimedia resources as an example, that the "select from historical creations" method in the first input prompt information 61 can be implemented through the third entry component.
[0211] When multiple input entries for basic multimedia resources are displayed in the multimedia generation interface, the user can select a certain input entry according to usage habits to input basic multimedia resources into the first display window, so as to realize multiple media resource input methods for the first display window, thereby improving the human-computer interaction efficiency of the multimedia generation interface.
[0212] After acquiring the first multimedia resource, the terminal uploads the first multimedia resource to the server. The server generates first description information based on the first multimedia resource, and adds the first description component carrying the first description information as a candidate description component to the candidate component panel associated with the user, so that the candidate component panel includes the first description component.
[0213] In step 703, the terminal obtains the fourth multimedia resource input on the multimedia generation interface, and displays the fourth multimedia resource input in the second display window on the multimedia generation interface.
[0214] The fourth multimedia resource is an additional multimedia resource input by the user. The user inputs at least one fourth multimedia resource on the multimedia generation interface, and each fourth multimedia resource is displayed in a second display window.
[0215] The user can use the fourth entry component or the fifth entry component to input the fourth multimedia resource into at least one second display window. Figure 8 and Figure 9In the multimedia generation interface 300 shown, the user uses the fourth entry component 70 to input the fourth multimedia resource A2 and the fourth multimedia resource A3 into two second display windows R2 respectively. Among them, the process of inputting a fourth multimedia resource into the second display window using the fourth entry component can refer to the process of inputting the first multimedia resource into the first display window using the first entry component.
[0216] Alternatively, the user can also use a fifth entry component to input the fourth multimedia resource into the second display window. This process can refer to the process of the user using the second entry component to input the first multimedia resource into the first display window, which will not be elaborated here. Or, using multiple fifth entry components to input the fourth multimedia resource into multiple second display windows will not be elaborated here.
[0217] After obtaining any input fourth multimedia resource, the terminal displays the fourth multimedia resource in a second display window, and the terminal uploads the fourth multimedia resource to the server. The server generates second description information based on the fourth multimedia resource, and uses the second description component as a candidate description component and adds it to the candidate component panel associated with the user, so that the candidate component panel includes the second description component.
[0218] Among them, the second description component is an interactive component carrying the second description information, that is, the second description component includes the second description information, and the second description component is used to describe the fourth multimedia resource. The second description information is the description information of the fourth multimedia resource, used to describe the fourth multimedia resource, and the fourth description information includes the multimedia characteristics and resource identifier of the fourth multimedia resource.
[0219] Step 703 is an optional step. For example, in the case where the current editing mode is the deletion mode, the user will not input the fourth multimedia resource in the multimedia generation interface, so the terminal does not need to execute step 703. After executing step 702, the terminal executes steps 704 and 706. In the case where the current editing mode is the modification mode or the addition mode, if the user inputs the fourth multimedia resource in the multimedia generation interface, the terminal executes step 703. After executing step 703, the terminal executes steps 705 and 706.
[0220] In step 704, the terminal displays the target prompt information in the multimedia generation interface based on the first multimedia resource and the input statement obtained on the multimedia generation interface, and the syntax structure of the input statement matches that of the reference statement.
[0221] After the user finishes inputting the first multimedia resource, the user inputs the target prompt information in the prompt information input box. During the process of inputting the target prompt information, the user can input the first description component in the candidate component panel into the prompt information input box and input text information in the prompt information input box, so that the first description component and the text information input by the user form an input statement.
[0222] The user can input the first description component in the candidate component panel into the prompt information input box through a component addition operation. Among them, the component addition operation can be inputting a specific character (such as @) in the prompt information input box, or a long-press operation on the prompt information input box. Here, the present disclosure does not limit the component addition operation.
[0223] For example, before or during the process of inputting text information in the prompt information input box, the user performs a component addition operation in the prompt information input box. The component addition operation indicates adding a description component in the prompt information input box. The terminal responds to the component addition operation in the prompt information input box and displays the candidate component panel. The user performs a selection operation on the first description component in the candidate component panel. The terminal responds to the selection operation on the first description component and displays the first description component in the prompt information input box. For example, the first description component is displayed at the current input position of the user to realize inputting the first description component in the prompt information input box. After that, the user can continue to input text information at the position after the first description component in the prompt information input box. Of course, the user can also not continue to input text information. During the information input process, the user can input one or more first description components in the prompt information input box in a similar manner, which will not be elaborated here. During the information input process, the terminal displays the text information input by the user and the first description component in the prompt information input box. The text information and the first description component in the prompt information input box form the user's input statement. After the user finishes inputting the input statement, the user performs a selection operation on the generation component in the multimedia generation interface. The terminal responds to this selection operation and uses the input statement in the prompt information input box as the target prompt information.
[0224] Take Figure 11For example, assume that the current editing mode is the deletion mode. After the user inputs the multimedia resource A1 into the first display window R1 and inputs information in the prompt message input box 40, the terminal deletes the prompt template 3 in the prompt message input box 40 in response to the input operation of the user in the prompt message input box 40. The user inputs the text message "Delete" in the prompt message input box 40. Through the component addition operation, the first description component T1 in the candidate component panel is input to the position after the text message "Delete" in the prompt message input box 40. The user continues to input the text message "the girl in", so that the input text message and the first description component T1 form the input statement G1, and the terminal displays the input statement G1 in the prompt message input box 40. The user performs a selection operation on the generation component K, and the terminal, in response to this selection operation, uses the input statement G1 as the target prompt message. Among them, the first description component T1 includes the thumbnail T12 of the multimedia resource A1 and the resource identifier T22, and the thumbnail T12 is the multimedia feature of the multimedia resource A1. Since the editing mode expressed by the input statement G1 is the deletion mode, the input statement G1 matches the reference statement S3.
[0225] In step 705, based on the first multimedia resource, the fourth multimedia resource, and the input statement obtained on the multimedia generation interface, the terminal displays the target prompt message in the multimedia generation interface, and the syntax structure of the input statement matches that of the reference statement.
[0226] Among them, the first multimedia resource is the multimedia resource displayed in the first display window. The fourth multimedia resource is the multimedia resource displayed in the second display window, and there is at least one fourth multimedia resource. The input statement is the statement input by the user in the prompt message input box. The input statement includes the first description information and the second description information. The first description information and the second description information in the input statement can be carried by the first description component and the second description component respectively, that is, the input statement includes the first description component and the second description component.
[0227] Among them, this step 705 is the same as the above step 704. The difference is that in this step 705, when the user inputs the target prompt message, the user can also, through the component addition operation, input the second description component in the candidate component panel into the prompt message input box, so that in the prompt message input box, the first description component, the second description component, and the text message input by the user are mixed into the input statement, and the terminal uses the input statement as the target prompt message.
[0228] Take Figure 12For example, assume that the current editing mode is the addition editing mode. After inputting multimedia resources A1 to A3 into the multimedia generation interface 300, the user inputs the first description component T1, the second description component T2, the second description component T3, and the text information T4 in the prompt information input box 40. The first description component T1, the second description component T2, the second description component T3, and the text information T4 are mixed into the input statement G2. The user performs a selection operation on the generation component K, and the terminal, in response to this selection operation, uses the input statement G2 as the target prompt information. Among them, the second description component T2 includes the thumbnail T21 of the multimedia resource A2 and the resource identifier T22, and the second description component T3 includes the thumbnail T31 of the multimedia resource A3 and the resource identifier T32. The thumbnail T21 is the multimedia feature of the multimedia resource A2, and the thumbnail T31 is the multimedia feature of the multimedia resource A3. Since the editing mode expressed by the input statement G2 is the addition mode, the input statement G2 matches the reference statement S2.
[0229] Assume that the current editing mode is the modification editing mode. As Figure 13 shown, similar to Figure 12 , the first description component T1, the second description component T2, the second description component T3, and the text information T5 in the prompt information input box 40 are mixed into the input statement G3. Since the editing mode expressed by the input statement G3 is the modification mode, the input statement G3 matches the reference statement S1. The terminal, in response to the selection operation on the generation component K, uses the input statement G3 as the target prompt information.
[0230] It should be understood that according to the above steps 704 and 705, the multimedia resources in the first display window and the second display window are both used to generate the target prompt information. The difference is that the resource types of the multimedia resources in the second display window and the multimedia resources in the first display window are different.
[0231] In step 706, the terminal, in response to the multimedia resource generation operation, displays the second multimedia resource generated based on the target prompt information.
[0232] Among them, this step 706 is the same as the above step 203. The difference is that when the user inputs the fourth multimedia resource and the target prompt information includes the second description component, the server inputs the first multimedia resource, each fourth multimedia resource, and the target prompt information into the artificial intelligence model. The artificial intelligence model processes the input first multimedia resource based on the target prompt information and each fourth multimedia resource, and obtains and outputs the second multimedia resource.
[0233] The method provided by the embodiments of the present disclosure, in addition to being able to achieve Figure 2In addition to the beneficial effects achieved by the method provided in the embodiment, different input entries are provided for multimedia resources of different resource types through the multimedia generation interface, so that the user can input multimedia resources of multiple resource types into the multimedia generation interface through different input entries. Subsequently, the terminal performs multimodal editing based on the input multiple multimedia resources to generate new multimedia resources. The above-described description components (such as the first description component or the second description component) include thumbnails of reference multimedia resources. When the user inputs multiple reference multimedia resources and mixes and arranges the target prompt information using multiple description components, the thumbnails in the description components facilitate the user to distinguish different reference multimedia resources without the user having to carefully distinguish different description components, so that the user can quickly mix and arrange the target prompt information, thereby further improving the efficiency of multimedia resource generation and also improving the user experience.
[0234] The above is described by taking the user inputting description components in the prompt information input box as an example. In some other embodiments, after the terminal obtains the reference multimedia resources input by the user, it can also provide description components for describing the reference multimedia resources in the prompt information box for the user to use. Next, in combination with Figure 14 , this implementation method will be introduced.
[0235] Figure 14 is a flowchart of another multimedia resource generation method shown according to an exemplary embodiment. Refer to Figure 14 , this multimedia resource generation method is applied to a terminal, and the terminal can be the terminal 101 in the above-described implementation environment. This multimedia resource generation method includes the following steps.
[0236] In step 1401, the terminal displays a multimedia generation interface, and a prompt template is displayed in the multimedia generation interface. The prompt template includes a reference statement for generating prompt information, and the prompt information indicates the editing method of the multimedia resource.
[0237] Among them, this step 1401 is the same as the above step 201, and will not be elaborated here.
[0238] Step 1402: The terminal obtains the input first multimedia resource on the multimedia generation interface and displays the input first multimedia resource in the first display window on the multimedia generation interface.
[0239] Among them, this step 1402 is the same as the above step 702, and will not be elaborated here.
[0240] Step 1403: The terminal displays a first description component in the prompt information input box based on the first multimedia resource obtained on the multimedia generation interface. The first description component is used to describe the first multimedia resource.
[0241] After the terminal sends the acquired first multimedia resource to the server, the server synchronizes a candidate component panel including the first description component to the terminal. The terminal adds the first description component in the candidate component panel to the prompt information input box in the multimedia interaction interface, so that the first description component is displayed in the prompt information input box.
[0242] Taking the current editing mode as the deletion mode as an example, as Figure 15 shown in the multimedia generation interface 300, after the user inputs the multimedia resource A1 into the first display window R1, the terminal deletes the prompt template 3 in the prompt information input box 40, and the first description component T1 for the first description component is displayed in the prompt information input box 40.
[0243] Step 1404: The terminal displays the acquired input statement in the prompt information input box, and the input statement in the prompt information input box and the first description component form the target prompt information.
[0244] Among them, the input statement is the information input by the user in the prompt information input box.
[0245] Taking the current editing mode as the deletion mode as an example, as Figure 16 shown in the multimedia generation interface 300, the user inputs the text information T6 at a position before or after the first description component T1 in the prompt information input box 40, and the terminal displays the input text information T5 in the prompt information input box 40. In the prompt information input box 40, the text information T6 and the first description component T1 are mixed into the target prompt information G4. Among them, the text information T6 is the input statement.
[0246] The first description component in the prompt information input box supports movement. During the input process of the input statement, the user can move the position of the first description component in the prompt information input box to embed the first description component into a suitable position in the input statement, so that the user can mix the first description component with the input statement to form the target prompt information by moving the first description component.
[0247] Exemplarily, the user performs a movement operation on the first description component in the first description component in the prompt information input box. In response to this movement operation on the first description component, the terminal moves the first description component from the first position in the prompt information input box to the second position in the input statement. Among them, the movement operation indicates moving the first description component from the first position in the prompt information input box to the second position in the prompt information input box. The first position is the current position of the first description component in the prompt information input box, and the second position is any position in the prompt information input box different from the first position.
[0248] Taking the current editing mode as the deletion mode as an example, as Figure 17The multimedia generation interface 300 shown. The terminal displays the first description component T1 at the position x1 in the prompt information input box 40. The user inputs text information T6 (i.e., the input statement) at a position after the position x1 in the prompt information input box 40. The user performs a moving operation F on the first description component T1, and the moving operation F indicates moving the first description component T1 to the position x2 in the text information T6. In response to the moving operation F, the terminal moves the first description component T1 from the position x1 to the position x2, and the first description component T1 and the text information T6 are arranged in a mixed manner to form the target prompt information G4.
[0249] In the case where the current editing mode is the modification mode or the addition mode, after step 1401, the terminal obtains the input fourth multimedia resource on the multimedia generation interface and displays the input fourth multimedia resource in the second display window on the multimedia generation interface. Similar to the above step 1403, based on each fourth multimedia resource obtained on the multimedia generation interface, the terminal displays at least one second description component in the prompt information input box, and each second description component is used to describe a fourth multimedia resource. Similar to the above step 1404, the terminal displays the obtained input statement in the prompt information input box, and the input statement, the first description component, and each second description component in the prompt information input box form the target prompt information. This will not be elaborated here. The second description component in the prompt information input box also supports moving, and the method of moving the second description component in the prompt information input box can refer to the method of moving the first description component in the prompt information input box, which will not be elaborated here.
[0250] During the input process of the input statement, if the user needs more first description components to describe their creativity, the user can add the first description components in the candidate component panel to the prompt information input box to meet the user's creative needs.
[0251] Exemplarily, during the process of the user inputting information in the prompt information input box, the user performs a component addition operation in the prompt information input box, and the terminal displays the candidate component panel in response to the component addition operation in the prompt information input box. The user performs a selection operation on the first description component in the candidate component panel, and the terminal adds the first description component to the prompt information input box in response to the selection operation on the first description component in the candidate component panel (refer to the relevant introduction in the above text).
[0252] Similarly, during the input process of the input statement, if the user needs more second description components to describe their creativity, the user can add the second description components in the candidate component panel to the prompt information input box to meet the user's creative needs.
[0253] In some embodiments, the multimedia generation interface further includes a candidate component area for displaying the description components deleted by the user from the prompt information input box. During the input of the input statement, if the user deletes a certain description component (such as the first description component or the second description component) in the prompt information input box, the terminal displays the deleted description component in the candidate component area without deleting the multimedia resource described by the description component in the multimedia generation interface. Subsequently, when inputting the target prompt information, if the user has a need for the deleted description component, the user can add the description component back to the prompt information input box from the candidate component area through a preset operation, without having to call out the candidate component panel to add the description component from the candidate component panel, which improves the efficiency of adding the deleted description component, thereby increasing the human-computer interaction efficiency and also improving the efficiency of multimedia resource generation.
[0254] Taking the deletion of the first description component as an example, the user performs a deletion operation on the first description component in the prompt information input box. In response to the deletion operation on the first description component in the prompt information input box, the terminal displays the deleted first description component in the candidate component area of the multimedia generation interface. Subsequently, the user can also perform a selection operation on the first description component in the candidate component area. In response to the selection operation on the first description component in the candidate component area, the terminal adds the first description component back to the prompt information input box.
[0255] Taking the current editing mode as the addition editing mode as an example, for the multimedia generation interface 300 shown as Figure 18 in the figure, the user performs a deletion operation on the first description component T1 in the prompt information input box 40. In response to the deletion operation, the terminal deletes the first description component T1 in the prompt information input box 40 and displays the deleted first description component T1 in the candidate component area W. If the user performs a selection operation on the first description component T1 in the candidate component area W, in response to this selection operation, the terminal adds the first description component T1 to the position after the information input in the prompt information input box 40 and deletes the first description component T1 in the candidate component area W. In the case where the current editing mode is the modification mode or the deletion mode, the first description component in the prompt information input box 40 can also be deleted and added back from the candidate component area W in a similar manner.
[0256] In the case where the current editing mode is the modification mode or the deletion mode, the first description component in the prompt information input box can also be deleted and the deleted first description component can be added back to the prompt information input box from the candidate component area in a similar manner. In the case where the current editing mode is the modification mode or the addition mode, the second description component in the prompt information input box can also be deleted and the deleted second description component can be added back to the prompt information input box from the candidate component area W in a similar manner.
[0257] In step 1405, the terminal displays a second multimedia resource generated based on the target prompt information in response to the multimedia resource generation operation.
[0258] Specifically, this step 1405 is the same as step 706 described above and will not be elaborated here.
[0259] The method provided by the embodiments of the present disclosure, in addition to being able to achieve Figure 2 the beneficial effects achieved by the method provided by the embodiments, before inputting the prompt information into the prompt information input box, it also displays a description component (such as a first description component, or a first description component and a second description component) in the prompt information input box, so that when the user inputs the target prompt information, the user can directly input and use the description component in the prompt information input box, without adding a description component from the candidate component panel, thereby simplifying the input process of the prompt information and being able to improve the efficiency of multimedia resource generation.
[0260] In the above embodiment, after displaying and inputting the first multimedia resource into the multimedia generation interface, the terminal deletes the prompt template in the prompt information input box, so that the user can construct the target prompt information in the prompt information input box. In some other embodiments, after obtaining the first multimedia resource from the multimedia generation interface, the first terminal can also update the reference statement in the prompt template based on the first object in the first multimedia resource, so that the updated reference statement includes the object identifier of the first object. Wherein, the first object may be a key object in the first multimedia resource, and the key object is, for example, the main character in the first multimedia resource, and the object identifier of the first object is used to indicate the first object, such as the name of the first object. Or, the first terminal can also update the text information related to the basic multimedia resource in the reference statement and the first entry component in the prompt template to the first description component.
[0261] After updating the reference statement in the prompt template, the user adjusts the updated reference statement in the prompt information input box. Since the modified reference statement includes the object identifier and the first description component for describing the first object, if the user's creativity involves the first object, the user does not need to input the object identifier of the first object in the prompt information input box, and also does not need to input the first description component. The user can slightly adjust the updated reference statement to make the adjusted reference statement describe the user's creativity. Subsequently, the adjusted reference statement can be used as the target prompt information and input to the artificial intelligence model to generate a new multimedia resource, thereby being able to reduce the amount of information input by the user when constructing the target prompt information, improving the user experience, and being able to improve the generation efficiency of the multimedia resource.
[0262] Taking the current editing method as the deletion method as an example, such as Figure 19For the multimedia generation interface 300 shown, initially, in the prompt information input box 40 in the multimedia generation interface 300, a prompt template 3 is displayed. After the user inputs a multimedia resource into the first display window R1, the terminal replaces the first entry component in the prompt template 3 with the first description component T1 and modifies the reference statement S3 in the prompt template 3 to obtain a reference statement S4. The reference statement S4 includes the object identifier "female student" of the first object. If the user's idea happens to be to delete the female students in the multimedia resource A1, the user can use the reference statement 4 as the target prompt information without inputting the target prompt information. Therefore, the efficiency of multimedia resource generation can be improved. In the case where the current editing mode is the addition mode or the deletion mode, the reference statement can also be updated in a similar manner, which will not be elaborated here.
[0263] In some other embodiments, after obtaining the first multimedia resource and the fourth multimedia resource from the multimedia generation interface, the first terminal can also update the reference statement in the prompt template based on the first object in the first multimedia resource and the second object in the fourth multimedia resource, so that the updated reference statement includes the object identifier describing the first object and the object identifier of the second object. Among them, the second object can be the key object in the second multimedia resource. Or, the first terminal can also update the text information related to the basic multimedia resource and the first entry component in the reference statement to the first description component, and the first terminal can also update the text information related to the additional multimedia resource and the second entry component in the reference statement to the second description component.
[0264] After updating the reference statement in the prompt template, the user adjusts the content in the updated reference statement in the prompt information input box. Since the modified reference statement includes the object identifier of the first object, the object identifier of the second object, the first description component, and the second description component, if the user's idea involves the first object and the second object, the user does not need to input the object identifier of the first object and the object identifier of the second object in the prompt information input box, nor does the user need to input the first description component and the second description component. By making fine adjustments to the updated reference statement, the user can make the adjusted reference statement describe the user's idea (similar to Figure 18 which will not be elaborated here), and then the adjusted reference statement can be used as the target prompt information and input to the artificial intelligence model to generate a new multimedia resource, thereby reducing the amount of information input by the user when mixing and arranging the target prompt information, improving the user experience, and also improving the generation efficiency of the multimedia resource.
[0265] In some embodiments, after the terminal displays the input first multimedia resource in the multimedia interaction interface, the user may further replace the input first multimedia resource with a third multimedia resource, where the third multimedia resource is different from the first multimedia resource. In this case, the terminal may update the multimedia features of the first multimedia resource in the first description component to the multimedia features of the third multimedia resource.
[0266] Exemplarily, the user performs a deletion operation on the first multimedia resource displayed in the multimedia interaction interface, and the terminal deletes the first multimedia resource in the multimedia interaction interface in response to the deletion operation. The user inputs the third multimedia resource to the multimedia generation interface in the same way as the first multimedia resource is input, and the terminal displays the third multimedia resource in the first display window in response to obtaining the third multimedia resource on the multimedia generation interface. The terminal updates the multimedia features of the first multimedia resource in the first description component to the multimedia features of the third multimedia resource in response to obtaining the third multimedia resource on the multimedia generation interface. For example, the terminal updates the multimedia features of the first multimedia resource on the first description component in the candidate component panel to the multimedia features of the third multimedia resource. If the first description component is displayed in the prompt message input box in the multimedia generation interface, the terminal updates the multimedia features of the first multimedia resource on the first description component in the prompt message input box to the multimedia features of the third multimedia resource. The first description component with updated multimedia features can describe the third multimedia resource.
[0267] In the case of obtaining the third multimedia resource on the multimedia generation interface, by updating the multimedia features of the first multimedia resource on the first description component to the multimedia features of the third multimedia resource, the updated first description component can be matched with the base multimedia resource currently input by the user, so that the user can use the new first description component to construct the target prompt message subsequently, thereby meeting the user's requirement of updating the base multimedia resource at any time during the multimodal editing process.
[0268] After the terminal displays the fourth multimedia resource in a certain second display window in the multimedia interaction interface, the user may further replace the fourth multimedia resource with a fifth multimedia resource, where the fifth multimedia resource is different from the fourth multimedia resource. In this case, the terminal may also update the multimedia features of the fourth multimedia resource in the second description component to the multimedia features of the fifth multimedia resource in a similar manner to meet the user's requirement of updating the additional multimedia resource at any time during the multimodal editing process, which will not be elaborated here.
[0269] In the case where the base multimedia resource is a video, the present disclosure also provides another method for generating a multimedia resource, which is used to generate a new multimedia resource based on the video. As Figure 20 shown, Figure 20is a flowchart of another method for generating a multimedia resource shown according to an exemplary embodiment. Refer to Figure 20 , this method for generating a multimedia resource is applied to a terminal, which may be terminal 101 in the above-mentioned implementation environment. This method for generating a multimedia resource includes the following steps.
[0270] In step 2001, the terminal displays a multimedia generation interface, which includes a prompt information input box for inputting prompt information that indicates the editing method for the video.
[0271] Among them, this step 2001 is the same as step 201 above. The difference is that in the above text, the prompt information indicates the editing method for the multimedia resource, while in this step 2001, the prompt information indicates the editing method for the video, which is a type of multimedia resource.
[0272] In step 2002, based on the first video obtained on the multimedia generation interface, the terminal displays a first description component in the prompt information input box. The first description component includes the video features and video identifier of the first video.
[0273] Among them, the first video is the basic multimedia resource input by the user to the multimedia generation interface, and it is also the first multimedia resource input by the user during the process of generating the multimedia resource this time. The video features of the first video are the multimedia features of the first video. For example, the thumbnail of the first video or other features that can describe the first video. The video identifier of the first video indicates the first video, and the video identifier can be the resource identifier of the first video, such as "video" or the video number of the first video. This step 2002 is the same as step 1403 above and will not be elaborated here.
[0274] Before step 2002, the terminal obtains the input first video on the multimedia generation interface and displays the input first video in the first display window on the multimedia generation interface. This process is the same as step 702 above, so that the terminal can execute this step 2001 based on the first video in the first display window.
[0275] In some embodiments, the terminal obtains the input target picture on the multimedia generation interface and displays the input target picture in the second display window on the multimedia generation interface (this process is the same as step 703 above). The target picture is the fourth multimedia resource input by the user. For example, when the current editing method is the modification method or the deletion method, the user uses the fourth entry component or the fifth entry component on the multimedia generation interface to input the target picture into the second display window, so as to generate the target multimedia resource based on the target picture, the first video, and the target prompt information later.
[0276] In this case, based on the target picture obtained on the multimedia generation interface, the terminal displays a second description component in the input prompt information box on the multimedia generation interface. The second description component includes a thumbnail of the target picture and a picture identifier to describe the target picture. The picture identifier indicates the target picture, and the picture identifier may be the resource identifier of the fourth multimedia resource when the fourth multimedia resource is a picture. Among them, based on the target picture obtained on the multimedia generation interface, the process of displaying the second description component in the input prompt information box on the multimedia generation interface is the same as the above process of displaying the second description component in the input prompt information box on the multimedia generation interface based on the fourth multimedia resource obtained on the multimedia generation interface.
[0277] In step 2003, the terminal displays the obtained input statement in the prompt information input box, and the input statement in the prompt information input box and the first description component form the target prompt information.
[0278] Among them, this step 2003 is the same as the above step 1404, and will not be elaborated here.
[0279] In some embodiments, when a target picture is input on the multimedia generation interface, the prompt information input box includes a second description component for describing the target picture, and the user can mix and arrange the input statement, the first description component, and the second description component into the target prompt information. For specific reference, see the relevant description in the above step 705, and will not be elaborated here.
[0280] In step 2004, the terminal displays the target multimedia resource generated based on the target prompt information in response to the multimedia resource generation operation.
[0281] Among them, the target multimedia resource is a new multimedia resource generated based on the first video and the target prompt information. The target multimedia resource may be a video or a picture. Here, the present disclosure embodiment does not limit the resource type of the target multimedia resource. In the above method embodiment, when the first multimedia resource is the first video, the target multimedia resource may be the above-mentioned second multimedia resource.
[0282] The method provided by the embodiments of the present disclosure, before the user inputs prompt information in the prompt information input box of the multimedia generation interface, based on the video input by the user on the multimedia generation interface, displays a description component in the prompt information input box. Since the description component includes the video features and video identifier of the input video, the description component can match the input video, so that when the user inputs prompt information, the description component can be used in combination, and the input statement and the description component are mixed to form the target prompt information, making the target prompt information more clearly and accurately describe the user's creativity. Furthermore, the multimedia resource generated based on the target prompt information can meet the user's creativity, thus eliminating the need for the user to input prompt information multiple times on the multimedia generation interface and submit the task of regenerating the multimedia resource multiple times, thereby improving the generation efficiency of the multimedia resource. Additionally, the video feature of the video in the description component is a thumbnail, which is convenient for the user to distinguish the description component matching the first video without the user having to carefully distinguish different description components, so that the user can quickly mix the target prompt information, further improving the generation efficiency of the multimedia resource and enhancing the user experience.
[0283] In some embodiments, all the basic multimedia resources in all the previous alternative technical solutions may be defined as videos, the additional multimedia resources may be defined as pictures, the first multimedia resource may be defined as the first video, the fourth multimedia resource may be defined as the target picture, the second multimedia resource may be defined as the target multimedia resource, the third multimedia resource may be defined as the second video, and the candidate multimedia resources of the basic multimedia resources (such as the candidate multimedia resources in the candidate multimedia area) may be defined as candidate videos, obtaining multiple new alternative solutions. Each new alternative solution can be combined arbitrarily to form alternative embodiments of the present disclosure, which will not be elaborated one by one here. Figure 20
[0284] The related technical solutions for generating videos based on images as a type of multimedia resource do not support: inputting images or videos in the prompt information input box, or, displaying a description component for describing the video or image in the prompt information input box, or, displaying a description component including a thumbnail in the prompt information input box. However, the solution of the present disclosure for generating multimedia resources based on videos supports: inputting images or videos in the prompt information input box, displaying a description component for describing the video or image in the prompt information input box, and displaying a description component including a thumbnail in the prompt information input box. Inputting images or videos in the prompt information input box so that the user can input multimedia resources into the multimedia generation interface in the prompt information input box; displaying a description component in the prompt information input box so that the user can mix and arrange the target prompt information with the description component and the input text information in the prompt information input box to better describe the user's creativity. Displaying a description component including a thumbnail in the prompt information input box so that when the user inputs multiple multimedia resources (such as videos and at least one image), it is easier to distinguish the multimedia resources described by different description components, so that the user can use multiple description components to mix and arrange the target prompt information.
[0285] Any combination of the above all optional technical solutions can form an optional embodiment of the present disclosure, which will not be elaborated here one by one.
[0286] Figure 21 It is a logical structural block diagram of a multimedia resource generation device shown according to an exemplary embodiment. Figure 21 The device 2100 shown is a multimedia resource generation device. Refer to Figure 21 , the device 2100 includes:
[0287] The first display unit 2101 is configured to execute displaying a multimedia generation interface, and display a prompt template in the multimedia generation interface. The prompt template includes a reference statement for generating prompt information, and the prompt information indicates an editing method for multimedia resources;
[0288] The second display unit 2102 is configured to execute, based on the first multimedia resource and the input statement obtained on the multimedia generation interface, display target prompt information in the multimedia generation interface, and the input statement matches the syntax structure of the reference statement;
[0289] The third display unit 2103 is configured to execute in response to a multimedia resource generation operation, display a second multimedia resource generated based on the target prompt information.
[0290] Optionally, the multimedia generation interface includes a prompt information input box for inputting the target prompt information; the second display unit 2102 is configured to execute:
[0291] Based on the first multimedia resource obtained on the multimedia generation interface, a first description component is displayed in the prompt information input box, and the first description component is used to describe the first multimedia resource;
[0292] The input statement entered is displayed in the prompt information input box, and the input statement and the first description component in the prompt information input box form the target prompt information.
[0293] Optionally, the device 2100 further includes:
[0294] A moving unit, configured to perform a moving operation in response to a movement operation on the first description component, and move the first description component from a first position in the prompt information input box to a second position in the input statement.
[0295] Optionally, the first description component includes the multimedia features and resource identifier of the first multimedia resource.
[0296] Optionally, the device 2100 further includes:
[0297] A first update unit, configured to perform an update operation in response to obtaining a third multimedia resource on the multimedia generation interface, and update the multimedia features of the first multimedia resource in the first description component to the multimedia features of the third multimedia resource.
[0298] Optionally, the device 2100 further includes:
[0299] A deletion unit, configured to perform a deletion operation in response to a deletion operation on the first description component in the prompt information input box, and display the deleted first description component in the candidate component area of the multimedia generation interface;
[0300] A first addition unit, configured to perform an addition operation in response to a selection operation on the first description component in the candidate component area, and add the first description component back to the prompt information input box.
[0301] Optionally, the device 2100 further includes:
[0302] A fourth display unit, configured to perform a display operation in response to a component addition operation in the prompt information input box, and display a candidate component panel, where the candidate component panel includes the first description component;
[0303] A second addition unit, configured to perform an addition operation in response to a selection operation on the first description component in the candidate component panel, and add the first description component to the prompt information input box.
[0304] Optionally, the multimedia generation interface includes a plurality of editing components, and different editing components indicate different editing methods for the multimedia resource; the first display unit 2101 is further configured to perform:
[0305] In response to a selection operation on any one of the editing components, in the multimedia generation interface, display the prompt template related to the editing method indicated by the editing component.
[0306] Optionally, the device 2100 further includes:
[0307] A second update unit, configured to perform updating the reference statement in the prompt template based on a first object in the first multimedia resource obtained on the multimedia generation interface.
[0308] Optionally, the device 2100 further includes:
[0309] A third update unit, configured to perform updating the reference statement in the prompt template based on a first object in the first multimedia resource and a second object in the fourth multimedia resource, where both the first multimedia resource and the fourth multimedia resource can be obtained on the multimedia generation interface.
[0310] Optionally, the first display unit 2101 is further configured to perform displaying the prompt template in the prompt information input box in the multimedia generation interface.
[0311] Optionally, the device 2100 further includes:
[0312] A fifth display unit, configured to perform, in response to a multimedia resource input operation in the prompt information input box, displaying the input first multimedia resource in a first display window in the multimedia generation interface.
[0313] Optionally, the prompt template further includes a first entry component, and the first entry component is used to input a multimedia resource into the first display window; the fifth display unit is configured to perform:
[0314] In response to a selection operation on the first entry component, display the input first multimedia resource in a first display window in the multimedia generation interface.
[0315] Optionally, the fifth display unit is configured to perform:
[0316] In response to an invocation operation for the first entry component in the prompt information input box, display the first entry component in the prompt information input box, where the first entry component is used to input a multimedia resource into the first display window;
[0317] In response to a selection operation on the first entry component, display the input first multimedia resource in a first display window in the multimedia generation interface.
[0318] Optionally, the multimedia generation interface further includes a second entry component, which is used to input the multimedia resource into the first display window, and the second entry component is located outside the prompt information input box.
[0319] Optionally, the multimedia generation interface further includes a candidate multimedia area, in which a third entry component corresponding to a candidate multimedia resource is displayed, and the third entry component is used to input the corresponding candidate multimedia resource into the first display window; the apparatus 2100 further includes:
[0320] A sixth display unit, configured to perform, in response to a selection operation on the third entry component, display the candidate multimedia resource in the first display window in the multimedia generation interface, and the candidate multimedia resource in the first display window is the first multimedia resource.
[0321] Optionally, the multimedia generation interface further includes a fourth entry component and / or a fifth entry component. The fourth entry component is located in the prompt template, and the fifth entry component is located outside the prompt information input box. Both the fourth entry component and the fifth entry component are used to input the multimedia resource into a second display window in the multimedia generation interface. The multimedia resource in the second display window is used to generate the target prompt information, and the resource types of the multimedia resource in the second display window and the multimedia resource in the first display window are different.
[0322] Regarding the apparatus in the above embodiments, the specific manners in which each unit performs operations have been described in detail in the embodiments related to the multimedia resource generation method, and will not be elaborated here.
[0323] Figure 22 It is a logical structural block diagram of a multimedia resource generation apparatus shown according to an exemplary embodiment. Figure 22 The shown apparatus 2200 is a multimedia resource generation apparatus. Refer to Figure 22 , the apparatus 2200 includes:
[0324] A first display unit 2201, configured to perform displaying a multimedia generation interface, where the multimedia generation interface includes a prompt information input box for inputting prompt information, and the prompt information indicates an editing method for a video.
[0325] A second display unit 2202, configured to display a first description component in the prompt information input box based on a first video obtained on the multimedia generation interface, where the first description component includes video features and a video identifier of the first video;
[0326] A third display unit 2203, configured to display an obtained input statement in the prompt information input box, where the input statement in the prompt information input box and the first description component form target prompt information;
[0327] A fourth display unit 2204, configured to display a target multimedia resource generated based on the target prompt information in response to a multimedia resource generation operation.
[0328] Optionally, the device 2200 further includes:
[0329] A fifth display unit, configured to display an input first video in a first display window on the multimedia generation interface in response to a multimedia resource input operation in the prompt information input box.
[0330] Optionally, the prompt information input box includes a first entry component for inputting a video into the first display window; the fifth display unit is configured to display the input first video in the first display window on the multimedia generation interface in response to a selection operation on the first entry component.
[0331] Optionally, the fifth display unit is configured to perform:
[0332] In response to an invocation operation on the first entry component in the prompt information input box, display the first entry component in the prompt information input box, where the first entry component is used to input a video into the first display window;
[0333] In response to a selection operation on the first entry component, display the input first video in the first display window on the multimedia generation interface.
[0334] Optionally, the multimedia generation interface further includes a second entry component for inputting a video into the first display window, and the second entry component is located outside the prompt information input box.
[0335] Optionally, the multimedia generation interface further includes a candidate multimedia area, where a third entry component corresponding to a candidate video is displayed in the candidate multimedia area, and the third entry component is used to input the corresponding candidate video into the first display window; the device 2200 further includes:
[0336] A fifth display unit, configured to execute a selection operation in response to the third entry component, and display the candidate video in the first display window in the multimedia generation interface, where the candidate video in the first display window is the first video.
[0337] Optionally, the multimedia generation interface further includes a fourth entry component and / or a fifth entry component. The fourth entry component is located in the prompt information input box, and the fifth entry component is located outside the prompt information input box. Both the fourth entry component and the fifth entry component are used to input pictures into the second display window in the multimedia generation interface, and the pictures in the second display window are used to generate the target prompt information.
[0338] Optionally, the apparatus 2200 further includes:
[0339] A moving unit, configured to execute a moving operation in response to the first description component, and move the first description component from the first position in the prompt information input box to the second position in the input statement.
[0340] Optionally, the apparatus 2200 further includes:
[0341] A first update unit, configured to execute an update operation in response to obtaining a second video on the multimedia generation interface, and update the video feature of the first video in the first description component to the video feature of the second video.
[0342] Optionally, the apparatus 2200 further includes:
[0343] A deletion unit, configured to execute a deletion operation in response to the first description component in the prompt information input box, and display the deleted first description component in the candidate component area of the multimedia generation interface;
[0344] A first addition unit, configured to execute an addition operation in response to the selection operation of the first description component in the candidate component area, and add the first description component back to the prompt information input box.
[0345] Optionally, the apparatus 2200 further includes:
[0346] A sixth display unit, configured to execute a component addition operation in response to the prompt information input box, and display a candidate component panel, where the candidate component panel includes the first description component;
[0347] A second addition unit, configured to execute an addition operation in response to the selection operation of the first description component in the candidate component panel, and add the first description component to the prompt information input box.
[0348] Optionally, the device 2200 further includes:
[0349] A seventh display unit, configured to execute displaying a prompt template in the multimedia generation interface, where the prompt template includes a reference statement for generating the prompt information, and the syntax structure of the input statement matches that of the reference statement.
[0350] Optionally, the multimedia generation interface includes a plurality of editing components, and different editing components indicate different editing methods for the video;
[0351] The seventh display unit is configured to execute, in response to a selection operation on any one of the editing components, displaying the prompt template related to the editing method indicated by the editing component in the multimedia generation interface.
[0352] Optionally, the device 2200 further includes:
[0353] A second update unit, configured to execute updating the reference statement in the prompt template based on a first object in a first video obtained on the multimedia generation interface.
[0354] Optionally, the device 2200 further includes:
[0355] A third update unit, configured to execute updating the reference statement in the prompt template based on the first object in the first video and a second object in a target picture, where the target picture can be obtained on the multimedia generation interface.
[0356] Optionally, the seventh display unit is configured to execute displaying the prompt template in the prompt information input box in the multimedia generation interface.
[0357] Optionally, the prompt template includes a first entry component, and the first entry component is used to input a video into a first display window in the multimedia generation interface.
[0358] Regarding the device in the above embodiments, the specific manners in which each unit performs operations have been described in detail in the embodiments related to the multimedia resource generation method, and will not be elaborated here.
[0359] Figure 23 is a block diagram of the structure of an electronic device shown according to an exemplary embodiment, as Figure 23The electronic device 2300 shown can be configured as the above-mentioned terminal. The electronic device 2300 can be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, or a desktop computer. The electronic device 2300 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0360] Generally, the electronic device 2300 includes a processor 2301 and a memory 2302.
[0361] The processor 2301 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 2301 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 2301 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 2301 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 2301 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0362] The memory 2302 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 2302 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 2302 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 2301 to implement the multimedia resource generation method provided in each embodiment of the present disclosure.
[0363] In some embodiments, the electronic device 2300 may further optionally include: a peripheral device interface 2303 and at least one peripheral device. The processor 2301, the memory 2302, and the peripheral device interface 2303 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 2303 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 2304, a touch display screen 2305, a camera assembly 2306, an audio circuit 2307, and a power supply 2308.
[0364] The peripheral device interface 2303 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 2301 and the memory 2302. In some embodiments, the processor 2301, the memory 2302, and the peripheral device interface 2303 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 2301, the memory 2302, and the peripheral device interface 2303 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0365] The radio frequency circuit 2304 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 2304 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 2304 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 2304 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 2304 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: a metropolitan area network, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 2304 may further include a circuit related to NFC (Near Field Communication), and this disclosure does not limit this.
[0366] The display screen 2305 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 2305 is a touch display screen, the display screen 2305 also has the ability to collect touch signals on or above the surface of the display screen 2305. The touch signals can be input to the processor 2301 as control signals for processing. At this time, the display screen 2305 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 2305, which is disposed on the front panel of the electronic device 2300; in other embodiments, there may be at least two display screens 2305, which are respectively disposed on different surfaces of the electronic device 2300 or are in a foldable design; in some embodiments, the display screen 2305 may be a flexible display screen, which is disposed on a curved surface or a foldable surface of the electronic device 2300. Even further, the display screen 2305 can also be set to an irregular non-rectangular shape, that is, an irregular-shaped screen. The display screen 2305 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0367] The camera module 2306 is used to capture images or videos. Optionally, the camera module 2306 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are respectively any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fused shooting functions. In some embodiments, the camera module 2306 may further include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. A dual-color-temperature flash refers to a combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.
[0368] The audio circuit 2307 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 2301 for processing, or input to the radio frequency circuit 2304 to implement voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 2300. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 2301 or the radio frequency circuit 2304 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 2307 may further include a headphone jack.
[0369] The power supply 2308 is used to supply power to each component in the electronic device 2300. The power supply 2308 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 2308 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0370] In some embodiments, the electronic device 2300 further includes one or more sensors 2310. The one or more sensors 2310 include but are not limited to: an acceleration sensor 2311, a gyroscope sensor 2312, a pressure sensor 2313, an optical sensor 2314, and a proximity sensor 2315.
[0371] The acceleration sensor 2311 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the electronic device 2300. For example, the acceleration sensor 2311 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 2301 can control the touch display screen 2305 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 2311. The acceleration sensor 2311 can also be used for collecting game or user's motion data.
[0372] The gyroscope sensor 2312 can detect the body direction and rotation angle of the electronic device 2300. The gyroscope sensor 2312 can cooperate with the acceleration sensor 2311 to collect the 3D actions of the user on the electronic device 2300. According to the data collected by the gyroscope sensor 2312, the processor 2301 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0373] The pressure sensor 2313 can be disposed on the side frame of the electronic device 2300 and / or the lower layer of the touch display screen 2305. When the pressure sensor 2313 is disposed on the side frame of the electronic device 2300, it can detect the holding signal of the user on the electronic device 2300, and the processor 2301 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 2313. When the pressure sensor 2313 is disposed on the lower layer of the touch display screen 2305, the processor 2301 can control the operable controls on the UI interface according to the pressure operation of the user on the touch display screen 2305. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0374] The optical sensor 2314 is used to collect the ambient light intensity. In one embodiment, the processor 2301 can control the display brightness of the touch display screen 2305 according to the ambient light intensity collected by the optical sensor 2314. Specifically, when the ambient light intensity is high, the display brightness of the touch display screen 2305 is increased; when the ambient light intensity is low, the display brightness of the touch display screen 2305 is decreased. In another embodiment, the processor 2301 can also dynamically adjust the shooting parameters of the camera module 2306 according to the ambient light intensity collected by the optical sensor 2314.
[0375] The proximity sensor 2315, also known as the distance sensor, is usually disposed on the front panel of the electronic device 2300. The proximity sensor 2315 is used to collect the distance between the user and the front of the electronic device 2300. In one embodiment, when the proximity sensor 2315 detects that the distance between the user and the front of the electronic device 2300 is gradually decreasing, the processor 2301 controls the touch display screen 2305 to switch from the lit state to the off state; when the proximity sensor 2315 detects that the distance between the user and the front of the electronic device 2300 is gradually increasing, the processor 2301 controls the touch display screen 2305 to switch from the off state to the lit state.
[0376] Those skilled in the art can understand that Figure 23 the structure shown in does not constitute a limitation on the electronic device 2300, and it may include more or fewer components than shown in the figure, or combine some components, or adopt a different component layout.
[0377] In an exemplary embodiment, there is also provided a computer-readable storage medium including at least one instruction, such as a memory including at least one instruction. The at least one instruction can be executed by a processor in an electronic device to complete the multimedia resource generation method in the above embodiment. Optionally, the computer-readable storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may include ROM (Read-Only Memory), RAM (Random-Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage devices, etc.
[0378] In an exemplary embodiment, there is also provided a computer program product including one or more instructions, and the one or more instructions can be executed by a processor of an electronic device to complete the multimedia resource generation method provided in each of the above embodiments.
[0379] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the multimedia resources involved in this disclosure are all obtained under full authorization.
[0380] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0381] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for generating multimedia resources, characterized in that, Including: Display a multimedia generation interface, the multimedia generation interface includes a prompt information input box for inputting prompt information, and the prompt information indicates an editing method for a video; Based on a first video obtained on the multimedia generation interface, display a first description component in the prompt information input box, and the first description component includes video features and a video identifier of the first video; Display an obtained input statement in the prompt information input box, and the input statement in the prompt information input box and the first description component form target prompt information; In response to a multimedia resource generation operation, display a target multimedia resource generated based on the target prompt information.
2. The multimedia resource generation method according to claim 1, wherein After displaying the multimedia generation interface, the method further includes: In response to a multimedia resource input operation in the prompt information input box, display the input first video in a first display window on the multimedia generation interface.
3. The multimedia resource generation method according to claim 2, wherein The prompt information input box includes a first entry component for inputting a video into the first display window; The step of displaying the input first video in the first display window on the multimedia generation interface in response to a multimedia resource input operation in the prompt information input box includes: In response to a selection operation on the first entry component, display the input first video in the first display window on the multimedia generation interface.
4. The multimedia resource generation method according to claim 2, wherein The step of displaying the input first video in the first display window on the multimedia generation interface in response to a multimedia resource input operation in the prompt information input box includes: In response to an invocation operation on the first entry component in the prompt information input box, display the first entry component in the prompt information input box, and the first entry component is used for inputting a video into the first display window; In response to a selection operation on the first entry component, display the input first video in the first display window on the multimedia generation interface.
5. The multimedia resource generation method according to claim 2, wherein The multimedia generation interface further includes a second entry component for inputting a video into the first display window, and the second entry component is located outside the prompt information input box.
6. The multimedia resource generation method according to claim 2, wherein The multimedia generation interface further includes a candidate multimedia area, and a third entry component corresponding to a candidate video is displayed in the candidate multimedia area, and the third entry component is used for inputting the corresponding candidate video into the first display window; After displaying the multimedia generation interface, the method further includes: In response to a selection operation on the third entry component, display the candidate video in the first display window on the multimedia generation interface, and the candidate video in the first display window is the first video.
7. The multimedia resource generation method according to any one of claims 1-6, characterized in that The multimedia generation interface further includes a fourth entry component and / or a fifth entry component. The fourth entry component is located in the prompt information input box, and the fifth entry component is located outside the prompt information input box. Both the fourth entry component and the fifth entry component are used to input pictures into a second display window in the multimedia generation interface, and the pictures in the second display window are used to generate the target prompt information.
8. The multimedia resource generation method according to any one of claims 1-6, characterized in that, After displaying the obtained input statement in the prompt information input box, the method further includes: In response to a moving operation on the first description component, moving the first description component from a first position in the prompt information input box to a second position in the input statement.
9. The multimedia resource generation method according to any one of claims 1-6, characterized in that After displaying the first description component in the prompt information input box, the method further includes: In response to obtaining a second video on the multimedia generation interface, updating the video feature of the first video in the first description component to the video feature of the second video.
10. The multimedia resource generation method according to any one of claims 1-6, characterized in that, After displaying the first description component in the prompt information input box, the method further includes: In response to a deletion operation on the first description component in the prompt information input box, displaying the deleted first description component in the candidate component area of the multimedia generation interface; In response to a selection operation on the first description component in the candidate component area, adding the first description component back to the prompt information input box.
11. The multimedia resource generation method according to any one of claims 1-6, characterized in that, After displaying the first description component in the prompt information input box, the method further includes: In response to a component addition operation in the prompt information input box, displaying a candidate component panel, and the candidate component panel includes the first description component; In response to a selection operation on the first description component in the candidate component panel, adding the first description component to the prompt information input box.
12. The multimedia resource generation method according to any one of claims 1-6, characterized in that, Before displaying the first description component in the prompt information input box based on the first video obtained on the multimedia generation interface, the method further includes: Displaying a prompt template in the multimedia generation interface, where the prompt template includes a reference statement for generating the prompt information, and the syntax structure of the input statement matches that of the reference statement.
13. The multimedia resource generation method according to claim 12, wherein The multimedia generation interface includes a plurality of editing components, and different editing components indicate different editing methods for the video; The displaying the prompt template in the multimedia generation interface includes: In response to a selection operation on any editing component, displaying, in the multimedia generation interface, the prompt template related to the editing method indicated by the editing component.
14. The multimedia resource generation method according to claim 12, wherein After displaying the prompt template in the multimedia generation interface, the method further includes: Updating the reference statement in the prompt template based on the first object in the first video obtained on the multimedia generation interface.
15. The multimedia resource generation method according to claim 12, characterized in that, After displaying the prompt template in the multimedia generation interface, the method further includes: Updating the reference statement in the prompt template based on the first object in the first video and the second object in the target picture, where the target picture can be obtained on the multimedia generation interface.
16. The multimedia resource generation method according to claim 12, wherein Displaying a prompt template in the multimedia generation interface includes: Displaying the prompt template in the prompt information input box in the multimedia generation interface.
17. The multimedia resource generation method according to claim 12, wherein The prompt template includes a first entry component for inputting a video into a first display window in the multimedia generation interface.
18. A multimedia resource generation device, characterized in that Comprising: A first display unit configured to display a multimedia generation interface, the multimedia generation interface including a prompt information input box for inputting prompt information indicating an editing method for a video; A second display unit configured to display a first description component in the prompt information input box based on a first video obtained on the multimedia generation interface, the first description component including video features and a video identifier of the first video; A third display unit configured to display an obtained input statement in the prompt information input box, the input statement in the prompt information input box and the first description component constituting target prompt information; A fourth display unit configured to display a target multimedia resource generated based on the target prompt information in response to a multimedia resource generation operation.
19. An electronic device, characterized in that, Comprising: One or more processors; One or more memories for storing executable instructions of the one or more processors; Wherein, the one or more processors are configured to execute the instructions to implement the multimedia resource generation method according to any one of claims 1 to 17.
20. A computer-readable storage medium, characterized in that, When at least one instruction in the computer-readable storage medium is executed by one or more processors of an electronic device, the computer device is enabled to execute the multimedia resource generation method according to any one of claims 1 to 17.
21. A computer program product, characterized in that, Including one or more instructions, the one or more instructions are executed by one or more processors of an electronic device, enabling the computer device to execute the multimedia resource generation method according to any one of claims 1 to 17.
Citation Information
Patent Citations
Multimedia input method, device and system based on mobile terminal
CN102541452A
Knowledge question and answer method, device and equipment and storage medium
CN116561277A
Information interaction method and device and readable storage medium
CN116962626A
Content display method and device, electronic equipment and storage medium
CN116975330A
Information identification method and device based on artificial intelligence and computer equipment
CN117520544A
Cited By
Interaction method and device, electronic equipment and storage medium
CN120897092A
Interaction method and apparatus, electronic device, and storage medium
CN120897092B
Media content processing method and device, equipment, storage medium and program product
CN121284345A